Yini Huang

Hello! This is Yini Huang~

About

My name is Yini Huang(黄依妮). I am an MPhil student in Data Science and Analytics at The Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Jiaheng Wei. Before that, I received my Bachelor degree from Donghua University in 2025.

I am expected to graduate from HKUST(GZ) in May 2027 and actively looking for PhD openings for Fall 2027.

My current interests include trustworthy AI and multimodal agent. Recently, I have also been exploring Embodied-AI. Welcome to reach out for potential discussions and research collaborations.

Experience

Education

2025 - 2027

MPhil in Data Science and Analytics

The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China

2021 - 2025

BSc in Statistics

Donghua University, Shanghai, China

2021 - 2025

BFA in Fashion and Accessory Design

Donghua University, Shanghai, China

Work

2025.6 - 2025.9

AI Data Operation Intern

AI Data Service and Operation Department

ByteDance, Shanghai, China

2024.11 - 2025.5

Data Analysis Intern

Data and Technology Department

Publicis Groupe, Shanghai, China

Publications

What Your Posts Reveal paper figure
Paper figure · assets/publications/posts-reveal.jpg

2026

What Your Posts Reveal: A Benchmark and Agentic Framework for User-Level Privacy Leakage on Social Media

Zifan Peng*, Yini Huang*, Aiwen Lu*, Qiming Ye, Peixian Zhang, Jingyi Zheng, Yule Liu, Xuechao Wang, Xinlei He†, Jiaheng Wei† ( * Equal contribution.† Corresponding author.)

arXiv

Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulative and cross-post: cues that appear harmless in isolation may jointly expose a user's home, workplace, or routine. However, current research lacks a unified benchmark for user-level multimodal privacy leakage and an evaluation metric that captures exposure severity beyond binary accuracy.

To address these gaps, we propose SopriBench, a synthetic benchmark guided by leakage patterns abstracted from a private reference corpus of Rednote and Instagram accounts, covering 50 user profiles and 1,569 images with attributes, contextual sensitivity, granularity, leakage type, inference difficulty, and supporting evidence. We further introduce the Privacy Exposure Score (PES), which weights value granularity by contextual sensitivity. Inspired by abductive reasoning, we introduce Argus, a training-free agentic framework for cumulative leakage inference. Argus forms hypotheses from accumulated evidence, verifies supporting evidence, and aggregates cross-post cues into privacy profiles, achieving 0.55 PES, a 25% improvement over the strongest baseline, with the largest gain on cross-post leakage.

UniCA paper figure
Paper figure · assets/publications/unica.jpg

2026

UniCA: Bi-directional Cross-Attention with Positive Similarity Loss for Robust Multi-Modal Retrieval

Yini Huang, Wenlong Zhang

arXiv

Multi-modal retrieval has become increasingly critical for handling the growing volume of integrated visual-textual data in real-world applications, but existing frameworks rely on implicit fusion via text encoder self-attention, limiting explicit cross-modal semantic alignment.

To address this gap, this paper proposes UniCA (Unified Cross-Attention Encoder), a multi-modal retrieval model with four key innovations: 1) a bi-directional cross-attention (Bi-CA) block that enables active semantic exchange between visual and textual tokens prior to concatenation, capturing inter-modal correlations more efficiently. 2) a Positive Similarity Loss that optimizes absolute semantic proximity between query and positive candidate embeddings. 3) a streamlined dataset UMR-S10 (Universal Multimodal Retrieval Sample 10%) to reduce computational costs while retaining semantic diversity and task representativeness. 4) an experimental validation on the WebQA benchmark demonstrates that UniCA outperforms the baseline model across Hybrid and Image-Text tasks, achieving improvements of up to 4.09% in Recall@5, 3.28% in Recall@10, and 3.96% in MRR@1 for the hybrid task. UniCA provides an efficient and robust solution for multi-modal retrieval, lowering deployment barriers through its lightweight dataset and enhanced fusion mechanism.

Projects

August 2026 – Current

High-Quality Text-to-Image SFT Data Construction

Work In Progress

Background: Text-to-Image(T2I) generative models demand high-quality image-text pairs for SFT, yet web-scale data suffer from misalignment, hallucinated content, and noisy captions.

Key Contributions:

  1. Built a large-scale data curation pipeline over existing web-scale datasets, performing systematic quality filtering to remove misaligned and low-fidelity samples.
  2. Developed a semantic consistency verification framework that decomposes captions into atomic assertions and cross-checks them against image evidence, enabling fine-grained diagnosis of hallucination and relational errors.

Jan 2026 – Current

Emotional Multimodal Agent for Companion Robots

GitHub

Background: This research focuses on building an emotional multimodal agent for indoor companion robot, providing long-term, personalized companion for people who need emotion support.

Key Contributions:

  1. Developed a mobile companion agent with long-term memory and proactive event reminders, providing natural, always-available conversation and personalized emotional support in daily life.
  2. Integrated robot sensor video/audio streams to infer environmental context and user state (facial expressions, body movements, and interaction cues), achieving understanding of user states.
  3. Built the agent as the physical robot's digital avatar, sharing unified memory, personality, and reasoning core across physical and digital domains, so users can keep chatting via phone even when away from the robot.

Sep 2025 – Dec 2025

Avatar-Based IELTS Speaking Assessment and Practice

GitHub

Background: Existing IELTS oral training lacks immersive human-like interaction and standardized automated scoring, hence building avatar-based intelligent practice system aligned with official examination standards.

Key Contributions:

  1. Co-developed a multimodal AI digital human-based IELTS Speaking Simulator for Part 1-3 practice, integrating LLM, ASR/TTS, and virtual avatar technology to realize realistic voice interaction and intelligent question generation.
  2. Constructed an automatic assessment system based on official IELTS criteria, completing multi-dimensional scoring and detailed feedback report generation for oral responses.