← Search

Yifei Huang

38 accepted papers

2026

Beyond “Made with AI”: Visualizing Provenance Density to Mitigate the Transparency Penalty

IJCAI 2026

As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary "Made with AI" labels r

Cited by 0Scholar
2026

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

CVPR 2026

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely provide the fine-grained, grounded evidence needed to rigorously

Cited by 0SourceScholar
2026

Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy Unfolding

AAAI 2026technical

Learning procedural-aware video representations is a key step towards building agents that can reason about and execute complex tasks. Existing methods typically address this problem by aligning visual content with textual descriptions at the task and step levels to inject procedural semantics into

Cited by 0SourcePDFScholar
2026

Multi-speaker Attention Alignment for Multimodal Social Interaction

CVPR 2026

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures.While Multimodal Large Language Models (MLLMs) are natural candidates, simply adding visual inputs yields surprisingly inconsi

Cited by 0SourcecodeScholar
2026

UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking

CVPR 2026

Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening.However, modeling the listener is exceptionally challenging: direct audio-driven training fails, producing stiff, static listening motions. This

Cited by 0SourcecodeScholar
2025

Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition

ICCV 2025poster

Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recent efforts to incorporate additional modalities, particularly text. However, the subtle variations in human posture, moti…

Cited by 0SourcePDFScholar
2025

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

ICLR 2025poster

The existing video understanding benchmarks for multimodal large language models (MLLMs) mainly focus on short videos. The few benchmarks for long video understanding often rely on multiple-choice questions (MCQs). Due to the limitations of MCQ evaluations and the advanced reasoning abilities of MLL…

Cited by 5SourcePDFScholar
2025

EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

ICLR 2025poster

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and…

Cited by 0SourcePDFScholar
2025

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

NeurIPS 2025poster

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences. Despite rapid progress in multimodal large language models (MLLMs)…

Cited by 0SourceScholar
2025

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

NeurIPS 2025poster

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current multimodal large language models (MLLMs), which excel at vis…

Cited by 0SourceScholar
2025

Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance

ICCV 2025poster

This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IM…

Cited by 0SourcePDFScholar
2025

Egocentric Object-Interaction Anticipation with Retentive and Predictive Learning

IJCAI 2025

Egocentric object-interaction anticipation is critical for applications like augmented reality and robotics, but existing methods struggle with misaligned egocentric encoding, insufficient supervision, and underutilized historical context. These limitations stem from a lack of focus on retention, i.

Cited by 0SourcePDFScholar
2025

Imitation Learning with Limited Actions via Diffusion Planners and Deep Koopman Controllers

ICRA 2025

Recent advances in diffusion-based robot policies have demonstrated significant potential in imitating multi-modal behaviors. However, these approaches typically require large quantities of demonstration data paired with corresponding robot action labels, creating a substantial data collection burde

Cited by 4SourcecodeScholar
2025

Learning Streaming Video Representation via Multitask Training

ICCV 2025poster

Understanding continuous video streams plays a fundamental role in real-time applications, including embodied AI and autonomous driving. Unlike offline video processing, streaming video understanding requires the ability to process video streams frame by frame, preserve historical information, and m…

Cited by 0SourcePDFScholar
2025

MAGRET: Machine-generated Text Detection with Rewritten Texts

COLING 2025main

With the quick advancement in text generation ability of Large Language Mode(LLM), concerns about the misuse of machine-generated content have grown, raising potential violations of legal and ethical standards. Some existing studies concentrate on detecting machine-generated text in open-source mode…

Cited by 0SourcePDFScholar
2025

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

ICLR 2025poster

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate d…

2025

SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training

ICLR 2025poster

We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully…

2025

TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation

ICML 2025poster

Text-to-image (T2I) generation has made remarkable progress in producing high-quality images, but a fundamental challenge remains: creating backgrounds that naturally accommodate text placement without compromising image quality. This capability is non-trivial for real-world applications like graph…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

CVPR 2024poster

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability we introduce EgoExoLearn a large-scale dataset that emulates the human demonstration following process in which ind…

2024

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

ECCV 2024poster

"We introduce , a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue. Our core design is a progressive training approach that unifies the masked video modeling, crossmodal contrastive learning, and…

2024

Retrieval-Augmented Egocentric Video Captioning

CVPR 2024poster

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper (1) we develop EgoI…

Cited by 38SourcePDFScholar
2023

3D Segmenter: 3D Transformer based Semantic Segmentation via 2D Panoramic Distillation

ICLR 2023poster

Recently, 2D semantic segmentation has witnessed a significant advancement thanks to the huge amount of 2D image datasets available. Therefore, in this work, we propose the first 2D-to-3D knowledge distillation strategy to enhance 3D semantic segmentation model with knowledge embedded in the latent…

Cited by 4SourcePDFScholar
2023

Memory-and-Anticipation Transformer for Online Action Understanding

ICCV 2023poster

Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal modeling for memory dependency. Nevertheless, an obvious weakness of this paradigm is that it can only model limited his…

Cited by 43PDFcodeScholar
2023

Pretraining Language Models with Text-Attributed Heterogeneous Graphs

EMNLP 2023long findings

In many real-world scenarios (e.g., academic networks, social platforms), different types of entities are not only associated with texts but also connected by various relationships, which can be abstracted as Text-Attributed Heterogeneous Graphs (TAHGs). Current pretraining tasks for Language Models…

Cited by 0SourcecodeScholar
2023

Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction

CVPR 2023poster

The Multiplane Image (MPI), containing a set of fronto-parallel RGBA layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed structure limits the performance, especially for surfaces imaged at oblique angles. We introduce the Structural MPI (S-MPI),…

Cited by 10SourcePDFScholar
2023

Weakly Supervised Temporal Sentence Grounding With Uncertainty-Guided Self-Training

CVPR 2023poster

The task of weakly supervised temporal sentence grounding aims at finding the corresponding temporal moments of a language description in the video, given video-language correspondence only at video-level. Most existing works select mismatched video-language pairs as negative samples and train the m…

Cited by 31SourcePDFScholar
2022

CLRNet: Cross Layer Refinement Network for Lane Detection

CVPR 2022poster

Lane is critical in the vision navigation system of the intelligent vehicle. Naturally, lane is a traffic sign with high-level semantics, whereas it owns the specific local pattern which needs detailed low-level features to localize accurately. Using different feature levels is of great importance f…

Cited by 259PDFcodeScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Interact Before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition

CVPR 2022poster

Unsupervised domain adaptive video action recognition aims to recognize actions of a target domain using a model trained with only out-of-domain (source) annotations. The inherent complexity of videos makes this task challenging but also provides ground for leveraging multi-modal inputs (e.g., RGB,…

Cited by 49PDFScholar
2021

Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling

AAAI 2021technical

Visual storytelling is a task of generating relevant and interesting stories for given image sequences. In this work we aim at increasing the diversity of the generated stories while preserving the informative content from the images. We propose to foster the diversity and informativeness of a gener…

Cited by 49SourcePDFScholar
2021

FACIAL: Synthesizing Dynamic Talking Face With Implicit Attribute Learning

ICCV 2021poster

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and eye blinks that are in-sync with the input audio signal. We…

Cited by 161PDFcodeScholar
2021

Goal-Oriented Gaze Estimation for Zero-Shot Learning

CVPR 2021poster

Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial f…

Cited by 172PDFcodeScholar
2021

Precise Multi-Modal In-Hand Pose Estimation using Low-Precision Sensors for Robotic Assembly

ICRA 2021poster

In industrial assembly tasks, the in-hand pose of grasped objects needs to be known with high precision for subsequent manipulation tasks such as insertion. This problem (in-hand-pose estimation) has traditionally been addressed using visual recognition or tactile sensing. On the one hand, while vis…

Cited by 36SourceScholar
2020

Learn to Recover Visible Color for Video Surveillance in a Day

ECCV 2020poster

In silicon sensors, the interference between visible and near-infrared (NIR) signals is a crucial problem. For all-day video surveillance, commercial camera systems usually adopt auxiliary NIR cut filter and NIR LED illumination to selectively block or enhance NIR signal according to the surrounding…

2018

Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition

ECCV 2018poster

We present a new computational model for gaze prediction in egocentric videos by exploring patterns in temporal shift of gaze fixations (attention transition) that are dependent on egocentric manipulation tasks. Our assumption is that the high-level context of how a task is completed in a certain wa…