← Search

Zihui Xue

20 accepted papers

2025

REG: Rectified Gradient Guidance for Conditional Diffusion Models

ICML 2025poster

Guidance techniques are simple yet effective for improving conditional generation in diffusion models. Albeit their empirical success, the practical implementation of guidance diverges significantly from its theoretical motivation. In this paper, we reconcile this discrepancy by replacing the scaled…

Cited by 0SourcePDFScholar
2025

Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning

CVPR 2025poster

Egocentric and exocentric perspectives of human action differ significantly, yet overcoming this extreme viewpoint gap is critical for applications in augmented reality and robotics. We propose ViewpointRosetta, an approach that unlocks large-scale unpaired ego and exo video data to learn clip-level…

Cited by 0SourcePDFScholar
2025

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

NeurIPS 2025poster

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application to video understanding remains underexplored. This paper pr…

Cited by 0SourceScholar
2024

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

ECCV 2024oral

"Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio during training, yet many sounds happen off-screen and have weak…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness

NeurIPS 2024poster

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made in video editing recently, these models often fall short in handling the intric…

Cited by 8SourcePDFScholar
2024

Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos

ECCV 2024poster

"We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this end, we propose a generative framework called Exo2Ego that dec…

Cited by 18SourcePDFScholar
2023

Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment

NeurIPS 2023poster

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to learning view-invariant features from paired synchronized vi…

Cited by 37SourcePDFScholar
2023

The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation

ICLR 2023top-5%

Crossmodal knowledge distillation (KD) extends traditional knowledge distillation to the area of multimodal learning and demonstrates great success in various applications. To achieve knowledge transfer across modalities, a pretrained network from one modality is adopted as the teacher to provide su…

2022

Co-Advise: Cross Inductive Bias Distillation

CVPR 2022poster

The inductive bias of vision transformers is more relaxed that cannot work well with insufficient data. Knowledge distillation is thus introduced to assist the training of transformers. Unlike previous works, where merely heavy convolution-based teachers are provided, in this paper, we delve into th…

Cited by 82PDFcodeScholar
2021

Anytime Depth Estimation with Limited Sensing and Computation Capabilities on Mobile Devices

CoRL 2021poster

Depth estimation is a safety critical and energy sensitive method for environment sensing. However, in real applications, the depth estimation may be halted at any time, due to the random interruptions or low energy capacity of battery when using powerful sensors like 3D LiDAR. To address this probl…

Cited by 4SourceScholar
2021

On Feature Decorrelation in Self-Supervised Learning

ICCV 2021poster

In self-supervised representation learning, a common idea behind most of the state-of-the-art approaches is to enforce the robustness of the representations to predefined augmentations. A potential issue of this idea is the existence of completely collapsed solutions (i.e., constant features), which…

Cited by 236PDFcodeScholar
2021

What Makes Multi-Modal Learning Better than Single (Provably)

NeurIPS 2021poster

The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is aggregated. Recently, joining the success of deep learning, there is an influential line of work on deep multi-modal le…

Cited by 340SourcePDFScholar