← Search

Angela Yao

89 accepted papers

2026

AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation

AAAI 2026technical

Optimization‐based text‑to‑3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "seman

Cited by 0SourcePDFScholar
2026

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

ICML 2026poster

Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and evicting old key-value(KV) caches is required for unbounded streams. Secondly, due to the high cost of collecting and tra…

Cited by 0SourceScholar
2026

Ego-Grounding for Personalized Question-Answering in Egocentric Videos

CVPR 2026

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce MyEgo, the first egocentric VideoQA dataset designed to evalua

Cited by 0SourcecodeScholar
2026

HumanBA: Human-Aware Bundle Adjustment via Global Human-Camera Decoupling

CVPR 2026

Recovering global human and camera motion from monocular video is essential for world-coordinate human reconstruction but remains challenging due to entangled motions in image space. Traditional SLAM methods estimate monocular camera motion but fail in scenes dominated by foreground objects such as

Cited by 0SourcecodeScholar
2026

InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression

ICLR 2026oral

Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which rigidly compress all content at a fixed rate, leading to redunda…

Cited by 0SourcecodeScholar
2026

Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing

ICLR 2026poster

Textured 3D morphing seeks to generate smooth and plausible transitions between two 3D assets, preserving both structural coherence and fine-grained appearance. This ability is crucial not only for advancing 3D generation research but also for practical applications in animation, editing, and digita…

Cited by 0SourcecodeScholar
2026

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

ICML 2026poster

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 7K+ timestamped questions for diagnosing User-centric Continual Spatial intelligence in egocentric video streams. UCS-Bench targets a new problem that emphasizes dynamic spatial reasoning, long-term memory, …

Cited by 0SourceScholar
2026

Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph Optimization

CVPR 2026

Learning-based structure-from-motion methods such as ACE-Zero have demonstrated strong performance in estimating camera poses and scene coordinates from unordered image collections without requiring ground truth supervision. However, the lack of global and multi-view consistency constraints in ACE-Z

Cited by 0SourceScholar
2026

LightAVSeg: Lightweight Audio-Visual Segmentation

ICML 2026poster

Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic computational cost, limiting their suitability for resource efficient deployment. Most efficiency oriented methods focu…

Cited by 0SourceScholar
2026

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

CVPR 2026

Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV-caching stores the Key-Value (KV) of the historical tokens via LLM prefill and enables more efficient streaming QA. However, existing methods cache every one or t

Cited by 0SourceScholar
2026

On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have advanced open-world action understanding and can be adapted as generative classifiers for closed-set settings by autoregressively generating action labels as text. However, this approach is inefficient, and shared subwords across action labels introduce…

Cited by 0SourcecodeScholar
2026

TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction

ICRA 2026poster

Pre-defined 3D object templates are widely used in 3D reconstruction of hand-object interactions. However, they often require substantial manual efforts to capture or source, and inherently restrict the adaptability of models to unconstrained interaction scenarios, e.g., heavily-occluded objects. To…

2026

The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy

CVPR 2026

Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challenging. We identify a key underlying cause: attention collapse in existing attention sharing mechanisms, where either posi

Cited by 0SourcecodeScholar
2026

VA-p: Variational Policy Alignment for Pixel-Aware Autoregressive Generation

CVPR 2026

Autoregressive (AR) visual generation relies on tokenizers to map images to and from discrete sequences. However, tokenizers are trained to reconstruct clean images from ground-truth tokens, while AR generators are optimized only for token likelihood. This misalignment leads to generated token seque

Cited by 0SourcecodeScholar
2026

reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency Regularization

ICLR 2026poster

Visual autoregressive (AR) generation offers a promising path toward unifying vision and language models, yet its performance remains suboptimal against diffusion models. Prior work often attributes this gap to tokenizer limitations and rasterization ordering. In this work, we identify a core bottle…

Cited by 0SourceScholar
2025

A Constrained Optimization Approach for Gaussian Splatting from Coarsely-posed Images and Noisy Lidar Point Clouds

ICCV 2025poster

3D Gaussian Splatting (3DGS) is a powerful reconstruction technique; however, it requires initialization from accurate camera poses and high-fidelity point clouds. Typically, the initialization is taken from Structure-from-Motion (SfM) algorithms; however, SfM is time-consuming and restricts the app…

Cited by 0SourcePDFScholar
2025

Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimation

CVPR 2025poster

Recent synthetic 3D human datasets for the face, body, and hands have pushed the limits on photorealism. Face recognition and body pose estimation have achieved state-of-the-art performance using synthetic training data alone, but for the hand, there is still a large synthetic-to-real gap. This pape…

2025

Context-Enhanced Memory-Refined Transformer for Online Action Detection

CVPR 2025poster

Online Action Detection (OAD) detects actions in streaming videos using past observations. State-of-the-art OAD approaches model past observations and their interactions with an anticipated future. The past is encoded using short- and long-term memories to capture immediate and long-range dependenci…

2025

EgoBlind: Towards Egocentric Visual Assistance for the Blind

NeurIPS 2025poster

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It…

Cited by 0SourcecodeScholar
2025

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

CVPR 2025poster

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are d…

2025

ExtPose: Robust and Coherent Pose Estimation by Extending ViTs

ICML 2025poster

Vision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel…

Cited by 0SourcePDFScholar
2025

Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen Categories

ICCV 2025poster

View-Guided Point Cloud Completion (VG-PCC) aims to reconstruct complete point clouds from partial inputs by referencing single-view images. While existing VG-PCC models perform well on in-class predictions, they exhibit significant performance drops when generalizing to unseen categories. We identi…

Cited by 0SourcePDFScholar
2025

Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh Recovery

ICCV 2025poster

Accurate camera motion estimation is essential for recovering global human motion in world coordinates from RGB video inputs. While SLAM is widely used for estimating camera trajectory and point cloud, monocular SLAM does so only up to an unknown scale factor. Previous works estimate the scale facto…

2025

Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object Categories

ICCV 2025poster

Language-guided Affordance Segmentation (LASO) aims to identify actionable object regions based on text instructions. At the core of its practicality is learning generalizable affordance knowledge that captures functional regions across diverse objects. However, current LASO solutions struggle to ex…

2025

On the Consistency of Video Large Language Models in Temporal Comprehension

CVPR 2025poster

Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency -- a key indicator for robustness and trustworthiness of…

2025

Streaming VideoLLMs for Real-Time Procedural Video Understanding

ICCV 2025poster

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens -- verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded…

Cited by 9SourcePDFScholar
2025

The Devil is in the Spurious Correlations: Boosting Moment Retrieval with Dynamic Learning

ICCV 2025poster

Given a textual query along with a corresponding video, the objective of moment retrieval aims to localize the moments relevant to the query within the video. While commendable results have been demonstrated by existing transformer-based approaches, predicting the accurate temporal span of the targe…

2025

Visual Intention Grounding for Egocentric Assistants

ICCV 2025poster

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are egocentric, and objects may be referred to implicitly through needs a…

2024

AID: Attention Interpolation of Text-to-Image Diffusion

NeurIPS 2024poster

Conditional diffusion models can create unseen images in various settings, aiding image interpolation. Interpolation in latent spaces is well-studied, but interpolation with specific conditions like text or image is less understood. Common approaches interpolate linearly in the conditioning space bu…

2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

Can I Trust Your Answer? Visually Grounded Video Question Answering

CVPR 2024highlight

We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video- language understanding. Specifically by forcing vision- language models (VLMs) to answer questions and simultane- ously provide visual evidence we seek to ascertain the extent to whic…

2024

Enhancing Video Super-Resolution via Implicit Resampling-based Alignment

CVPR 2024highlight

In video super-resolution it is common to use a frame-wise alignment to support the propagation of information over time. The role of alignment is well-studied for low-level enhancement in video but existing works overlook a critical step -- resampling. We show through extensive experiments that for…

Cited by 15SourcePDFScholar
2024

Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment

ECCV 2024poster

"Procedural activity videos often exhibit a long-tailed action distribution due to varying action frequencies and durations. However, state-of-the-art temporal action segmentation methods overlook the long tail and fail to recognize tail actions. Existing long-tail methods make class-independent ass…

2024

Make Me a BNN: A Simple Strategy for Estimating Bayesian Uncertainty from Pre-trained Models

CVPR 2024poster

Deep Neural Networks (DNNs) are powerful tools for various computer vision tasks yet they often struggle with reliable uncertainty quantification -a critical requirement for real-world applications. Bayesian Neural Networks (BNN) are equipped for uncertainty estimation but cannot scale to large DNNs…

Cited by 8SourcePDFScholar
2024

NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model

ECCV 2024oral

"Modeling the physical contacts between the hand and object is standard for refining inaccurate hand poses and generating novel human grasp in 3D hand-object reconstruction. However, existing methods rely on geometric constraints that cannot be specified or controlled. This paper introduces a novel…

Cited by 2SourcePDFScholar
2024

Pairwise Distance Distillation for Unsupervised Real-World Image Super-Resolution

ECCV 2024poster

"Standard single-image super-resolution creates paired training data from high-resolution images through fixed downsampling kernels. However, real-world super-resolution (RWSR) faces unknown degradations in the low-resolution inputs, all the while lacking paired training data. Existing methods appro…

2024

Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement

ICLR 2024poster

Activation shaping has proven highly effective for identifying out-of-distribution (OOD) samples post-hoc. Activation shaping prunes and scales network activations before estimating the OOD energy score; such an extremely simple approach achieves state-of-the-art OOD detection with minimal in-distri…

2024

WAVE: Warping DDIM Inversion Features for Zero-shot Text-to-Video Editing

ECCV 2024poster

"Text-driven video editing has emerged as a prominent application based on the breakthroughs of image diffusion models. Existing state-of-the-art methods focus on zero-shot frameworks due to limited training data and computing resources. To preserve structure consistency, previous frameworks usually…

2023

Analyzing and Diagnosing Pose Estimation With Attributions

CVPR 2023poster

We present Pose Integrated Gradient (PoseIG), the first interpretability technique designed for pose estimation. We extend the concept of integrated gradients for pose estimation to generate pixel-level attribution maps. To enable comparison across different pose frameworks, we unify different pose…

2023

DropIT: Dropping Intermediate Tensors for Memory-Efficient DNN Training

ICLR 2023poster

A standard hardware bottleneck when training deep neural networks is GPU memory. The bulk of memory is occupied by caching intermediate tensors for gradient computation in the backward pass. We propose a novel method to reduce this footprint - Dropping Intermediate Tensors (DropIT). DropIT drops mi…

2023

Improving Deep Regression with Ordinal Entropy

ICLR 2023poster

In computer vision, it is often observed that formulating regression problems as a classification task yields better performance. We investigate this curious phenomenon and provide a derivation to show that classification, with the cross-entropy loss, outperforms regression with a mean squared error…

2023

Opening the Vocabulary of Egocentric Actions

NeurIPS 2023poster

Human actions in egocentric videos often feature hand-object interactions composed of a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations — sparsity of action compositions and a closed set of interacting objects. Thi…

2023

Overcoming the Trade-Off Between Accuracy and Plausibility in 3D Hand Shape Reconstruction

CVPR 2023poster

Direct mesh fitting for 3D hand shape reconstruction estimates highly accurate meshes. However, the resulting meshes are prone to artifacts and do not appear as plausible hand shapes. Conversely, parametric models like MANO ensure plausible hand shapes but are not as accurate as the non-parametric m…

Cited by 9SourcePDFScholar
2022

A Generalized & Robust Framework for Timestamp Supervision in Temporal Action Segmentation

ECCV 2022poster

"In temporal action segmentation, Timestamp supervision requires only a handful of labeled frames per video sequence. For unlabelled frames, Timestamp works rely on assigning hard labels and performance rapidly collapses under subtle violations of the annotation assumptions. We propose a novel Expec…

Cited by 28SourcePDFScholar
2022

Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

CVPR 2022poster

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101…

Cited by 246PDFcodeScholar
2022

Comprehensive Regularization in a Bi-directional Predictive Network for Video Anomaly Detection

AAAI 2022technical

Video anomaly detection aims to automatically identify unusual objects or behaviours by learning from normal videos. Previous methods tend to use simplistic reconstruction or prediction constraints, which leads to the insufficiency of learned representations for normal data. As such, we propose a no…

Cited by 79SourcePDFScholar
2022

Iterative Contrast-Classify for Semi-supervised Temporal Action Segmentation

AAAI 2022technical

Temporal action segmentation classifies the action of each frame in (long) video sequences. Due to the high cost of frame-wise labeling, we propose the first semi-supervised method for temporal action segmentation. Our method hinges on unsupervised representation learning, which, for temporal action…

2022

Leveraging Action Affinity and Continuity for Semi-Supervised Temporal Action Segmentation

ECCV 2022poster

"We present a semi-supervised learning approach to the temporal action segmentation task. The goal of the task is to temporally detect and segment actions in long, untrimmed procedural videos, where only a small set of videos are densely labelled, and a large collection of videos are unlabelled. To…

Cited by 19SourcePDFScholar
2022

Perception-Distortion Balanced ADMM Optimization for Single-Image Super-Resolution

ECCV 2022poster

"In image super-resolution, both pixel-wise accuracy and perceptual fidelity are desirable. However, most deep learning methods only achieve high performance in one aspect due to the perception-distortion trade-off, and works that successfully balance the trade-off rely on fusing results from separa…

2022

Video as Conditional Graph Hierarchy for Multi-Granular Question Answering

AAAI 2022technical

Video question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been focused on designing sophisticated cross-modal interactions to fuse the information from two modalities, while encoding the…

2021

NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions

CVPR 2021poster

We introduce NExT-QA, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions. Based on the dataset, we set up multi-choice and open-ended QA tasks targeting at causal action reasoning, temporal action reaso…

Cited by 463PDFcodeScholar
2021

Towards Compact Single Image Super-Resolution via Contrastive Self-distillation

IJCAI 2021poster

Convolutional neural networks (CNNs) are highly successful for super-resolution (SR) but often require sophisticated architectures with heavy memory cost and computational overhead significantly restricts their practical deployments on resource-limited devices. In this paper, we proposed a novel con…

2020

Dual Grid Net: Hand Mesh Vertex Regression from Single Depth Maps

ECCV 2020poster

We aim to recover the dense 3D surface of the hand from depth maps and propose a network that can predict mesh vertices, transformation matrices for every joint and joint coordinates in a single forward pass. Use fully convolutional architectures, we first map depth image features to the mesh grid a…

Cited by 31SourcePDFScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

Temporal Aggregate Representations for Long-Range Video Understanding

ECCV 2020poster

Future prediction, especially in long-range videos, requires reasoning from current and past observations. In this work, we address questions of temporal extent, scaling, and level of semantic abstraction with a flexible multi-granular temporal aggregation framework. We show that it is possible to a…

2017

Crossing Nets: Combining GANs and VAEs With a Shared Latent Space for Hand Pose Estimation

CVPR 2017spotlight

State-of-the-art methods for 3D hand pose estimation from depth images require large amounts of annotated training data. We propose modelling the statistical relationship of 3D hand poses and corresponding depth images using two deep generative models with a shared latent space. By design, our archi…

Cited by 184PDFScholar