← Search

Yayuan Li

6 accepted papers

2026

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

CVPR 2026

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent motion due to fixed-length temporal inputs and discrete fram

Cited by 0SourcecodeScholar
2026

Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos

CVPR 2026

We introduce Mistake Attribution (MATT), a new task for fine-grained understanding of human mistakes in egocentric videos. While prior work detects whether a mistake occurs, MATT attributes the mistake to what part of the instruction is violated (semantic role), when in the video the deviation becom

Cited by 0SourcecodeScholar
2026

When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging

ICML 2026poster

Model merging combines multiple fine-tuned models into a single model by $\textit{adding}$ their weight updates, providing a lightweight alternative to retraining. Existing methods primarily target resolving conflicts between task updates, leaving the failure mode of over-counting shared knowledge u…

Cited by 0SourceScholar
2026

When to Think and When to Look: Uncertainty-Guided Lookback

CVPR 2026

Test-time "thinking" (i.e., generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision-language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how t

Cited by 0SourcecodeScholar
2025

Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP

AAAI 2025technical

Contrastive Language-Image Pretraining (CLIP) has been widely used in vision tasks. Notably, CLIP has demonstrated promising performance in few-shot learning (FSL). However, existing CLIP-based methods in training-free FSL (i.e., without the requirement of additional training) mainly learn different…

2025

Transparent and Coherent Procedural Mistake Detection

EMNLP 2025

Procedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a procedural text). Despite significant recent efforts, machine performance in the wild remains nonviable, and the reason

Cited by 0SourcePDFScholar