← Search

Yuxuan Ding

4 accepted papers

2025

TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

ICLR 2025poster

Existing benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understanding. However, *how well do the models truly perform visual temporal reasoning?* Our study of existing benchmarks shows th…

2025

TReF-6: Inferring Task-Relevant Frames from a Single Demonstration for One-Shot Skill Generalization

CoRL 2025poster

Robots often struggle to generalize from a single demonstration due to the lack of a transferable and interpretable spatial representation. In this work, we introduce TReF-6, a method that infers a simplified, abstracted 6DoF Task-Relevant Frame from a single trajectory. Our approach identifies an i…

Cited by 0SourceScholar