← Search

Yiming Bao

3 accepted papers

2026

APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

AAAI 2026technical

Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and resource constraints during both training and inference. Alth

Cited by 0SourcePDFScholar
2026

DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos

RSS 2026poster

Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier of manipulation knowledge, offer significant potential for scaling up robot lea…

Cited by 0SourceScholar
2026

UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos

CVPR 2026

Dexterous manipulation remains challenging due to the cost of collecting real-robot teleoperation data, the heterogeneity of hand embodiments, and the high dimensionality of control. We present UniDex, a robot foundation suite that couples a large-scale robot-centric dataset with a unified vision-la

Cited by 0SourcecodeScholar