← Search

Yuting Mei

1 accepted papers

2025

EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining

NeurIPS 2025poster

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most previous works learn from 1D text or 2D visual cues, such as bo…

Cited by 0SourcecodeScholar