← Search

Minghao Lai

2 accepted papers

2026

Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

CVPR 2026

Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We in

Cited by 0SourcecodeScholar
2025

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

ICCV 2025poster

Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streaming video or queries using visual cues. To fill this gap, we present a new task named Online Video Grounding with Hybrid-…

Cited by 0SourcePDFScholar