← Search

Mingtong Dai

2 accepted papers

2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

CVPR 2026

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE)

Cited by 0SourcecodeScholar
2026

PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing

ICLR 2026poster

Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggl…

Cited by 0SourceScholar