← Search

Zhen Yuan

2 accepted papers

2026

MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet their positional encoding mechanisms remain suboptimal. Existing approaches uniformly assign positional indices to all tokens, overlooking variations in information density within and across modalities,

Cited by 0SourceScholar
2026

Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions

ICRA 2026poster

Manipulation involving rigid-deformable interactions, such as hanging clothes or dressing humans, is essential for household robots. Compared to single-object manipulation or interactions between rigid bodies, these tasks are particularly challenging due to the rich multi-point contacts and the comp…

Cited by 0Scholar