← Search

Kaiyan Xiao

2 accepted papers

2026

GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding

CVPR 2026

Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame sampling to extract video information, resulting in a sparse distribution of ke

Cited by 0SourcecodeScholar
2025

Reinforcement Learning-based Optimization of Humanoid Joint Motion Control via Text-driven Human Motion Mapping

IROS 2025

Human motion retargeting for humanoid robots, transferring human motion data to robots for imitation, presents significant challenges but offers considerable potential for real-world applications. Traditionally, this process relies on human demonstrations captured through pose estimation or motion c

Cited by 0SourceScholar