← Search

Jialong Xue

2 accepted papers

2026

Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention

CVPR 2026

Video large language models (Video-LLMs) face high computational costs due to large volumes of visual tokens. Existing token compression methods typically adopt a two-stage spatiotemporal compression strategy, relying on stage-specific metrics and an implicit assumption of spatiotemporal separabilit

Cited by 0SourceScholar
2025

High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects

IROS 2025

High-precision tiny object alignment remains a common and critical challenge for humanoid robots in real world. To address this problem, this paper proposes a vision-based framework for precisely estimating and controlling the relative position between a handheld tool and a target object for humanoi

Cited by 1SourceScholar