← Search

Kele Shao

4 accepted papers

2026

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

CVPR 2026

Omnimodal large language models (OmniLLMs) have attracted increasing research attention of late towards unified audio-video understanding. However, the high computational cost of processing longer joint audio-video token sequences has become a key bottleneck. Existing token compression methods have

Cited by 0SourcecodeScholar
2025

HoliTom: Holistic Token Merging for Fast Video Large Language Models

NeurIPS 2025poster

Video large language models (video LLMs) excel at video comprehension but face significant computational inefficiency due to redundant video tokens. Existing token pruning methods offer solutions. However, approaches operating within the LLM (inner-LLM pruning), such as FastV, incur intrinsic comput…

Cited by 0SourcecodeScholar
2025

LI-GS: Gaussian Splatting With LiDAR Incorporated for Accurate Large-Scale Reconstruction

RA-L 2025

Large-scale 3D reconstruction is critical in the field of robotics, and the potential of 3D Gaussian Splatting (3DGS) for achieving accurate object-level reconstruction has been demonstrated. However, ensuring geometric accuracy in outdoor and unbounded scenes remains a significant challenge. This s

Cited by 32SourceScholar