← Search

Chenfei Liao

4 accepted papers

2026

Accelerating Streaming Video Large Language Models via Hierarchical Token Compression

CVPR 2026

Streaming Video Large Language Models (VideoLLMs) have demonstrated impressive performance across various video understanding tasks, but they face significant challenges in real-time deployment due to the high computational cost of processing dense visual tokens from continuous video streams. In str

Cited by 0SourcecodeScholar
2026

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

ICML 2026poster

Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human creativity. Current reasoning-augmented paradigms mostly rely on explicit thought processes, where intermediate reasoni…

Cited by 0SourceScholar
2025

OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation

ICCV 2025poster

Segment Anything Model 2 (SAM2) has emerged as a strong base model in various pinhole imaging segmentation tasks. However, when applying it to 360^\circ domain, the significant field-of-view (FoV) gap between pinhole (70^\circ x70^\circ) and panoramic images (180^\circ x360^\circ) poses unique chall…

Cited by 0SourcePDFScholar
2025

UMSSS: A Visual Scene Semantic Segmentation Dataset for Underground Mines

ICASSP 2025accepted

Specialized datasets designed for mining scenarios are the essential foundation for the development, operation, and research of intelligent mines. Currently, the available datasets focus primarily on open-pit mines, with a lack of specialized datasets for underground mines. This gap severely hinders…

Cited by 0SourceScholar