← Search

Yuheng Shi

7 accepted papers

2026

Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation

ICLR 2026poster

Robotic manipulation with Vision-Language-Action models requires efficient inference over long-horizon multi-modal context, where attention to dense visual tokens dominates computational cost. Existing methods optimize inference speed by reducing visual redundancy within VLA models, but they overloo…

Cited by 0SourcecodeScholar
2026

Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception

ICLR 2026poster

Multimodal Large Language Models (MLLMs) require high-resolution visual information to perform fine-grained perception, yet processing entire high-resolution images is computationally prohibitive. While recent methods leverage a Region-of-Interest (RoI) mechanism to focus on salient areas, they typ…

Cited by 0SourcecodeScholar
2025

Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation

ICCV 2025poster

While CLIP has advanced open-vocabulary predictions, its performance on semantic segmentation remains suboptimal. This shortfall primarily stems from its spatial-invariant semantic features and constrained resolution. While previous adaptations addressed spatial invariance semantic by modifying the…

2024

Event-based Few-shot Fine-grained Human Action Recognition

IROS 2024poster

Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as…

Cited by 1SourceScholar
2024

Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model

NeurIPS 2024poster

Despite the significant achievements of Vision Transformers (ViTs) in various vision tasks, they are constrained by the quadratic complexity. Recently, State Space Models (SSMs) have garnered widespread attention due to their global receptive field and linear complexity with respect to the input len…

2023

YOLOV: Making Still Image Object Detectors Great at Video Object Detection

AAAI 2023technical

Video object detection (VID) is challenging because of the high variation of object appearance as well as the diverse deterioration in some frames. On the positive side, the detection in a certain frame of a video, compared with that in a still image, can draw support from other frames. Hence, how t…