← Search

Yiyang Huang

4 accepted papers

2026

SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense

ICLR 2026poster

Large Vision-Language Models (LVLMs) excel in diverse cross-modal tasks. However, object hallucination, where models produce plausible but inaccurate object descriptions, remains a significant challenge. In contrast to previous work focusing on LLM components, this paper is the first to trace LVLM h…

Cited by 0SourcecodeScholar
2025

D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition

EMNLP 2025

Video large language models (Vid-LLMs), which excel in diverse video-language tasks, can be effectively constructed by adapting image-pretrained vision-language models (VLMs). However, this adaptation remains challenging, as it requires processing dense and temporally extended visual inputs that exc

2025

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

ICCV 2025poster

Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation…