← Search

Yuqi Pang

3 accepted papers

2026

MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention

AAAI 2026technical

Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual models. However, these approaches face high training and inference costs, as well as challenges in extracting visual det

Cited by 0SourcePDFScholar
2025

Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning

ICASSP 2025accepted

Although Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large Language Models (MLLMs), however, is resource-intensive and constrained by various training limitations. In this paper,…

Cited by 0SourceScholar
2025

Object-Based Video Tampering Localization via Trace Consistency Analysis

ICASSP 2025accepted

With the rapid advancement of object-based video inpainting and splicing tampering techniques, the dissemination of malicious videos on the internet poses significant risks. Existing localization methods, however, exhibit limitations such as restriction to specific datasets, limited performance in d…

Cited by 0SourceScholar