← Search

Jianjiang Yang

4 accepted papers

2026

CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning

AAAI 2026technical

Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it also introduces severe image token redundancy. Because image tokens contain sparse information, many contribute little to r

Cited by 0SourcePDFScholar
2026

Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning

AAAI 2026technical

Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context dem

Cited by 0SourcePDFScholar
2025

ReLoop: “Seeing Twice and Thinking Backwards” via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

EMNLP 2025

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in open-ended visual question answering, they remain vulnerable to hallucinations. These are outputs that contradict or misrepresent input semantics, posing a critical challenge to the reliability and factual consistenc

2025

TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration

EMNLP 2025

Multimodal in-context learning (ICL) has emerged as a key mechanism for harnessing the capabilities of large vision–language models (LVLMs). However, its effectiveness remains highly sensitive to the quality of input ICL sequences, particularly for tasks involving complex reasoning or open-ended gen

Cited by 0SourcePDFScholar