← Search

Chuanqi Cheng

7 accepted papers

2026

ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models

ICLR 2026poster

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations of existing methods: supervised fine-tuning (SFT) often com…

Cited by 0SourcecodeScholar
2025

A Survey on Personalized Alignment—The Missing Piece for Large Language Models in Real-World Applications

ACL 2025finding

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their transition to real-world applications reveals a critical limitation: the inability to adapt to individual preferences while maintaining alignment with universal human values. Current alignment techniques adopt a one-si…

Cited by 0SourcePDFScholar
2025

DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions

ACL 2025long

In this paper, we propose contextualized and situated text-to-speech (CS-TTS), a novel TTS task to promote more accurate and customized speech generation using prompts with Dialogues, Narratives, and Actions (DNA). While prompt-based TTS methods facilitate controllable speech generation, existing TT…

2025

Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

ICML 2025poster

Long-form video processing fundamentally challenges vision-language models (VLMs) due to the high computational costs of handling extended temporal sequences. Existing token pruning and feature merging methods often sacrifice critical temporal dependencies or dilute semantic information. We introduc…

2025

Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains

ACL 2025long

Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features. In this…

Cited by 0SourcePDFScholar
2024

From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis

EMNLP 2024main

We explore multi-step reasoning in vision-language models (VLMs). The problem is challenging, as reasoning data consisting of multiple steps of visual and language processing are barely available. To overcome the challenge, we first introduce a least-to-most visual reasoning paradigm, which interlea…

2024

“In-Dialogues We Learn”: Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue Learning

EMNLP 2024main

Personalized dialogue systems have gained significant attention in recent years for their ability to generate responses in alignment with different personas. However, most existing approaches rely on pre-defined personal profiles, which are not only time-consuming and labor-intensive to create but a…

Cited by 2SourcePDFScholar