← Search

Delin Chen

7 accepted papers

2026

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

CVPR 2026

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train VLMs to render explicit images, but the heavy image-generation pre-training oft

Cited by 0SourcecodeScholar
2025

CausalEval: Towards Better Causal Reasoning in Language Models

NAACL 2025long

Causal reasoning (CR) is a crucial aspect of intelligence, essential for problem-solving, decision-making, and understanding the world. While language models (LMs) can generate rationales for their outputs, their ability to reliably perform causal reasoning remains uncertain, often falling short in…

Cited by 0SourcePDFScholar
2025

Scaling Autonomous Agents via Automatic Reward Modeling And Planning

ICLR 2025poster

Large language models (LLMs) have demonstrated remarkable capabilities across a range of text-generation tasks. However, LLMs still struggle with problems requiring multi-step decision-making and environmental feedback, such as online shopping, scientific reasoning, and mathematical problem-solving.…

Cited by 3SourcePDFScholar
2025

VCA: Video Curious Agent for Long Video Understanding

ICCV 2025poster

Long video understanding poses unique challenges due to its temporal complexity and low information density. Recent works address this task by sampling numerous frames or incorporating auxiliary tools using LLMs, both of which result in high computational costs. In this work, we introduce a curiosit…

Cited by 0SourcePDFScholar
2024

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

ICLR 2024poster

A remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large vision-language foundation models (VLMs) fall short of such compositional abilities due to their ``bag-of-words" behaviors and inability to cons…

Cited by 16SourcePDFScholar
2024

FlexAttention for Efficient High-Resolution Vision-Language Models

ECCV 2024poster

"Current high-resolution vision-language models encode images as high-resolution image tokens and exhaustively take all these tokens to compute attention, which significantly increases the computational cost. To address this problem, we propose , a flexible attention mechanism for efficient high-res…

Cited by 13SourcePDFScholar
2023

Scratch Each Other's Back: Incomplete Multi-Modal Brain Tumor Segmentation via Category Aware Group Self-Support Learning

ICCV 2023poster

Although Magnetic Resonance Imaging (MRI) is very helpful for brain tumor segmentation and discovery, it often lacks some modalities in clinical practice. As a result, degradation of prediction performance is inevitable. According to current implementations, different modalities are considered to be…

Cited by 19PDFcodeScholar