← Search

Xinlei Yu

9 accepted papers

2026

DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag Editing

ICLR 2026poster

Drag-based image editing has long suffered from distortions in the target region, largely because the priors of earlier base models, Stable Diffusion, are insufficient to project optimized latents back onto the natural image manifold. With the shift from UNet-based DDPMs to more scalable DiT with fl…

Cited by 0SourcecodeScholar
2026

Dual Latent Memory for Visual Multi-agent System

ICML 2026poster

While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure …

Cited by 0SourceScholar
2026

LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules

AAAI 2026technical

Diagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been made in using multimodal large language models for analyzing

Cited by 0SourcePDFScholar
2026

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

ICML 2026poster

Humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings. However, existing omnimodal models still exhibit substantial performance degradation on visual tasks when the audio modality is incorporated. We identify this …

Cited by 0SourceScholar
2026

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning

AAAI 2026technical

Current multimodal large language models (MLLMs) still face significant challenges in complex visual tasks (e.g., spatial understanding, fine-grained perception). Prior methods have tried to incorporate visual reasoning, however, they fail to leverage attention correction with spatial cues to iterat

Cited by 0SourcePDFScholar
2026

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views

CVPR 2026

Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limited views remains a significant challenge. Previous reasoning methods typically rely on pure text (e.g., topological cogn

Cited by 0SourcecodeScholar
2026

VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models

CVPR 2026

Despite the remarkable success of Vision-Language Models (VLMs), their performance on a range of complex visual tasks is often hindered by a "visual processing bottleneck": a propensity to lose grounding in visual evidence and exhibit a deficit in contextualized visual experience during prolonged ge

Cited by 0SourcecodeScholar
2026

Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling

CVPR 2026

The dominant paradigm of monolithic scaling in Vision-Language Models (VLMs) is failing for understanding and reasoning in documents, yielding diminishing returns as it struggles with the inherent need of this domain for document-based procedural reasoning, cognitive complexity, and factual accuracy

Cited by 0SourcecodeScholar
2026

Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow

ICLR 2026poster

Multi-Agent System (MAS) powered by Visual Language Models (VLMs) enables challenging tasks but suffers from a novel failure term, multi-agent visual hallucination snowballing, where hallucinations are seeded in a single agent and amplified by following ones due to the over-reliance on textual flow…

Cited by 0SourcecodeScholar