← Search

Ziyue Wang

31 accepted papers

2026

3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis

ICML 2026poster

3D CT analysis spans a continuum from low-level perception to high-level clinical understanding. Existing 3D-oriented analysis methods adopt either isolated task-specific modeling or task-agnostic end-to-end paradigms to produce one-hop outputs, impeding the systematic accumulation of perceptual evi…

Cited by 2SourceScholar
2026

Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning

ICLR 2026poster

The professionalism of a human doctor in outpatient service depends on two core abilities: the ability to make accurate medical decisions and the medical consultation skill to conduct strategic, empathetic patient inquiry. Existing Large Language Models (LLMs) have achieved remarkable accuracy on me…

Cited by 0SourcecodeScholar
2026

KNNDA: A New Perspective of Alignment Recovery for Partially View-Aligned Clustering

AAAI 2026technical

In multi-view clustering (MVC), complementary and consistent information from multiple views is integrated to improve clustering performance. However, inter-view sample correspondences may be partially missing in practice, making it difficult to learn cross-view consistency, which leads to the parti

Cited by 0SourcePDFScholar
2026

MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow

ICLR 2026poster

Modern clinical diagnosis relies on the comprehensive analysis of multi-modal patient data, drawing on medical expertise to ensure systematic and rigorous reasoning. Recent advances in Vision–Language Models (VLMs) and agent-based methods are reshaping medical diagnosis by effectively integrating mu…

Cited by 0SourcecodeScholar
2026

OmniDenseCap: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

ICML 2026poster

This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine-grained, and structured audio-visual narratives with explicit timestamps. To ensure dense semantic coverage, we introduce a six-dimensional structural schema to create "script-like" captions, enabling reade…

Cited by 0SourceScholar
2026

PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

AAAI 2026technical

While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding. Existing alignment methods struggle to capture fine-grained corr

Cited by 0SourcePDFScholar
2026

Video-KTR: Reinforcing Video Reasoning via Key Token Attribution

ICLR 2026poster

Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models (MLLMs), yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token selection. Such approaches neglect fine-grained links among visual input…

Cited by 5SourcecodeScholar
2025

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

ACL 2025long

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked.…

2025

BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model

IROS 2025

Localization is one of the core parts of modern robotics. Classic localization methods typically follow the retrieve-then-register paradigm, achieving remarkable success. Recently, the emergence of end-to-end localization approaches has offered distinct advantages, including a streamlined system arc

Cited by 1SourcecodeScholar
2025

CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models

CVPR 2025poster

Vision-Language Models (VLMs) have recently witnessed significant progress in visual comprehension. As the permitting length of image context grows, VLMs can now comprehend a broader range of views and spaces. Current benchmarks provide insightful analysis of VLMs in tasks involving complex visual i…

2025

Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models

EMNLP 2025

Object hallucinations in Large Vision-Language Models (LVLMs) significantly impede their real-world applicability. As the primary component for accurately interpreting visual information, the choice of visual encoder is pivotal. We hypothesize that the diverse training paradigms employed by differen

2025

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

EMNLP 2025

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to the lack of relevant datasets, research on semantic understanding of Dongba hieroglyphs has progressed slowly. To this end

2025

Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information

IJCAI 2025

Recently, multi-view data has gradually attracted attention. However, real-world applications often face Partial View-aligned Problem (PVP) and Partially Sample-missing Problem (PSP) due to data loss or corruption. Existing methods addressing PVP typically focus only on learning from the information

Cited by 0SourcePDFScholar
2025

EgoLife: Towards Egocentric Life Assistant

CVPR 2025poster

We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one we…

2025

How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape Game

ICCV 2025poster

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require coordinating multiple abilities, including visual perception, visual reasoning, spatial awareness, and target deduction.…

2025

Incomplete and Unpaired Multi-View Graph Clustering with Cross-View Feature Fusion

AAAI 2025technical

Due to its effectiveness and efficiency, graph-based multi-view clustering has recently attracted much attention. However, the multi-view data are often incomplete and unpaired in real-world applications as a consequence of data loss or corruption. Although efforts have been made through a series of…

Cited by 0SourcePDFScholar
2025

MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. Due to their strong performance in image-text alignment, MLLMs can effectively understand image-text pairs with clear meanings. However, effectively resolving the inherent ambiguiti

2025

Multi-scale Context Intertwining for Panoramic Renal Pathology Segmentation

ICASSP 2025accepted

Panoramic segmentation of renal pathological tissues plays a crucial role in diagnosing renal carcinoma and other kidney-related diseases. The multi-scale nature of kidney tissues, which requires different magnification levels for accurate analysis, presents a significant challenge for segmentation…

Cited by 0SourceScholar
2025

Perspective Transition of Large Language Models for Solving Subjective Tasks

ACL 2025finding

Large language models (LLMs) have revolutionized the field of natural language processing, enabling remarkable progress in various tasks. Different from objective tasks such as commonsense reasoning and arithmetic question-answering, the performance of LLMs on subjective tasks is still limited, wher…

2025

The Four Color Theorem for Cell Instance Segmentation

ICML 2025poster

Cell instance segmentation is critical to analyzing biomedical images, yet accurately distinguishing tightly touching cells remains a persistent challenge. Existing instance segmentation frameworks, including detection-based, contour-based, and distance mapping-based approaches, have made significan…

2024

Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion

ACL 2024long

With the bloom of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) that incorporate LLMs with pre-trained vision models have recently demonstrated impressive performance across diverse vision-language tasks. However, they fall short to comprehend context involving multiple imag…

2024

CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models

ACL 2024long

Multimodal large language models (MLLMs) have demonstrated promising results in a variety of tasks that combine vision and language. As these models become more integral to research and applications, conducting comprehensive evaluations of their capabilities has grown increasingly important. However…

Cited by 8SourcePDFScholar
2024

Graph-Structured Speculative Decoding

ACL 2024findings

Speculative decoding has emerged as a promising technique to accelerate the inference of Large Language Models (LLMs) by employing a small language model to draft a hypothesis sequence, which is then validated by the LLM. The effectiveness of this approach heavily relies on the balance between perfo…

2024

Model Composition for Multimodal Large Language Models

ACL 2024long

Recent developments in Multimodal Large Language Models (MLLMs) have shown rapid progress, moving towards the goal of creating versatile MLLMs that understand inputs from various modalities. However, existing methods typically rely on joint training with paired multimodal instruction data, which is…

2024

Octopus: Embodied Vision-Language Programmer from Environmental Feedback

ECCV 2024poster

"Large vision-language models (VLMs) have achieved substantial progress in multimodal perception and reasoning. When integrated into an embodied agent, existing embodied VLM works either output detailed action sequences at the manipulation level or only provide plans at an abstract level, leaving a…

2023

Filling the Image Information Gap for VQA: Prompting Large Language Models to Proactively Ask Questions

EMNLP 2023long findings

Large Language Models (LLMs) demonstrate impressive reasoning ability and the maintenance of world knowledge not only in natural language tasks, but also in some vision-language tasks such as open-domain knowledge-based visual question answering (OK-VQA). As images are invisible to LLMs, researchers…

Cited by 0SourcecodeScholar
2023

Imperceptible Adversarial Attack via Invertible Neural Networks

AAAI 2023technical

Adding perturbations via utilizing auxiliary gradient information or discarding existing details of the benign images are two common approaches for generating adversarial examples. Though visual imperceptibility is the desired property of adversarial examples, conventional adversarial attacks still…

2023

Scaling Law Analysis for Covariance Based Activity Detection in Cooperative Multi-Cell Massive Mimo

ICASSP 2023accepted

This paper studies the covariance based activity detection problem in a multi-cell massive multiple-input multiple-output (MIMO) system, where the active devices transmit their signature sequences to multiple base stations (BSs), and the BSs cooperatively detect the active devices based on the recei…

Cited by 0SourceScholar
2022

Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval

ECCV 2022poster

"In this paper we revisit feature fusion, an old-fashioned topic, in the new context of text-to-video retrieval. Different from previous research that considers feature fusion only at one end, let it be video or text, we aim for feature fusion for both ends within a unified framework. We hypothesize…

2021

An Efficient Active Set Algorithm for Covariance Based Joint Data and Activity Detection for Massive Random Access with Massive MIMO

ICASSP 2021accepted

This paper proposes a computationally efficient algorithm to solve the joint data and activity detection problem for massive random access with massive multiple-input multiple-output (MIMO). The BS acquires the active devices and their data by detecting the transmitted preassigned nonorthogonal sign…

Cited by 0SourceScholar