← Search

Jinglu Wang

21 accepted papers

2026

CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework

ICLR 2026poster

Large visual language models (VLMs) have shown strong multi-modal medical reasoning ability, but most operate as end-to-end black boxes, diverging from clinicians’ evidence-based, staged workflows and hindering clinical accountability. Complementarily, expert visual grounding models can accurately l…

Cited by 0SourceScholar
2026

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

ICLR 2026poster

Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across diverse medical specialties, limiting their performance. Recent efforts introduce multi-agent collaboration frameworks insp…

Cited by 0SourceScholar
2026

Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images

ICLR 2026poster

The rapid advancement of AI-generated content (AIGC) has enabled the synthesis of visually convincing images; however, many such outputs exhibit subtle \textbf{semantic anomalies}, including unrealistic object configurations, violations of physical laws, or commonsense inconsistencies, which comprom…

Cited by 0SourceScholar
2025

StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image Streams

ICCV 2025poster

The advent of 3D Gaussian Splatting (3DGS) has advanced 3D scene reconstruction and novel view synthesis. With the growing interest of interactive applications that need immediate feedback, online 3DGS reconstruction in real-time is in high demand. However, none of existing methods yet meet the dema…

Cited by 0SourcePDFScholar
2024

QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition

CVPR 2024poster

Audiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved establishing robust correspondences between audio and visual contents poses unique challeng…

2024

R^2-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations

ECCV 2024poster

"Referring perception, which aims at grounding visual objects with multimodal referring guidance, is essential for bridging the gap between humans, who provide instructions, and the environment where intelligent systems perceive. Despite progress in this field, the robustness of referring perception…

Cited by 3SourcePDFScholar
2023

Efficient View Synthesis with Neural Radiance Distribution Field

ICCV 2023poster

Recent work on Neural Radiance Fields (NeRF) has demonstrated significant advances in high-quality view synthesis. A major limitation of NeRF is its low rendering efficiency due to the need for multiple network forwardings to render a single pixel. Existing methods to improve NeRF either reduce the…

Cited by 1PDFcodeScholar
2023

High-Fidelity and Freely Controllable Talking Head Video Generation

CVPR 2023poster

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face often has unexpected deformation and severe distortions. Sec…

Cited by 36SourcePDFScholar
2023

PaintSeg: Painting Pixels for Training-free Segmentation

NeurIPS 2023poster

The paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which creates a contrast between the original image and a painted image in which a masked area is painted using off-the-shelf ge…

2023

Robust Referring Video Object Segmentation with Cyclic Structural Consensus

ICCV 2023poster

Referring Video Object Segmentation (R-VOS) is a challenging task that aims to segment an object in a video based on a linguistic expression. Most existing R-VOS methods have a critical assumption: the object referred to must appear in the video. This assumption, which we refer to as "semantic conse…

Cited by 35PDFScholar
2023

Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction

CVPR 2023poster

The Multiplane Image (MPI), containing a set of fronto-parallel RGBA layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed structure limits the performance, especially for surfaces imaged at oblique angles. We introduce the Structural MPI (S-MPI),…

Cited by 10SourcePDFScholar
2023

Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text

EMNLP 2023long main

Linguistic communication is prevalent in Human-Computer Interaction (HCI). Speech (spoken language) serves as a convenient yet potentially ambiguous form due to noise and accents, exposing a gap compared to text. In this study, we investigate the prominent HCI task, Referring Video Object Segmentati…

Cited by 0SourceScholar
2023

Two-Shot Video Object Segmentation

CVPR 2023poster

Previous works on video object segmentation (VOS) are trained on densely annotated videos. Nevertheless, acquiring annotations in pixel level is expensive and time-consuming. In this work, we demonstrate the feasibility of training a satisfactory VOS model on sparsely annotated videos--we merely req…

2022

Hybrid Instance-Aware Temporal Fusion for Online Video Instance Segmentation

AAAI 2022technical

Recently, transformer-based image segmentation methods have achieved notable success against previous solutions. While for video domains, how to effectively model temporal context with the attention of object instances across frames remains an open problem. In this paper, we propose an online video…

Cited by 21SourcePDFScholar
2022

Neural Capture of Animatable 3D Human from Monocular Video

ECCV 2022poster

"We present a novel paradigm of building an animatable 3D human representation from a monocular video input, such that it can be rendered in any unseen poses and views. Our method is based on a dynamic Neural Radiance Field (NeRF) rigged by a mesh-based parametric 3D human model serving as a geometr…

Cited by 28SourcePDFScholar
2022

Reliable Propagation-Correction Modulation for Video Object Segmentation

AAAI 2022technical

Error propagation is a general but crucial problem in online semi-supervised video object segmentation. We aim to suppress error propagation through a correction mechanism with high reliability. The key insight is to disentangle the correction from the conventional mask propagation process with re…

2021

Weakly-supervised Temporal Action Localization by Uncertainty Modeling

AAAI 2021technical

Weakly-supervised temporal action localization aims to learn detecting temporal intervals of action classes with only video-level labels. To this end, it is crucial to separate frames of action classes from the background frames (i.e., frames not belonging to any action classes). In this paper, we p…

2020

Joint Semantic Segmentation and Boundary Detection Using Iterative Pyramid Contexts

CVPR 2020poster

In this paper, we present a joint multi-task learning framework for semantic segmentation and boundary detection. The critical component in the framework is the iterative pyramid context module (PCM), which couples two tasks and stores the shared latent semantics to interact between the two tasks. F…

Cited by 169PDFScholar
2017

Progressive Large Scale-Invariant Image Matching in Scale Space

ICCV 2017poster

The power of modern image matching approaches is still fundamentally limited by the abrupt scale changes in images. In this paper, we propose a scale-invariant image matching approach to tackling the very large scale variation of views. Drawing inspiration from the scale space theory, we start with…

Cited by 47PDFScholar
2015

Higher-Order CRF Structural Segmentation of 3D Reconstructed Surfaces

ICCV 2015poster

In this paper, we propose a structural segmentation algorithm to partition multi-view stereo reconstructed surfaces of large-scale urban environments into structural segments. Each segment corresponds to a structural component describable by a surface primitive of up to the second order. This segmen…

Cited by 18PDFScholar