← Search

Jianzhe Gao

6 accepted papers

2026

Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

CVPR 2026

Clinical video diagnosis, in which physicians assess dynamic tissue responses across procedural stages, is critical for detecting diseases such as cervical and colorectal cancers. Recent spatiotemporal models map visual progressions directly to diagnostic outputs, yet overlook two hallmarks of exper

Cited by 0SourceScholar
2026

History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

AAAI 2026technical

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agent

Cited by 0SourcePDFScholar
2026

SAMT: Generating Structured Avatar Meshes and Textures from a Single Image

ICML 2026poster

Despite rapid progress in 3D generative models, producing production-grade 3D face assets from a single image remains challenging. To reconstruct facial micro-structures and fine-grained multiview-consistent textures, this work presents a two-stage framework named SAMT for monocular 3D avatar genera…

Cited by 0SourceScholar
2026

Uncertainty-Aware 3D Reconstruction for Dynamic Underwater Scenes

ICLR 2026poster

Underwater 3D reconstruction remains challenging due to the intricate interplay between light scattering and environment dynamics. While existing methods yield plausible reconstruction with rigid scene assumptions, they struggle to capture temporal dynamics and remain sensitive to observation noise.…

Cited by 0SourceScholar
2026

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

ICLR 2026poster

Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient evidence for reliable grounding or ambiguity in interpreting spatial cues, yet th…

Cited by 0SourceScholar
2025

3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation

ICCV 2025poster

Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough scene understanding. While existing works equip agents with various scene representations to enhance spatial awareness, they often neglect the compl…