← Search

Zihang Jiang

12 accepted papers

2026

MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text

AAAI 2026technical

Artificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major challenge. Existing medical multi-modal benchmarks are typically limi

Cited by 0SourcePDFScholar
2026

MedGMAE: Gaussian Masked Autoencoders for Medical Volumetric Representation Learning

ICLR 2026poster

Self-supervised pre-training has emerged as a critical paradigm for learning transferable representations from unlabeled medical volumetric data. Masked autoencoder based methods have garnered significant attention, yet their application to volumetric medical image faces fundamental limitations from…

Cited by 0SourcecodeScholar
2025

AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIP

CVPR 2025poster

Anomaly detection (AD) identifies outliers for applications like defect and lesion detection. While CLIP shows promise for zero-shot AD tasks due to its strong generalization capabilities, its inherent Anomaly-Unawareness leads to limited discrimination between normal and abnormal features. To addre…

2025

Self-Supervised Diffusion MRI Denoising via Iterative and Stable Refinement

ICLR 2025poster

Magnetic Resonance Imaging (MRI), including diffusion MRI (dMRI), serves as a ``microscope'' for anatomical structures and routinely mitigates the influence of low signal-to-noise ratio scans by compromising temporal or spatial resolution. However, these compromises fail to meet clinical demands for…

2024

CARZero: Cross-Attention Alignment for Radiology Zero-Shot Classification

CVPR 2024poster

The advancement of Zero-Shot Learning in the medical domain has been driven forward by using pre-trained models on large-scale image-text pairs focusing on image-text alignment. However existing methods primarily rely on cosine similarity for alignment which may not fully capture the complex relatio…

2023

OmniAvatar: Geometry-Guided Controllable 3D Head Synthesis

CVPR 2023poster

We present OmniAvatar, a novel geometry-guided 3D head synthesis model trained from in-the-wild unstructured images that is capable of synthesizing diverse identity-preserved 3D heads with compelling dynamic details under full disentangled control over camera poses, facial expressions, head shapes,…

Cited by 29SourcePDFScholar
2023

TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration

ICCV 2023poster

We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information pro…

Cited by 69PDFcodeScholar
2022

Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning

CVPR 2022poster

Class Incremental Learning (CIL) aims at learning a classifier in a phase-by-phase manner, in which only data of a subset of the classes are provided at each phase. Previous works mainly focus on mitigating forgetting in phases after the initial one. However, we find that improving CIL at its initia…

Cited by 89PDFcodeScholar
2021

All Tokens Matter: Token Labeling for Training Better Vision Transformers

NeurIPS 2021poster

In this paper, we present token labeling---a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViTs that computes the classification loss on an additional trainable class token, our proposed one takes advantage of all t…

2020

PAGAN: A Phase-Adapted Generative Adversarial Networks for Speech Enhancement

ICASSP 2020accepted

Deep neural networks (DNNs) are becoming more and more popular in speech enhancement. Most of DNN-based speech enhancement approaches currently operate on magnitude spectra and ignore the phase mismatch between noisy and clean speech which greatly limits the speech enhancement performance. This pape…

Cited by 0SourceScholar
2020

ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning

ICLR 2020poster

Recent powerful pre-trained language models have achieved remarkable performance on most of the popular datasets for reading comprehension. It is time to introduce more challenging datasets to push the development of this field towards more comprehensive reasoning of text. In this paper, we introduc…

Cited by 264SourcecodeScholar