← Search

Zehao Xiao

13 accepted papers

2026

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

ICLR 2026poster

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples across modalities while overlooking distributional differences.…

Cited by 0SourceScholar
2026

Towards Uniformity and Alignment for Multimodal Representation Learning

ICML 2026poster

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify …

Cited by 0SourceScholar
2025

DynaPrompt: Dynamic Test-Time Prompt Tuning

ICLR 2025poster

Test-time prompt tuning enhances zero-shot generalization of vision-language models but tends to ignore the relatedness among test samples during inference. Online test-time prompt tuning provides a simple way to leverage the information in previous test samples, albeit with the risk of prompt colla…

Cited by 0SourcePDFScholar
2025

Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes

ICML 2025poster

Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentatio…

Cited by 0SourcePDFScholar
2024

Any-Shift Prompting for Generalization over Distributions

CVPR 2024poster

Image-language models with prompt learning have shown remarkable advances in numerous downstream vision tasks. Nevertheless conventional prompt learning methods overfit the training distribution and lose the generalization ability on the test distributions. To improve the generalization across vario…

Cited by 17SourcePDFScholar
2024

GO4Align: Group Optimization for Multi-Task Alignment

NeurIPS 2024poster

This paper proposes **GO4Align**, a multi-task optimization approach that tackles task imbalance by explicitly aligning the optimization across tasks. To achieve this, we design an adaptive group risk minimization strategy, comprising two techniques in implementation: (i) dynamical group assignment,…

2023

Energy-Based Test Sample Adaptation for Domain Generalization

ICLR 2023poster

In this paper, we propose energy-based sample adaptation at test time for domain generalization. Where previous works adapt their models to target domains, we adapt the unseen target samples to source-trained models. To this end, we design a discriminative energy-based model, which is trained on sou…

2023

ProtoDiff: Learning to Learn Prototypical Networks by Task-Guided Diffusion

NeurIPS 2023poster

Prototype-based meta-learning has emerged as a powerful technique for addressing few-shot learning challenges. However, estimating a deterministic prototype using a simple average function from a limited number of examples remains a fragile process. To overcome this limitation, we introduce ProtoDif…

2022

Association Graph Learning for Multi-Task Classification with Category Shifts

NeurIPS 2022accept

In this paper, we focus on multi-task classification, where related classification tasks share the same label space and are learned simultaneously. In particular, we tackle a new setting, which is more realistic than currently addressed in the literature, where categories shift from training to test…

2022

Learning to Generalize across Domains on Single Test Samples

ICLR 2022poster

We strive to learn a model from a set of source domains that generalizes well to unseen target domains. The main challenge in such a domain generalization scenario is the unavailability of any target domain data during training, resulting in the learned model not being explicitly adapted to the unse…

2021

A Bit More Bayesian: Domain-Invariant Learning with Uncertainty

ICML 2021spotlight

Domain generalization is challenging due to the domain shift and the uncertainty caused by the inaccessibility of target domain data. In this paper, we address both challenges with a probabilistic framework based on variational Bayesian inference, by incorporating uncertainty into neural network wei…

2019

Crowd Counting and Density Estimation by Trellis Encoder-Decoder Networks

CVPR 2019poster

Crowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-…

Cited by 435PDFScholar
2019

Relational Attention Network for Crowd Counting

ICCV 2019poster

Crowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision…

Cited by 214PDFScholar