← Search

Xiantong Zhen

40 accepted papers

2026

Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization

ICML 2026poster

Group Relative Policy Optimization (GRPO) effectively scales LLM reasoning but incurs prohibitive computational costs due to its extensive group-based sampling requirement. While recent selective data utilization methods can mitigate this overhead, they could induce estimation bias by altering the u…

Cited by 0SourceScholar
2025

HYDEN: Hyperbolic Density Representations for Medical Images and Reports

COLING 2025main

In light of the inherent entailment relations between images and text, embedding point vectors in hyperbolic space has been employed to leverage its hierarchical modeling advantages for visual semantic representation learning. However, point vector embeddings struggle to address semantic uncertainty…

2025

LKA-ReID: Vehicle Re-Identification with Large Kernel Attention

ICASSP 2025accepted

With the rapid development of intelligent transportation systems and the popularity of smart city infrastructure, Vehicle Re-ID technology has become an important research field. The vehicle Re-ID task faces an important challenge, which is the high similarity between different vehicles. Existing me…

Cited by 0SourceScholar
2024

Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation

ICML 2024poster

Fine-tuning pretrained large models to downstream tasks is an important problem, which however suffers from huge memory overhead due to large-scale parameters. This work strives to reduce memory overhead in fine-tuning from perspectives of activation function and layer normalization. To this end, we…

2023

Energy-Based Test Sample Adaptation for Domain Generalization

ICLR 2023poster

In this paper, we propose energy-based sample adaptation at test time for domain generalization. Where previous works adapt their models to target domains, we adapt the unseen target samples to source-trained models. To this end, we design a discriminative energy-based model, which is trained on sou…

2023

Episodic Multi-Task Learning with Heterogeneous Neural Processes

NeurIPS 2023spotlight

This paper focuses on the data-insufficiency problem in multi-task learning within an episodic training setup. Specifically, we explore the potential of heterogeneous information across tasks and meta-knowledge among episodes to effectively tackle each task with limited data. Existing meta-learning…

2023

Implicit Diffusion Models for Continuous Super-Resolution

CVPR 2023poster

Image super-resolution (SR) has attracted increasing attention due to its wide applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces an Implicit Diffusion Model (IDM) for high-fidelity continu…

2023

Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models

ICCV 2023poster

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great effectiveness in transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to overfit to seen classes while failing to generalize to unsee…

Cited by 37PDFScholar
2023

Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples

ICCV 2023poster

Referring video object segmentation (RVOS), as a supervised learning task, relies on sufficient annotated data for a given scene. However, in more realistic scenarios, only minimal annotations are available for a new scene, which poses significant challenges to existing RVOS methods. With this in mi…

Cited by 3PDFcodeScholar
2023

Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning

ICLR 2023poster

Multimodal few-shot learning is challenging due to the large domain gap between vision and language modalities. Existing methods are trying to communicate visual concepts as prompts to frozen language models, but rely on hand-engineered task induction to reduce the hypothesis space. To make the whol…

2023

MetaModulation: Learning Variational Feature Hierarchies for Few-Shot Learning with Fewer Tasks

ICML 2023poster

Meta-learning algorithms are able to learn a new task using previously learned knowledge, but they often require a large number of meta-training tasks which may not be readily available. To address this issue, we propose a method for few-shot learning with fewer tasks, which we call MetaModulation.…

2023

Order-preserving Consistency Regularization for Domain Adaptation and Generalization

ICCV 2023poster

Deep learning models fail on cross-domain challenges if the model is oversensitive to domain-specific attributes, e.g., lightning, background, camera angle, etc. To alleviate this problem, data augmentation coupled with consistency regularization are commonly adopted to make the model less sensitive…

Cited by 16PDFcodeScholar
2023

SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail

CVPR 2023poster

Modern image classifiers perform well on populated classes while degrading considerably on tail classes with only a few instances. Humans, by contrast, effortlessly handle the long-tailed recognition challenge, since they can learn the tail representation based on different levels of semantic abstra…

Cited by 18SourcePDFScholar
2022

Association Graph Learning for Multi-Task Classification with Category Shifts

NeurIPS 2022accept

In this paper, we focus on multi-task classification, where related classification tasks share the same label space and are learned simultaneously. In particular, we tackle a new setting, which is more realistic than currently addressed in the literature, where categories shift from training to test…

2022

Hierarchical Variational Memory for Few-shot Learning Across Domains

ICLR 2022poster

Neural memory enables fast adaptation to new tasks with just a few training samples. Existing memory models store features only from the single last layer, which does not generalize well in presence of a domain shift between training and test distributions. Rather than relying on a flat memory, we p…

2022

Learning to Generalize across Domains on Single Test Samples

ICLR 2022poster

We strive to learn a model from a set of source domains that generalizes well to unseen target domains. The main challenge in such a domain generalization scenario is the unavailability of any target domain data during training, resulting in the learned model not being explicitly adapted to the unse…

2022

Variational Model Perturbation for Source-Free Domain Adaptation

NeurIPS 2022accept

We aim for source-free domain adaptation, where the task is to deploy a model pre-trained on source domains to target domains. The challenges stem from the distribution shift from the source to the target domain, coupled with the unavailability of any source data and labeled target data for optimiza…

2021

A Bit More Bayesian: Domain-Invariant Learning with Uncertainty

ICML 2021spotlight

Domain generalization is challenging due to the domain shift and the uncertainty caused by the inaccessibility of target domain data. In this paper, we address both challenges with a probabilistic framework based on variational Bayesian inference, by incorporating uncertainty into neural network wei…

2021

Learning to Learn Dense Gaussian Processes for Few-Shot Learning

NeurIPS 2021poster

Gaussian processes with deep neural networks demonstrate to be a strong learner for few-shot learning since they combine the strength of deep learning and kernels while being able to well capture uncertainty. However, it remains an open problem to leverage the shared knowledge provided by related ta…

Cited by 32SourcePDFScholar
2021

Meta-Learning with Variational Semantic Memory for Word Sense Disambiguation

ACL 2021long

A critical challenge faced by supervised word sense disambiguation (WSD) is the lack of large annotated datasets with sufficient coverage of words in their diversity of senses. This inspired recent research on few-shot WSD using meta-learning. While such work has successfully applied meta-learning t…

2021

MetaNorm: Learning to Normalize Few-Shot Batches Across Domains

ICLR 2021poster

Batch normalization plays a crucial role when training deep neural networks. However, batch statistics become unstable with small batch sizes and are unreliable in the presence of distribution shifts. We propose MetaNorm, a simple yet effective meta-learning normalization. It tackles the aforementio…

Cited by 83SourcePDFScholar
2021

Seminar Learning for Click-Level Weakly Supervised Semantic Segmentation

ICCV 2021poster

Annotation burden has become one of the biggest barriers to semantic segmentation. Approaches based on click-level annotations have therefore attracted increasing attention due to their superior trade-off between supervision and annotation cost. In this paper, we propose seminar learning, a new lear…

Cited by 41PDFScholar
2021

Variational Multi-Task Learning with Gumbel-Softmax Priors

NeurIPS 2021poster

Multi-task learning aims to explore task relatedness to improve individual tasks, which is of particular significance in the challenging scenario that only limited data is available for each task. To tackle this challenge, we propose variational multi-task learning (VMTL), a general probabilistic in…

2020

Few-Shot Semantic Segmentation with Democratic Attention Networks

ECCV 2020poster

Few-shot segmentation has recently generated great popularity, addressing a challenging yet important problem of segmenting objects from unseen categories with scarce annotated support images. The crux of few-shot segmentation is to extract object information from the support image and then propagat…

Cited by 254SourcePDFScholar
2020

Learning to Learn Kernels with Variational Random Features

ICML 2020poster

We introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the late…

Cited by 34SourcePDFScholar
2020

Learning to Learn Variational Semantic Memory

NeurIPS 2020poster

In this paper, we introduce variational semantic memory into meta-learning to acquire long-term knowledge for few-shot learning. The variational semantic memory accrues and stores semantic information for the probabilistic inference of class prototypes in a hierarchical Bayesian framework. The seman…

2020

Learning to Learn with Variational Information Bottleneck for Domain Generalization

ECCV 2020poster

Domain generalization models learn to generalize to previously unseen domains, but suffer from prediction uncertainty and domain shift. In this paper, we address both problems. We introduce a probabilistic meta-learning model for domain generalization, in which classifier parameters shared across do…

Cited by 196SourcePDFScholar
2020

Transductive Relation-Propagation Network for Few-shot Learning

IJCAI 2020poster

Few-shot learning, aiming to learn novel concepts from few labeled examples, is an interesting and very challenging problem with many practical advantages. To accomplish this task, one should concentrate on revealing the accurate relations of the support-query pairs. We propose a transductive relati…

Cited by 0SourcePDFScholar
2019

Crowd Counting and Density Estimation by Trellis Encoder-Decoder Networks

CVPR 2019poster

Crowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-…

Cited by 435PDFScholar
2019

Relational Attention Network for Crowd Counting

ICCV 2019poster

Crowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision…

Cited by 214PDFScholar
2018

Direct Shape Regression Networks for End-to-End Face Alignment

CVPR 2018poster

Face alignment has been extensively studied in computer vision community due to its fundamental role in facial analysis, but it remains an unsolved problem. The major challenges lie in the highly nonlinear relationship between face images and associated facial shapes, which is coupled by underlying…

2017

Learning Deep Match Kernels for Image-Set Classification

CVPR 2017poster

Image-set classification has recently generated great popularity due to its widespread applications in computer vision. The great challenges arise from effectively and efficiently measuring the similarity between image sets with high inter-class ambiguity and huge intra-class variability. In this pa…

Cited by 49PDFScholar
2016

Realistic human action recognition: When deep learning meets VLAD

ICASSP 2016accepted

Human action recognition from realistic scenarios is extremely challenging due to large intra-class variation and complex background clutters. In this paper, by leveraging the strength of deep learning and vector of locally aggregated descriptors (VLAD), we propose a new methods for human action rec…

Cited by 0SourceScholar