← Search

Zhigang Zeng

6 accepted papers

2025

Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization

ICLR 2025poster

Extracting a small subset of crucial rationales from the full input is a key problem in explainability research. The most widely used fundamental criterion for rationale extraction is the maximum mutual information (MMI) criterion. In this paper, we first demonstrate that MMI suffers from diminishin…

2025

Language-Driven Multi-Label Zero-Shot Learning with Semantic Granularity

ICCV 2025poster

Recent methods learn class-unified prompt contexts by image data to adapt CLIP to zero-shot multi-label image classification, which achieves impressive performance. However, simply tuning prompts is insufficient to deal with novel classes across different semantic granularity levels. This limitation…

2024

Harnessing Neural Unit Dynamics for Effective and Scalable Class-Incremental Learning

ICML 2024poster

Class-incremental learning (CIL) aims to train a model to learn new classes from non-stationary data streams without forgetting old ones. In this paper, we propose a new kind of connectionist model by tailoring neural unit dynamics that adapt the behavior of neural networks for CIL. In each training…

Cited by 4SourcePDFScholar
2024

Towards Continual Learning Desiderata via HSIC-Bottleneck Orthogonalization and Equiangular Embedding

AAAI 2024technical

Deep neural networks are susceptible to catastrophic forgetting when trained on sequential tasks. Various continual learning (CL) methods often rely on exemplar buffers or/and network expansion for balancing model stability and plasticity, which, however, compromises their practical value due to pri…

Cited by 9SourcePDFScholar
2023

DaFKD: Domain-Aware Federated Knowledge Distillation

CVPR 2023poster

Federated Distillation (FD) has recently attracted increasing attention for its efficiency in aggregating multiple diverse local models trained from statistically heterogeneous data of distributed clients. Existing FD methods generally treat these models equally by merely computing the average of th…

Cited by 78SourcePDFScholar