← Search

Zhengjue Wang

11 accepted papers

2025

Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models

CVPR 2025poster

Concept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image…

Cited by 1SourcePDFScholar
2025

Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification

CVPR 2025poster

Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequentl…

2024

MeaCap: Memory-Augmented Zero-shot Image Captioning

CVPR 2024poster

Zero-shot image captioning (IC) without well-paired image-text data can be categorized into two main types: training-free and text-only-training methods. While both types integrate pre-trained vision-language models such as CLIP for image-text similarity evaluation and a pre-trained language model (…

2023

ConZIC: Controllable Zero-Shot Image Captioning by Sampling-Based Polishing

CVPR 2023poster

Zero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC), ZeroCap abandons supervised training and sequentially searching every word in…

2021

EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering

ACL 2021long

Natural language processing (NLP) often faces the problem of data diversity such as different domains, themes, styles, and so on. Therefore, a single language model (LM) is insufficient to learn all knowledge from diverse samples. To solve this problem, we firstly propose an autoencoding topic model…

2021

Memory-Efficient Network for Large-Scale Video Compressive Sensing

CVPR 2021poster

Video snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimizatio…

Cited by 93PDFcodeScholar
2021

MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive Sensing

CVPR 2021poster

To capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-s…

Cited by 68PDFcodeScholar
2020

BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging

ECCV 2020poster

We consider the problem of video snapshot compressive imaging (SCI), where multiple high-speed frames are coded by different masks and then summed to a single measurement. This measurement and the modulation masks are fed into our Recurrent Neural Network (RNN) to reconstruct the desired high-speed…

2020

Deep Relational Topic Modeling via Graph Poisson Gamma Belief Network

NeurIPS 2020poster

To analyze a collection of interconnected documents, relational topic models (RTMs) have been developed to describe both the link structure and document content, exploring their underlying relationships via a single-layer latent representation with limited expressive capability. To better utilize th…

2020

Learning Dynamic Hierarchical Topic Graph with Graph Convolutional Network for Document Classification

AISTATS 2020poster

Constructing a graph with graph convolutional network (GCN) to explore the relational structure of the data has attracted lots of interests in various tasks. However, for document classification, existing graph based methods often focus on the straightforward word-word and word-document relations,…

2020

Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling

ICLR 2020poster

For bidirectional joint image-text modeling, we develop variational hetero-encoder (VHE) randomized generative adversarial network (GAN), a versatile deep generative model that integrates a probabilistic text decoder, probabilistic image encoder, and GAN into a coherent end-to-end multi-modality lea…

Cited by 1SourcecodeScholar