← Search

Cheng Zhong

9 accepted papers

2025

CODA: Repurposing Continuous VAEs for Discrete Tokenization

ICCV 2025poster

Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challenging, as it requires both compressing visual signals into a compact representation and discretizing them into a fixed set…

Cited by 0SourcePDFScholar
2025

EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving

NeurIPS 2025poster

We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 18…

Cited by 0SourceScholar
2025

Graph Pooling via Dropping Task-Irrelevant Nodes

ICASSP 2025accepted

Graph neural networks (GNNs) face scalability challenges. While recent approaches have adopted pooling strategies inspired by convolutional neural networks (CNNs) to reduce graph size and improve efficiency, these methods often focus on local information and are optimized for single graph-level task…

Cited by 0SourceScholar
2025

Melody Structure Transfer Network: Generating Music with Separable Self-Attention

ICASSP 2025accepted

Most existing symbolic music generation methods focus on generating short pieces, typically less than 8 bars and occasionally up to 32 bars. Generating long music sequences requires effective representation of coherent musical structures. Vanilla self-attention face challenges in capturing subtle lo…

Cited by 0SourceScholar
2025

Towards Green VAE: A Light Pixel-weighting Technique to Enhance Variational AutoEncoder

ICASSP 2025accepted

Variational autoencoders (VAEs) has been a popular generative model for its effectiveness, mathematical foundation, and its impact to other approaches in deep generative learning. For its relatively light-weights and easiness for training, compared with Generative Adversarial Networks (GANs) or othe…

Cited by 0SourceScholar
2024

Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video Analysis

AAAI 2024technical

Video Correlation Learning (VCL), which aims to analyze the relationships between videos, has been widely studied and applied in various general video tasks. However, applying VCL to instructional videos is still quite challenging due to their intrinsic procedural temporal structure. Specifically, p…

Cited by 5SourcePDFScholar
2024

MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

NeurIPS 2024spotlight

Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on short videos containing only a single event and simple causal…

2024

TimeCraft: Navigate Weakly-Supervised Temporal Grounded Video Question Answering via Bi-directional Reasoning

ECCV 2024poster

"Video reasoning typically operates within the Video Question-Answering (VQA) paradigm, which demands that the models understand and reason about video content from temporal and causal perspectives. Traditional supervised VQA methods gain this capability through meticulously annotated QA datasets, w…

2023

Learning How to Learn Domain-Invariant Parameters for Domain Generalization

ICASSP 2023accepted

Due to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs a…

Cited by 0SourceScholar