← Search

Cheng Yu

14 accepted papers

2026

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

CVPR 2026

Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This

Cited by 0SourceScholar
2026

SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation

CVPR 2026

Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear

Cited by 0SourcecodeScholar
2025

Information Retrieval Induced Safety Degradation in AI Agents

NeurIPS 2025poster

Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the growing integration of LLMs and AI agents with external information sources and real-world environments raises critical questions about ho…

Cited by 0SourceScholar
2025

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

EMNLP 2025

While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improv

2025

Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition

CVPR 2025poster

Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos.Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents ho…

Cited by 0SourcePDFScholar
2023

Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers

ICASSP 2023accepted

Previous works reveal that similar to CNNs, vision transformers (ViT) are also vulnerable to universal adversarial patch attacks. In this paper, we empirically reveal and mathematically explain that the shallow tokens in the transformer and the attention of the network can largely influence the clas…

Cited by 0SourceScholar
2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

ICASSP 2022accepted

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech synthesis, they are still lagging behind in speech enhancement. This work leverage…

Cited by 0SourceScholar
2022

Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech

ACL 2022long

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE (CUC-VAE) is proposed to estimate a posterior probability distribution of the latent prosody features for each phoneme b…

2022

MetricGAN-U: Unsupervised Speech Enhancement/ Dereverberation Based Only on Noisy/ Reverberated Speech

ICASSP 2022accepted

Most of the deep learning-based speech enhancement models are learned in a supervised manner, which implies that pairs of noisy and clean speech are required during training. Consequently, several noisy speeches recorded in daily life cannot be used to train the model. Although certain unsupervised…

Cited by 0SourceScholar
2022

Speech Recovery For Real-World Self-Powered Intermittent Devices

ICASSP 2022accepted

The incompleteness of speech inputs severely degrades the performance of all the related speech signal processing applications. Although many researches have been proposed to address this issue, they controlled the data missing conditions by simulation with self-defined masking lengths or sizes. Bes…

Cited by 0SourceScholar
2021

Defending Against Universal Adversarial Patches by Clipping Feature Norms

ICCV 2021poster

Physical-world adversarial attacks based on universal adversarial patches have been proved to be able to mislead deep convolutional neural networks (CNNs), exposing the vulnerability of real-world visual classification systems based on CNNs. In this paper, we empirically reveal and mathematically ex…

Cited by 37PDFScholar
2019

Information Entropy Based Feature Pooling for Convolutional Neural Networks

ICCV 2019poster

In convolutional neural networks (CNNs), we propose to estimate the importance of a feature vector at a spatial location in the feature maps by the network's uncertainty on its class prediction, which can be quantified using the information entropy. Based on this idea, we propose the entropy-based f…

Cited by 40PDFScholar
2019

MVSCRF: Learning Multi-View Stereo With Conditional Random Fields

ICCV 2019poster

We present a deep-learning architecture for multi-view stereo with conditional random fields (MVSCRF). Given an arbitrary number of input images, we first use a U-shape neural network to extract deep features incorporating both global and local information, and then build a 3D cost volume for the re…

Cited by 108PDFScholar