← Search

Cheng Lu

49 accepted papers

2026

DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation

CVPR 2026

Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes as geometric conditions in diffusion models for conditional

Cited by 0SourceScholar
2025

Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known Attributes

ICASSP 2025accepted

Adapting pre-trained models has become the dominant approach for anomalous sound detection (ASD), where classifying the attributes of machine working status is commonly chosen as the deputy task for fine-tuning. However, attributes might be intractable to collect for some machines, causing the label…

Cited by 0SourceScholar
2025

Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning

ICASSP 2025accepted

The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability o…

Cited by 0SourceScholar
2025

Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

This paper introduces our solution, the Task-Specific Feature Learning (TSFL) method, designed to address the second track of the MEIJU Challenge at ICASSP 2025, namely, Imbalanced Emotion and Intent Recognition (English). The TSFL method incorporates three core components: the use of LLM features t…

Cited by 0SourceScholar
2025

Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration Prediction

ICASSP 2025accepted

Zero-shot Emotional Voice Conversion (EVC) aims to transform a speaker’s emotional state to match a target emotion, even for speakers and emotion categories that were not encountered during training, thereby enhancing the generalization ability of traditional EVC systems. Despite advancements in the…

Cited by 4SourceScholar
2025

Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

This paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from t…

Cited by 0SourceScholar
2024

Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition

ICASSP 2024accepted

Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This p…

Cited by 0SourceScholar
2024

Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection

ICASSP 2024accepted

Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages…

Cited by 0SourceScholar
2024

Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution Adaptation

ICASSP 2024accepted

In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new s…

Cited by 0SourceScholar
2024

PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice Conversion

ICASSP 2024accepted

In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of con…

Cited by 0SourceScholar
2024

Privileged Prior Information Distillation for Image Matting

AAAI 2024technical

Performance of trimap-free image matting methods is limited when trying to decouple the deterministic and undetermined regions, especially in the scenes where foregrounds are semantically ambiguous, chromaless, or high transmittance. In this paper, we propose a novel framework named Privileged Prior…

Cited by 1SourcePDFScholar
2024

Progressively Learning from Macro-Expressions for Micro-Expression Recognition

ICASSP 2024accepted

Micro-expression (ME) recognition is challenging due to the low-intensity facial motions. An idea to overcome this is learning assisted by macro-expressions (MaEs). However, the intensity gap between MaE and ME is so huge that related works fail to effectively leverage MaE’s assistance in overcoming…

Cited by 0SourceScholar
2024

Score Regularized Policy Optimization through Diffusion Behavior

ICLR 2024poster

Recent developments in offline reinforcement learning have uncovered the immense potential of diffusion modeling, which excels at representing heterogeneous behavior policies. However, sampling from diffusion policies is considerably slow because it necessitates tens to hundreds of iterative inferen…

2024

Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition

ICASSP 2024accepted

Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above in…

Cited by 0SourceScholar
2024

The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing

ICLR 2024poster

We present a unified probabilistic formulation for diffusion-based image editing, where a latent variable is edited in a task-specific manner and generally deviates from the corresponding marginal distribution induced by the original stochastic or ordinary differential equation (SDE or ODE). Instead…

2024

Towards Efficient Exact Optimization of Language Model Alignment

ICML 2024poster

The alignment of language models with human preferences is vital for their application in real-world tasks. The problem is formulated as optimizing the model's policy to maximize the expected reward that reflects human preferences with minimal deviation from the initial policy. While considered as a…

2023

CMNet: Contrastive Magnification Network for Micro-Expression Recognition

AAAI 2023technical

Micro-Expression Recognition (MER) is challenging because the Micro-Expressions' (ME) motion is too weak to distinguish. This hurdle can be tackled by enhancing intensity for a more accurate acquisition of movements. However, existing magnification strategies tend to use the features of facial image…

Cited by 5SourcePDFScholar
2023

Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement Learning

ICML 2023poster

Guided sampling is a vital approach for applying diffusion models in real-world tasks that embeds human-defined guidance during the sampling procedure. This paper considers a general setting where the guidance is defined by an (unnormalized) energy function. The main challenge for this setting is th…

2023

DPM-Solver-v3: Improved Diffusion ODE Solver with Empirical Model Statistics

NeurIPS 2023poster

Diffusion probabilistic models (DPMs) have exhibited excellent performance for high-fidelity image generation while suffering from inefficient sampling. Recent works accelerate the sampling procedure by proposing fast ODE solvers that leverage the specific ODE form of DPMs. However, they highly rely…

2023

Domain Adaptation with Adversarial Training on Penultimate Activations

AAAI 2023technical

Enhancing model prediction confidence on target data is an important objective in Unsupervised Domain Adaptation (UDA). In this paper, we explore adversarial training on penultimate activations, i.e., input features of the final linear classification layer. We show that this strategy is more efficie…

2023

Gaussian Mixture Solvers for Diffusion Models

NeurIPS 2023poster

Recently, diffusion models have achieved great success in generative tasks. Sampling from diffusion models is equivalent to solving the reverse diffusion stochastic differential equations (SDEs) or the corresponding probability flow ordinary differential equations (ODEs). In comparison, SDE-based so…

2023

Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs

ICML 2023poster

Diffusion models have exhibited excellent performance in various domains. The probability flow ordinary differential equation (ODE) of diffusion models (i.e., diffusion ODEs) is a particular case of continuous normalizing flows (CNFs), which enables deterministic inference and exact likelihood evalu…

2023

Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

ICLR 2023poster

In offline reinforcement learning, weighted regression is a common method to ensure the learned policy stays close to the behavior policy and to prevent selecting out-of-sample actions. In this work, we show that due to the limited distributional expressivity of policy models, previous methods might…

2023

On Calibrating Diffusion Probabilistic Models

NeurIPS 2023poster

Recently, diffusion probabilistic models (DPMs) have achieved promising results in diverse generative tasks. A typical DPM framework includes a forward process that gradually diffuses the data distribution and a reverse process that recovers the data distribution from time-dependent data scores. In…

2023

POINTACL: Adversarial Contrastive Learning for Robust Point Clouds Representation Under Adversarial Attack

ICASSP 2023accepted

Adversarial contrastive learning (ACL) is considered an effective way to improve the robustness of pre-trained models. In contrastive learning, a projector which consists of multilayer perceptron (MLP) will project high dimension 3D point cloud feature into low dimension for calculating contrastive…

Cited by 0SourceScholar
2023

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

NeurIPS 2023spotlight

Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity problems. In this work, we propose to model the 3D parameter as a random variabl…

2023

Ultra Real-Time Portrait Matting via Parallel Semantic Guidance

ICASSP 2023accepted

Most existing portrait matting models either require expensive auxiliary information or try to decompose the task into sub-tasks that are usually resource-hungry. These challenges limit its application on low-power computing devices. In this paper, we propose an ultra-light-weighted portrait matting…

Cited by 0SourceScholar
2022

3DG-STFM: 3D Geometric Guided Student-Teacher Feature Matching

ECCV 2022poster

"We tackle the essential task of finding dense visual correspondences between a pair of images. This is a challenging problem due to various factors such as poor texture, repetitive patterns, illumination variation, and motion blur in practical scenarios. In contrast to methods that use dense corres…

2022

A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive Networks

ICASSP 2022accepted

Micro-Expression recognition (MER) is a challenging task due to the short duration and low intensity of Micro-Expressions. A popular method to tackle this is magnifying MEs so as to enlarge the expression intensity to make recognition easier. However, the single fixed magnification strategy, widely…

Cited by 0SourceScholar
2022

DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps

NeurIPS 2022accept

Diffusion probabilistic models (DPMs) are emerging powerful generative models. Despite their high-quality generation performance, DPMs still suffer from their slow sampling as they generally need hundreds or thousands of sequential function evaluations (steps) of large neural networks to draw a samp…

2022

Maximum Likelihood Training for Score-based Diffusion ODEs by High Order Denoising Score Matching

ICML 2022spotlight

Score-based generative models have excellent performance in terms of generation quality and likelihood. They model the data distribution by matching a parameterized score network with first-order data score functions. The score network can be used to define an ODE (“score-based diffusion ODE”) for e…

2022

SDETR: Attention-Guided Salient Object Detection with Transformer

ICASSP 2022accepted

Most existing CNN-based salient object detection methods can identify fine-grained segmentation details like hair and animal fur, but often mispredict the salient object due to lack of global contextual information caused by locality convolution layers. The limited training data of the current SOD t…

Cited by 0SourceScholar
2021

Virtual Multi-Modality Self-Supervised Foreground Matting for Human-Object Interaction

ICCV 2021poster

Most existing human matting algorithms tried to separate pure human-only foreground from the background. In this paper, we propose a Virtual Multi-modality Foreground Matting (VMFM) method to learn human-object interactive foreground (human and objects interacted with him or her) from a raw RGB imag…

Cited by 7PDFcodeScholar
2020

VFlow: More Expressive Generative Flows with Variational Data Augmentation

ICML 2020poster

Generative flows are promising tractable models for density modeling that define probabilistic distributions with invertible transformations. However, tractability imposes architectural constraints on generative flows. In this work, we study a previously overlooked constraint that all the intermedia…

2019

On the Equivalence of Semidifinite Relaxations for MIMO Detection with General Constellations

ICASSP 2019accepted

The multiple-input multiple-output (MIMO) detection problem is a fundamental problem in modern digital communications. Semidefinite relaxation (SDR) based algorithms are a popular class of approaches to solving the problem because the algorithms have a polynomial-time worst-case complexity and gener…

Cited by 0SourceScholar
2019

Staying up to Date with Online Content Changes Using Reinforcement Learning for Scheduling

NeurIPS 2019poster

From traditional Web search engines to virtual assistants and Web accelerators, services that rely on online information need to continually keep track of remote content changes by explicitly requesting content updates from remote sources (e.g., web pages). We propose a novel optimization objective…

2017

Model-Based Iterative Restoration for Binary Document Image Compression With Dictionary Learning

CVPR 2017poster

The inherent noise in the observed (e.g., scanned) binary document image degrades the image quality and harms the compression ratio through breaking the pattern repentance and adding entropy to the document images. In this paper, we design a cost function in Bayesian framework with dictionary learni…

Cited by 10PDFScholar
2016

Tag recommendation via robust probabilistic discriminative matrix factorization

ICASSP 2016accepted

Low-rank matrix factorization serves as a key technique in learning latent factor models for many applications in machine learning. However, in many applications, observed data often exhibits different levels of noise. To address this issue, we propose a Robust Probabilistic Discriminative Matrix Fa…

Cited by 0SourceScholar
2015

Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognition

ICASSP 2015accepted

Exploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold…

Cited by 0SourceScholar
2015

The THUEE system for the openKWS14 keyword search evaluation

ICASSP 2015accepted

The OpenKWS14 keyword search evaluation is one of the most challenging and influential evaluations in the field of speech recognition. Its goal is to build a high-performance keyword search system for a minority language with limited training data in a short period of time. We present the system of…

Cited by 0SourceScholar