← Search

Yujun Wang

18 accepted papers

2026

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM

AAAI 2026technical

Multimodal large language models (MLLMs) frequently hallucinate by over-committing to spurious visual cues. Prior remedies–Visual and Instruction Contrastive Decoding (VCD, ICD)–mitigate this issue, yet the mechanism remains opaque. We first empirically show that their improvements systematically co

Cited by 41SourcePDFScholar
2026

EchoRL: Reinforcement Learning via Rollout Echoing

ICML 2026poster

Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models. However, as training proceeds, the learning signal can collapse thus makes the training gain become marginal and ineffective. Specifically, a growin…

Cited by 0SourceScholar
2025

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering

ACL 2025long

Multimodal Large Language Models (MLLMs) enhance visual tasks by integrating visual representations into large language models (LLMs). The textual modality, inherited from LLMs, enables instruction following and in-context learning, while the visual modality boosts downstream task performance throug…

2024

CED: Consistent Ensemble Distillation for Audio Tagging

ICASSP 2024accepted

Augmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the widely recognized Audioset (AS) benchmark. Although both techniques are effective individually, their combined use, call…

Cited by 0SourceScholar
2024

TNFormer: Single-Pass Multilingual Text Normalization with a Transformer Decoder Model

ICASSP 2024accepted

Text Normalization (TN) is a pivotal pre-processing procedure in speech synthesis systems, which converts diverse forms of text into a canonical form suitable for correct synthesis. This work introduces a novel model, TNFormer, which innovatively transforms the TN task into a next token prediction p…

Cited by 0SourceScholar
2023

Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction

ICASSP 2023accepted

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio an…

Cited by 0SourceScholar
2023

Improving Weakly Supervised Sound Event Detection with Causal Intervention

ICASSP 2023accepted

Existing weakly supervised sound event detection (WSSED) work has not explored both types of co-occurrences simultaneously, i.e., some sound events often co-occur, and their occurrences are usually accompanied by specific background sounds, so they would be inevitably entangled, causing misclassific…

Cited by 0SourceScholar
2023

Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based Network

ICASSP 2023accepted

Electroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDANet) to classify the correct auditory stimulus evoking the EEG signal. It adopts the Attention-based Correlation Module (…

Cited by 0SourceScholar
2023

Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers

ICASSP 2023accepted

Keyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that adds additional capabilities in the form of audio tagging (AT) to a KWS model. However, previous work did not consider the…

Cited by 0SourceScholar
2022

Improving Emotional Speech Synthesis by Using SUS-Constrained VAE and Text Encoder Aggregation

ICASSP 2022accepted

Learning emotion embedding from reference audio is a straightforward approach for multi-emotion speech synthesis in encoder-decoder systems. But how to get better emotion embedding and how to inject it into TTS acoustic model more effectively are still under investigation. In this paper, we propose…

Cited by 0SourceScholar
2022

Learning Decoupling Features Through Orthogonality Regularization

ICASSP 2022accepted

Keyword spotting (KWS) and speaker verification (SV) are two important tasks in speech applications. Research shows that the state-of-art KWS and SV models are trained independently using different datasets since they expect to learn distinctive acoustic features. However, humans can distinguish lan…

Cited by 0SourceScholar
2022

MSDTRON: A High-Capability Multi-Speaker Speech Synthesis System for Diverse Data Using Characteristic Information

ICASSP 2022accepted

In multi-speaker speech synthesis, data from a number of speakers usually tend to have great diversity due to the fact that the speakers may differ largely in ages, speaking styles, emotions, and so on. It is important but challenging to improve the modeling capabilities for multi-speaker speech syn…

Cited by 0SourceScholar
2022

Multi-Scale Refinement Network Based Acoustic Echo Cancellation

ICASSP 2022accepted

Recently, deep encoder-decoder networks have shown outstanding performance in acoustic echo cancellation (AEC). However, the subsampling operations like convolution striding in the encoder layers significantly decrease the feature resolution lead to fine-grained information loss. This paper proposes…

Cited by 0SourceScholar
2022

PAMA-TTS: Progression-Aware Monotonic Attention for Stable SEQ2SEQ TTS with Accurate Phoneme Duration Control

ICASSP 2022accepted

Sequence expansion between encoder and decoder is a critical challenge in sequence-to-sequence TTS. Attention-based methods achieve great naturalness but suffer from unstable issues like missing and repeating phonemes, not to mention accurate duration control. Duration-informed methods, on the contr…

Cited by 0SourceScholar
2022

Pseudo Strong Labels for Large Scale Weakly Supervised Audio Tagging

ICASSP 2022accepted

Large-scale audio tagging datasets inevitably contain imperfect labels, such as clip-wise annotated (temporally weak) tags with no exact on- and offsets, due to a high manual labeling cost. This work proposes pseudo strong labels (PSL), a simple label augmentation framework that enhances the supervi…

Cited by 8SourceScholar
2021

AutoKWS: Keyword Spotting with Differentiable Architecture Search

ICASSP 2021accepted

Smart audio devices are gated by an always-on lightweight keyword spotting program to reduce power consumption. It is however challenging to design models that have both high accuracy and low latency for accurate and fast responsiveness. Many efforts have been made to develop end-to-end neural netwo…

Cited by 0SourceScholar