← Search

Guangzhi Sun

32 accepted papers

2026

AUDIO-CONDITIONED DIFFUSION LLMS FOR ASR AND DELIBERATION PROCESSING

ICASSP 2026poster

Diffusion-based large language models (DLLMs) have recently attracted growing interest as an alternative to autoregressive decoders. In this work, we present an empirical study on using the diffusion-based large language model LLaDA for automatic speech recognition (ASR). We first investigate its us…

Cited by 0SourcePDFScholar
2026

LOW-RANK AND SPARSE MODEL MERGING FOR MULTI-LINGUAL SPEECH RECOGNITION AND TRANSLATION

ICASSP 2026poster

Language diversity presents a significant challenge in speech-to-text (S2T) tasks, such as automatic speech recognition and translation. Traditional multi-lingual multi-task training approaches aim to address this by jointly optimising multiple speech recognition and translation tasks across various…

Cited by 0SourcePDFScholar
2026

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

ICML 2026poster

Recent progress in large language models (LLMs) has enabled understanding of both speech and non-speech audio, but has also exposed new safety risks arising from complex audio inputs that are inadequately handled by current safeguards. We introduce SACRED-Bench (Speech–Audio Composition for RED-team…

Cited by 0SourceScholar
2026

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

ICML 2026poster

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language model that processes over 3-hour videos at $1$ FPS and $360$p resolution, outper…

Cited by 0SourceScholar
2025

Audio-centric Video Understanding Benchmark without Text Shortcut

EMNLP 2025

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly depends on auditory information, as audio offers critical cont

2025

Bayesian WeakS-to-Strong from Text Classification to Generation

ICLR 2025poster

Advances in large language models raise the question of how alignment techniques will adapt as models become increasingly complex and humans will only be able to supervise them weakly. Weak-to-Strong mimics such a scenario where weak model supervision attempts to harness the full capabilities of a m…

Cited by 0SourcePDFScholar
2025

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

ICML 2025poster

Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety benchmarks often focus solely on the refusal of individual problematic queries, which overlooks the importance of the context where the query occurs and may caus…

Cited by 0SourcePDFScholar
2025

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

ICASSP 2025accepted

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced audito…

Cited by 0SourceScholar
2025

Improving LLM Video Understanding with 16 Frames Per Second

ICML 2025poster

Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual informat…

2025

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

NeurIPS 2025poster

In order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such as voice activity detectors, interrupters, conversation state predictors, or multiple LLMs. These systems, however, suff…

Cited by 0SourcecodeScholar
2025

SkillAggregation: Reference-free LLM-Dependent Aggregation

ACL 2025long

Large Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments. Single LLMs were used initially, however, recent work suggests using multiple LLMs as judges yields improved performance. An important step in exploiting multiple judgements…

2025

Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?

EMNLP 2025

Unlearning has emerged as a critical capability for large language models (LLMs) to support data privacy, regulatory compliance, and ethical AI deployment. Recent techniques often rely on obfuscation by injecting incorrect or irrelevant information to suppress knowledge. Such methods effectively con

2025

Wav2Prompt: End-to-End Speech Prompt Learning and Task-based Fine-tuning for Text-based LLMs

NAACL 2025long

Wav2Prompt is proposed which allows integrating spoken input with a text-based large language model (LLM). Wav2Prompt uses a straightforward training process with only the same data used to train an automatic speech recognition (ASR) model. After training, Wav2Prompt learns continuous representation…

Cited by 1SourcePDFScholar
2025

video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

ICML 2025poster

While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in gener…

2024

Connecting Speech Encoder and Large Language Model for ASR

ICASSP 2024accepted

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative s…

Cited by 0SourceScholar
2024

Enhancing Quantised End-to-End ASR Models Via Personalisation

ICASSP 2024accepted

Recent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Model quantisation is an effective solution that sometimes causes the word error rate (WER) to increase. In this paper, a…

Cited by 4SourceScholar
2024

Extending Large Language Models for Speech and Audio Captioning

ICASSP 2024accepted

Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achie…

Cited by 0SourceScholar
2024

M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset

ACL 2024long

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as the texts and pictures in the slides and possibly even the pap…

2024

Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation

ICASSP 2024accepted

Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-e…

Cited by 0SourceScholar
2024

SALMONN: Towards Generic Hearing Abilities for Large Language Models

ICLR 2024poster

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a…

2024

Speech-based Slot Filling using Large Language Models

ACL 2024findings

Recently, advancements in large language models (LLMs) have shown an unprecedented ability across various language tasks. This paper investigates the potential application of LLMs to slot filling with noisy ASR transcriptions, via both in-context learning and task-specific fine-tuning. Dedicated pro…

2024

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

ICML 2024poster

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences…

2023

End-to-End Spoken Language Understanding with Tree-Constrained Pointer Generator

ICASSP 2023accepted

End-to-end spoken language understanding (SLU) suffers from the long-tail word problem. This paper exploits contextual biasing, a technique to improve the speech recognition of rare words, in end-to-end SLU systems. Specifically, a tree-constrained pointer generator (TCPGen), a powerful and efficien…

Cited by 0SourceScholar
2023

Spectral Clustering-Aware Learning of Embeddings for Speaker Diarisation

ICASSP 2023accepted

In speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning of embeddings (SCALE) to address the mismatch. Specifically,…

Cited by 1SourceScholar
2022

Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech

ACL 2022long

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE (CUC-VAE) is proposed to estimate a posterior probability distribution of the latent prosody features for each phoneme b…

2021

Transformer Language Models with LSTM-Based Cross-Utterance Information Representation

ICASSP 2021accepted

The effective incorporation of cross-utterance information has the potential to improve language models (LMs) for automatic speech recognition (ASR). To extract more powerful and robust cross-utterance representations for the Transformer LM (TLM), this paper proposes the R-TLM which uses hidden stat…

Cited by 0SourceScholar
2020

Fully-Hierarchical Fine-Grained Prosody Modeling For Interpretable Speech Synthesis

ICASSP 2020accepted

This paper proposes a hierarchical, fine-grained and interpretable latent variable model for prosody based on the Tacotron 2 text-to-speech model. It achieves multi-resolution modeling of prosody by conditioning finer level representations on coarser level ones. Additionally, it imposes hierarchical…

Cited by 0SourceScholar
2020

Generating Diverse and Natural Text-to-Speech Samples Using a Quantized Fine-Grained VAE and Autoregressive Prosody Prior

ICASSP 2020accepted

Recent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting latent features at each input token (e.g., phonemes). However,…

Cited by 0SourceScholar