← Search

Yudong Yang

8 accepted papers

2026

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

ICML 2026poster

Recent progress in large language models (LLMs) has enabled understanding of both speech and non-speech audio, but has also exposed new safety risks arising from complex audio inputs that are inadequately handled by current safeguards. We introduce SACRED-Bench (Speech–Audio Composition for RED-team…

Cited by 0SourceScholar
2026

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

ICASSP 2026oral

Speech therapy is essential for rehabilitating speech disorders caused by neurological impairments such as stroke. However, traditional manual and computer-assisted systems are limited in real-time accessibility and articulatory motion feedback. Recent advances in multimodal large language models (M…

Cited by 0SourcePDFScholar
2026

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

ICML 2026poster

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language model that processes over 3-hour videos at $1$ FPS and $360$p resolution, outper…

Cited by 0SourceScholar
2025

Audio-centric Video Understanding Benchmark without Text Shortcut

EMNLP 2025

Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly depends on auditory information, as audio offers critical cont

2025

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

ICASSP 2025accepted

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced audito…

Cited by 0SourceScholar
2025

Improving LLM Video Understanding with 16 Frames Per Second

ICML 2025poster

Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual informat…

2025

video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

ICML 2025poster

While recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in gener…

2024

An Audio-Textual Diffusion Model for Converting Speech Signals into Ultrasound Tongue Imaging Data

ICASSP 2024accepted

Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the general patterns of tongue motions, and thus the quality of genera…

Cited by 0SourceScholar