← Search

Wenyi Yu

11 accepted papers

2026

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

AAAI 2026technical

Audio comprehension—including speech, non-speech sounds, and music—is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challen

Cited by 0SourcePDFScholar
2025

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

ICASSP 2025accepted

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced audito…

Cited by 0SourceScholar
2025

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

ACL 2025long

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing d…

2025

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

NeurIPS 2025poster

In order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such as voice activity detectors, interrupters, conversation state predictors, or multiple LLMs. These systems, however, suff…

Cited by 0SourcecodeScholar
2024

An Improved Empirical Fisher Approximation for Natural Gradient Descent

NeurIPS 2024poster

Approximate Natural Gradient Descent (NGD) methods are an important family of optimisers for deep learning models, which use approximate Fisher information matrices to pre-condition gradients during training. The empirical Fisher (EF) method approximates the Fisher information matrix empirically by…

Cited by 4SourcePDFScholar
2024

Connecting Speech Encoder and Large Language Model for ASR

ICASSP 2024accepted

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative s…

Cited by 0SourceScholar
2024

Extending Large Language Models for Speech and Audio Captioning

ICASSP 2024accepted

Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achie…

Cited by 0SourceScholar
2024

M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset

ACL 2024long

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as the texts and pictures in the slides and possibly even the pap…

2024

SALMONN: Towards Generic Hearing Abilities for Large Language Models

ICLR 2024poster

Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a…

2024

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

ICML 2024poster

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences…