← Search

Chao-Han Huck Yang

58 accepted papers

2026

Long Grounded Thoughts: Synthesizing Grounded Visual Problems and Distilling Reasoning Chains at Scale

ICML 2026poster

Despite rapid progress, multimodal reasoning still lacks a systematic approach to synthesize large-scale vision-centric datasets beyond visual math. We introduce a framework able to synthesize vision-centric problems spanning diverse levels of complexity, and the resulting dataset with over 1M high-…

Cited by 0SourceScholar
2026

Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning

ICASSP 2026oral

We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test audio-language models on interactive question-answering over di…

Cited by 0SourcePDFScholar
2026

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

ICLR 2026poster

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curati…

Cited by 0SourcecodeScholar
2026

Pre-training Tensor-Train Networks Facilitates Machine Learning with Variational Quantum Circuits

ICASSP 2026oral

Data encoding remains a fundamental bottleneck in quantum machine learning, where amplitude encoding of high-dimensional classical vectors into quantum states incurs exponential cost. In this work, we propose a pre-trained tensor-train (TT) encoding network (Pre-TT-Encoder) that significantly reduce…

Cited by 0SourcePDFScholar
2026

Test-Time Alignment for Large Language Models via Textual Model Predictive Control

ICLR 2026poster

Aligning Large Language Models (LLMs) with human preferences through finetuning is resource-intensive, motivating lightweight alternatives at test time. We address test-time alignment through the lens of sequential decision making, a perspective that reveals two fundamental challenges. When actions…

Cited by 0SourceScholar
2026

TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models

ICLR 2026poster

Recent advances in multimodal time series learning underscore a paradigm shift from analytics centered on basic patterns toward advanced time series understanding and reasoning. However, existing multimodal time series datasets mostly remain at the level of surface alignment and question answering,…

Cited by 0SourcecodeScholar
2026

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

ICLR 2026oral

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning. This paper introduces Unified Audio Language Model (UALM), w…

Cited by 0SourcecodeScholar
2025

A Quantum Circuit-Based Compression Perspective for Parameter-Efficient Learning

ICLR 2025poster

Quantum-centric supercomputing presents a compelling framework for large-scale hybrid quantum-classical tasks. Although quantum machine learning (QML) offers theoretical benefits in various applications, challenges such as large-size data encoding in the input stage and the reliance on quantum resou…

Cited by 5SourcePDFScholar
2025

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

NeurIPS 2025spotlight

We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning acros…

Cited by 0SourcecodeScholar
2025

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

ICLR 2025poster

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. Th…

Cited by 1SourcePDFScholar
2025

Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation

ICCV 2025poster

Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such…

Cited by 0SourcePDFScholar
2025

Chain-of-Thought Prompting for Speech Translation

ICASSP 2025accepted

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance…

Cited by 0SourceScholar
2025

CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Models

EMNLP 2025

Large language models (LLMs) can rewrite the N-best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot. Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving

2025

Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

ICASSP 2025accepted

Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires si…

Cited by 0SourceScholar
2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

ICLR 2025poster

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication…

2025

ESPnet-SpeechLM: An Open Speech Language Model Toolkit

NAACL 2025system demonstrations

We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing them as universal sequential modeling problems, encompassing a cohesive workflo…

2025

Extending Automatic Machine Translation Evaluation to Book-Length Documents

EMNLP 2025

Despite Large Language Models (LLMs) demonstrating superior translation performance and long-context capabilities, evaluation methodologies remain constrained to sentence-level assessment due to dataset limitations, token number restrictions in metrics, and rigid sentence boundary requirements. We i

2025

Fugatto 1: Foundational Generative Audio Transformer Opus 1

ICLR 2025poster

Fugatto is a versatile audio synthesis and transformation model capable of following free-form text instructions with optional audio inputs. While large language models (LLMs) trained with text on a simple next-token prediction objective can learn to infer instructions directly from the data, models…

2025

MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and Summarization

ACL 2025long

We introduce MISP-Meeting, a new real-world, multimodal dataset that covers subject-oriented long-form content. MISP-Meeting integrates information from speech, vision, and text modalities to facilitate automatic meeting transcription and summarization (AMTS). Challenging conditions in human meeting…

2025

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

ICML 2025poster

Neural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In this work, we introduce OWLS, an open-access, reproducible su…

Cited by 1SourcePDFScholar
2025

Projection Valued-based Quantum Machine Learning Adapting to Differential Privacy Algorithm for Word-level Lipreading

ICASSP 2025accepted

Deep neural network (DNN)-based lipreading models have achieved excellent recognition accuracy but are currently facing challenges related to user privacy. To address this, we propose a novel hybrid quantum-classical neural network (HQCNN) for lipreading that balances superior performance with enhan…

Cited by 0SourceScholar
2025

Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction

ICASSP 2025accepted

Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts…

Cited by 0SourceScholar
2025

SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models

ACL 2025long

We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models (LLM_Voice), designed to assess their voice understanding ability. Moving beyond popular voice understanding metrics such as word error ra…

2025

Towards Neural Scaling Laws for Time Series Foundation Models

ICLR 2025poster

Scaling laws offer valuable insights into the design of time series foundation models (TSFMs). However, previous research has largely focused on the scaling laws of TSFMs for in-distribution (ID) data, leaving their out-of-distribution (OOD) scaling behavior and the influence of model architectures…

2025

UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation

ICLR 2025poster

Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or gen…

Cited by 0SourcePDFScholar
2024

Bayesian Example Selection Improves In-Context Learning for Speech, Text and Visual Modalities

EMNLP 2024main

Large language models (LLMs) can adapt to new tasks through in-context learning (ICL) based on a few examples presented in dialogue history without any model parameter update. Despite such convenience, the performance of ICL heavily depends on the quality of the in-context examples presented, which…

2024

Exploiting A Quantum Multiple Kernel Learning Approach For Low-Resource Spoken Command Recognition

ICASSP 2024accepted

We propose a theoretical analysis of quantum projection learning (QPL) that employs multiple kernels, highlighting its advantages through representation error analysis. Building upon previous studies that utilized a single quantum kernel-based method, we further investigate a quantum projection fram…

Cited by 0SourceScholar
2024

FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model

EMNLP 2024industry

In this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction. Unlike other modalities such as vision or text, speech has unique temporal dependencies, making previous efficient inference works on other modalities not directly applicable. Further…

2024

From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment

EMNLP 2024main

Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual captions more descriptive, improving alignment with the visual context. However, while many studies focus on the benefi…

Cited by 3SourcePDFScholar
2024

GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators

ACL 2024long

Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. However, both translation tasks typically utilize beam search decoding and top-1 hypothesis se…

2024

Hot-Fixing Wake Word Recognition for End-to-End ASR Via Neural Model Reprogramming

ICASSP 2024accepted

This paper proposes two novel variants of neural reprogramming to enhance wake word recognition in streaming end-to-end ASR models without updating model weights. The first, "trigger-frame reprogramming", prepends the input speech feature sequence with the learned trigger-frames of the target wake w…

Cited by 0SourceScholar
2024

It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition

ICLR 2024poster

Recent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output. Specifically, an LLM is utilized to carry out a direct mapping from the N-best hypotheses list generated by…

Cited by 27SourcePDFScholar
2024

Large Language Models are Efficient Learners of Noise-Robust Speech Recognition

ICLR 2024spotlight

Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with "…

2024

Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue

ICASSP 2024accepted

Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking style, which are essential for achieving natural, human-like…

Cited by 0SourceScholar
2024

Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models

NeurIPS 2024poster

We propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) systems in diverse target domains, such as noise and accents. STAR is developed for prevalent speech foundation models based…

2024

Towards ASR Robust Spoken Language Understanding Through in-Context Learning with Word Confusion Networks

ICASSP 2024accepted

In the realm of spoken language understanding (SLU). numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with transcribed speech instead of conventional written text. In real-world scenarios, prior to input into an LLM. an automated…

Cited by 0SourceScholar
2023

A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command Recognition

ICASSP 2023accepted

We propose a quantum kernel learning (QKL) framework to address the inherent data sparsity issues often encountered in training large-scare acoustic models in low-resource scenarios. We project acoustic features based on classical-to-quantum feature encoding. Different from existing quantum convolut…

Cited by 0SourceScholar
2023

Certified Robustness of Quantum Classifiers Against Adversarial Examples Through Quantum Noise

ICASSP 2023accepted

Recently, quantum classifiers have been known to be vulnerable to adversarial attacks, where quantum classifiers are fooled by imperceptible noises to have misclassification. In this paper, we propose one first theoretical study that utilizing the added quantum random rotation noise can improve the…

Cited by 0SourceScholar
2023

From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition

ICASSP 2023accepted

In this work, we propose a new parameter-efficient learning framework based on neural model reprogramming for cross-lingual speech recognition, which can re-purpose well-trained English automatic speech recognition (ASR) models to recognize the other languages. We design different auxiliary neural a…

Cited by 0SourceScholar
2023

HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models

NeurIPS 2023poster

Advancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets. However, even state-of-the-art ASR systems experience performance degradation when confronted with adverse conditions, as a well-tr…

2023

Low-Resource Music Genre Classification with Cross-Modal Neural Model Reprogramming

ICASSP 2023accepted

Transfer learning (TL) approaches have shown promising results when handling tasks with limited training data. However, considerable memory and computational resources are often required for fine-tuning pre-trained neural networks with target domain data. In this work, we introduce a novel method fo…

Cited by 0SourceScholar
2023

Pessimistic Model Selection for Offline Deep Reinforcement Learning

UAI 2023poster

Deep Reinforcement Learning (DRL) has demonstrated great potentials in solving sequential decision making problems in many applications. Despite its promising performance, practical gaps exist when deploying DRL in real-world scenarios. One main barrier is the over-fitting issue that leads to poor g…

Cited by 5SourcePDFScholar
2023

Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition

EMNLP 2023short main

We introduce a new cross-modal fusion technique designed for generative error correction in automatic speech recognition (ASR). Our methodology leverages both acoustic information and external linguistic representations to generate accurate speech transcription contexts. This marks a step towards a…

Cited by 93SourcecodeScholar
2022

A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network Pruning

ICASSP 2022accepted

Audio-only based wake word spotting (WWS) is challenging under noisy conditions due to the environmental interference in signal transmission. In this paper, we investigate on designing a compact audio-visual WWS system by utilizing the visual information to alleviate the degradation. Specifically, i…

Cited by 0SourceScholar
2022

A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge Transfer

ICASSP 2022accepted

We propose a variational Bayesian (VB) approach to learning distributions of latent variables in deep neural network (DNN) models for cross-domain knowledge transfer, to address acoustic mismatches between training and testing conditions. Instead of carrying out point estimation in conventional maxi…

Cited by 0SourceScholar
2022

Mitigating Closed-Model Adversarial Examples with Bayesian Neural Modeling for Enhanced End-to-End Speech Recognition

ICASSP 2022accepted

In this work, we aim to enhance the system robustness of end-to-end automatic speech recognition (ASR) against adversarially-noisy speech examples. We focus on a rigorous and empirical "closed-model adversarial robustness" setting (e.g., on-device or cloud applications). The adversarial noise is onl…

Cited by 0SourceScholar
2022

Training a Resilient Q-network against Observational Interference

AAAI 2022technical

Deep reinforcement learning (DRL) has demonstrated impressive performance in various gaming simulators and real-world applications. In practice, however, a DRL agent may receive faulty observation by abrupt interferences such as black-out, frozen-screen, and adversarial perturbation. How to design…

2022

When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous Computing

ICASSP 2022accepted

The rapid development of quantum computing has demonstrated many unique characteristics of quantum advantages, such as richer feature representation and more secured protection on model parameters. This work proposes a vertical federated learning architecture based on variational quantum circuits to…

Cited by 0SourceScholar
2021

A Two-Stage Approach to Device-Robust Acoustic Scene Classification

ICASSP 2021accepted

To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks (CNNs) is proposed. Our two-stage system leverages on an ad-hoc score combination based on two C…

Cited by 0SourceScholar
2021

Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech Recognition

ICASSP 2021accepted

We propose a novel decentralized feature extraction approach in federated learning to address privacy-preservation issues for speech recognition. It is built upon a quantum convolutional neural network (QCNN) composed of a quantum circuit encoder for feature extraction, and a recurrent neural networ…

Cited by 0SourceScholar
2021

Voice2Series: Reprogramming Acoustic Models for Time Series Classification

ICML 2021spotlight

Learning to classify time series with limited data is a practical yet challenging problem. Current methods are primarily based on hand-designed feature extraction rules or domain-specific data augmentation. Motivated by the advances in deep speech processing models and the fact that voice data are u…

2020

Characterizing Speech Adversarial Examples Using Self-Attention U-Net Enhancement

ICASSP 2020accepted

Recent studies have highlighted adversarial examples as ubiquitous threats to the deep neural network (DNN) based speech recognition systems. In this work, we present a U-Net based attention model, UNet <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">At…

Cited by 0SourceScholar
2020

Enhanced Adversarial Strategically-Timed Attacks Against Deep Reinforcement Learning

ICASSP 2020accepted

Recent deep neural networks based techniques, especially those equipped with the ability of self-adaptation in the system level such as deep reinforcement learning (DRL), are shown to possess many advantages of optimizing robot learning systems (e.g., autonomous navigation and continuous robot arm c…

Cited by 0SourceScholar
2020

Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding

ICASSP 2020accepted

Performing driving behaviors based on causal reasoning is essential to ensure driving safety. In this work, we investigated how state-of-the-art 3D Convolutional Neural Networks (CNNs) perform on classifying driving behaviors based on causal reasoning. We proposed a perturbation-based visual explana…

Cited by 0SourceScholar
2020

Submodular Rank Aggregation on Score-Based Permutations for Distributed Automatic Speech Recognition

ICASSP 2020accepted

Distributed automatic speech recognition (ASR) requires to aggregate outputs of distributed deep neural network (DNN)-based models. This work studies the use of submodular functions to design a rank aggregation on score-based permutations, which can be used for distributed ASR systems in both superv…

Cited by 0SourceScholar
2020

Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train Network

ICASSP 2020accepted

We propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key idea is to cast the conventional deep neural network (DNN) based vector-to-vector regression formulation under a tensor…

Cited by 0SourceScholar
2020

Y-Net: Multi-Scale Feature Aggregation Network With Wavelet Structure Similarity Loss Function For Single Image Dehazing

ICASSP 2020accepted

Single image dehazing is the ill-posed two-dimensional signal reconstruction problem. Recently, deep convolutional neural networks (CNN) have been successfully used in many computer vision problems. In this paper, we propose a Y-net that is named for its structure. This network reconstructs clear im…

Cited by 0SourceScholar