← Search

Karren D Yang

5 accepted papers

2026

AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

CVPR 2026

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning, tracking who speaks, maintaining roles, and grounding events across time. These scenarios are central to multim

Cited by 0SourceScholar
2024

Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language Models

ICASSP 2024accepted

While Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be fine-tuned on data from these domains. However, target-domain data usually are not readily available in many scenarios. In this paper, we…

Cited by 0SourceScholar
2024

Probabilistic Speech-Driven 3D Facial Motion Synthesis: New Benchmarks Methods and Applications

CVPR 2024poster

We consider the task of animating 3D facial geometry from speech signal. Existing works are primarily deterministic focusing on learning a one-to-one mapping from speech signal to 3D face meshes on small datasets with limited speakers. While these models can achieve high-quality lip articulation for…

Cited by 14SourcePDFScholar
2023

Text is all You Need: Personalizing ASR Models Using Controllable Speech Synthesis

ICASSP 2023accepted

Adapting generic speech recognition models to specific individuals is a challenging problem due to the scarcity of personalized data. Recent works have proposed boosting the amount of training data using personalized text-to-speech synthesis. Here, we ask two fundamental questions about this strateg…

Cited by 0SourceScholar
2019

Scalable Unbalanced Optimal Transport using Generative Adversarial Networks

ICLR 2019poster

Generative adversarial networks (GANs) are an expressive class of neural generative models with tremendous success in modeling high-dimensional continuous measures. In this paper, we present a scalable method for unbalanced optimal transport (OT) based on the generative-adversarial framework. We for…