← Search

William Chen

19 accepted papers

2026

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

ICML 2026poster

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound effects. Compared to traditional audio processing tasks, audio sto…

Cited by 0SourcecodeScholar
2026

Learning Affordances at Inference-Time for Vision-Language-Action Models

ICRA 2026poster

Solving complex real-world control tasks often takes multiple tries: if we fail at first, we reflect on what went wrong, and change our strategy accordingly to avoid making the same mistake. In robotics, Vision-Language-Action models (VLAs) offer a promising path towards solving complex control task…

2026

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

RSS 2026poster

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is …

Cited by 26SourceScholar
2026

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

RSS 2026poster

Pretrained vision-language models (VLMs) can make semantic and visual inferences across diverse settings, providing valuable common-sense priors for robotic control. However, effectively grounding this knowledge in robot behaviors remains an open challenge. Prior methods often employ a hierarchical …

Cited by 0SourceScholar
2025

Bridging Speech and Text Foundation Models with ReShape Attention

ICASSP 2025accepted

This paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape A…

Cited by 0SourceScholar
2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

ICLR 2025poster

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication…

2025

ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems

NAACL 2025system demonstrations

Advancements in audio foundation models (FMs) have fueled interest in end-to-end (E2E) spoken dialogue systems, but different web interfaces for each system makes it challenging to compare and contrast them effectively. Motivated by this, we introduce an open-source, user-friendly toolkit designed t…

2025

ESPnet-SpeechLM: An Open Speech Language Model Toolkit

NAACL 2025system demonstrations

We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing them as universal sequential modeling problems, encompassing a cohesive workflo…

2025

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

ICML 2025poster

Neural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In this work, we introduce OWLS, an open-access, reproducible su…

Cited by 1SourcePDFScholar
2025

Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility

ICLR 2025poster

With the rise of large language models (LLMs), increasing research has recognized their risk of leaking personally identifiable information (PII) under malicious attacks. Although efforts have been made to protect PII in LLMs, existing methods struggle to balance privacy protection with maintaining…

Cited by 3SourcePDFScholar
2025

Training Strategies for Efficient Embodied Reasoning

CoRL 2025oral

Robot chain-of-thought reasoning (CoT) -- wherein a model predicts helpful intermediate representations before choosing actions -- provides an effective method for improving the generalization and performance of robot policies, especially vision-language-action models (VLAs). While such approaches h…

Cited by 0SourceScholar
2024

AugSumm: Towards Generalizable Speech Summarization Using Synthetic Labels from Large Language Models

ICASSP 2024accepted

Abstractive speech summarization (SSUM) aims to generate humanlike summaries from speech. Given variations in information captured and phrasing, recordings can be summarized in multiple ways. Therefore, it is more reasonable to consider a probabilistic distribution of all potential summaries rather…

Cited by 0SourceScholar
2024

Evaluating Self-Supervised Speech Representations for Indigenous American Languages

COLING 2024main

The application of self-supervision to speech representation learning has garnered significant interest in recent years, due to its scalability to large amounts of unlabeled data. However, much progress, both in terms of pre-training and downstream evaluation, has remained concentrated in monolingua…

Cited by 4SourcePDFScholar
2024

Indoor and Outdoor 3D Scene Graph Generation Via Language-Enabled Spatial Ontologies

RA-L 2024

This paper proposes an approach to build 3D scene graphs in arbitrary indoor and outdoor environments. Such extension is challenging; the hierarchy of concepts that describe an outdoor environment is more complex than for indoors, and manually defining such hierarchy is time-consuming and does not s

Cited by 46SourceScholar
2024

On the Evaluation of Speech Foundation Models for Spoken Language Understanding

ACL 2024findings

The Spoken Language Understanding Evaluation (SLUE) suite of benchmark tasks was recently introduced to address the need for openresources and benchmarking of complex spoken language understanding (SLU) tasks, including both classification and sequence generation tasks, on natural speech. The benchm…

Cited by 6SourcePDFScholar
2024

Robotic Control via Embodied Chain-of-Thought Reasoning

CoRL 2024poster

A key limitation of learned robot control policies is their inability to generalize outside their training data. Recent works on vision-language-action models (VLAs) have shown that the use of large, internet pre-trained vision-language models as the backbone of learned robot policies can substanti…

Cited by 58SourceScholar
2024

Towards Robust Speech Representation Learning for Thousands of Languages

EMNLP 2024main

Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world’s 7000+ languages. We propose XEUS, a Cross-lingual Encoder for Universal Speech, trained on over 1 million hours of d…

2024

Train Long and Test Long: Leveraging Full Document Contexts in Speech Processing

ICASSP 2024accepted

The quadratic memory complexity of self-attention has generally restricted Transformer-based models to utterance-based speech processing, preventing models from leveraging long-form contexts. A common solution has been to formulate long-form speech processing into a streaming problem, only using lim…

Cited by 0SourceScholar
2023

Improving Massively Multilingual ASR with Auxiliary CTC Objectives

ICASSP 2023accepted

Multilingual Automatic Speech Recognition (ASR) models have extended the usability of speech technologies to a wide variety of languages. With how many languages these models have to handle, however, a key to understanding their imbalanced performance across different languages is to examine if the…

Cited by 0SourceScholar