← Search

Mark Hasegawa-Johnson

59 accepted papers

2026

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding

ICML 2026spotlight

While on-policy distillation offers dense supervision for training small reasoning models, its optimization dynamics in the multimodal domain remain under-explored. In this work, we challenge the standard monolithic view of Vision-Language Model (VLM) distillation by mathematically decomposing the l…

Cited by 0SourceScholar
2026

IN-SYNC: ADAPTATION OF SPEECH AWARE LARGE LANGUAGE MODELS FOR ASR WITH WORD LEVEL TIMESTAMP PREDICTIONS

ICASSP 2026oral

Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and…

Cited by 0SourcePDFScholar
2026

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signal--rewarding the confidence growth in the ground-truth answer--effectively improves language reasoning training by provi

Cited by 0SourcecodeScholar
2025

Cohort-Sensitive Labeling: An Effective Approach for Enhancing ASR Performance

ICASSP 2025accepted

This paper proposes a cohort-sensitive labeling (CSL) for automatic speech recognition (ASR). CSL is a method that distinguishes data labels based on cohorts, allowing models to learn cohort-specific information. For evaluation, we applied CSL using gender information in the training data of LibriSp…

Cited by 0SourceScholar
2025

Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech Recognition

ICASSP 2025accepted

Automatic Speech Recognition (ASR) holds immense potential to provide an effective interface for assistive technologies, but its performance remains unsatisfactory for people with speech impairments such as dysarthria. Existing ASR systems struggle to accurately recognize dysarthric speech due to th…

Cited by 0SourceScholar
2025

Improved Recognition of the Speech of People with Parkinson's Who Stutter

ICASSP 2025accepted

Stuttering is a speech disorder often associated with neurological conditions, including Parkinson’s disease (PD). Despite advancements in modern automatic speech recognition (ASR) technologies, today’s systems still face challenges in accurately recognizing dysarthric speech, particularly when stut…

Cited by 0SourceScholar
2025

LIMMITS'25: Multilingual Streaming TTS With Neural Codecs for Indian Languages

ICASSP 2025accepted

This work provides a summary of the Multilingual streaming TTS with neural codecs for Indian languages challenge (LIMMITS’25), organized as part of the ICASSP 2025 signal processing grand challenge. Towards this, 278 hours of TTS data in 4 Indian languages - Gujarati, Indian English, Bhojpuri, and K…

Cited by 0SourceScholar
2025

Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition

ICASSP 2025accepted

In this paper, we present a speaker-independent dysarthric speech recognition system, with a focus on evaluating the recently released Speech Accessibility Project (SAP-1005) dataset, which includes speech data from individuals with Parkinson’s disease (PD). Despite the growing body of research in d…

Cited by 0SourceScholar
2025

Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More

ICASSP 2025accepted

With the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 A…

Cited by 0SourceScholar
2024

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

ICASSP 2024accepted

In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapp…

Cited by 0SourceScholar
2024

Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation Pipeline

COLING 2024main

The growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-…

2024

TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

ACL 2024findings

Reinforcement Learning from Human Feedback (RLHF) leverages human preference data to train language models to align more closely with human essence. These human preference data, however, are labeled at the sequence level, creating a mismatch between sequence-level preference labels and tokens, which…

2024

Unsupervised Speech Recognition with N-skipgram and Positional Unigram Matching

ICASSP 2024accepted

Training unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the power of lower-order N-skipgrams…

Cited by 0SourceScholar
2023

Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction

ICASSP 2023accepted

Audiovisual target speaker extraction is the task of separating, from an audio mixture, the speaker whose face is visible in an accompanying video. Published approaches typically upsample the video or downsample the audio, then fuse the two streams using concatenation, multiplication, or cross-modal…

Cited by 0SourceScholar
2023

INTapt: Information-Theoretic Adversarial Prompt Tuning for Enhanced Non-Native Speech Recognition

ACL 2023findings

Automatic Speech Recognition (ASR) systems have attained unprecedented performance with large speech models pre-trained based on self-supervised speech representation learning. However, these pre-trained speech models suffer from representational bias as they tend to better represent those prominent…

Cited by 1SourcePDFScholar
2023

Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-Speech

ICASSP 2023accepted

The Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi,…

Cited by 0SourceScholar
2023

Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition

ACL 2023findings

Existing supervised sign language recognition systems rely on an abundance of well-annotated data. Instead, an unsupervised speech-to-sign language recognition (SSR-U) system learns to translate between spoken and sign languages by observing only non-parallel speech and sign-language corpora. We pro…

2022

ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers

ICML 2022spotlight

Self-supervised learning in speech involves training a speech representation network on a large-scale unannotated speech corpus, and then applying the learned representations to downstream tasks. Since the majority of the downstream tasks of SSL learning in speech largely focus on the content inform…

2022

Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough Audio

ICASSP 2022accepted

The distinct cough sounds produced by a variety of respiratory diseases suggest the potential for the development of a new class of audio bio-markers for the detection of COVID-19. Accurate audio biomarker-based COVID-19 tests would be inexpensive, readily scalable, and non-invasive. Audio biomarker…

Cited by 3SourceScholar
2022

Equivariance Discovery by Learned Parameter-Sharing

AISTATS 2022poster

Designing equivariance as an inductive bias into deep-nets has been a prominent approach to build effective models, e.g., a convolutional neural network incorporates translation equivariance. However, incorporating these inductive biases requires knowledge about the equivariance properties of the da…

2022

Fast and Efficient MMD-Based Fair PCA via Optimization over Stiefel Manifold

AAAI 2022technical

This paper defines fair principal component analysis (PCA) as minimizing the maximum mean discrepancy (MMD) between the dimensionality-reduced conditional distributions of different protected classes. The incorporation of MMD naturally leads to an exact and tractable mathematical formulation of fair…

2022

Forget-free Continual Learning with Winning Subnetworks

ICML 2022spotlight

Inspired by Lottery Ticket Hypothesis that competitive subnetworks exist within a dense network, we propose a continual learning method referred to as Winning SubNetworks (WSN), which sequentially learns and selects an optimal subnetwork for each task. Specifically, WSN jointly learns the model weig…

2022

SMSMix: Sense-Maintained Sentence Mixup for Word Sense Disambiguation

EMNLP 2022finding

Word Sense Disambiguation (WSD) is an NLP task aimed at determining the correct sense of a word in a sentence from discrete sense choices. Although current systems have attained unprecedented performances for such tasks, the nonuniform distribution of word senses during training generally results in…

2022

Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech Recognition

ACL 2022long

Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to under-resourced speech technology. In this paper, we bridge the gap between the linguistic a…

Cited by 4SourcePDFScholar
2022

SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder Bottlenecks

ICASSP 2022accepted

SpeechSplit can perform aspect-specific voice conversion by disentangling speech into content, rhythm, pitch, and timbre using multiple autoencoders in an unsupervised manner. However, SpeechSplit requires careful tuning of the autoencoder bottlenecks, which can be time-consuming and less robust. Th…

Cited by 0SourceScholar
2022

Syn2Vec: Synset Colexification Graphs for Lexical Semantic Similarity

NAACL 2022long

In this paper we focus on patterns of colexification (co-expressions of form-meaning mapping in the lexicon) as an aspect of lexical-semantic organization, and use them to build large scale synset graphs across BabelNet’s typologically diverse set of 499 world languages. We introduce and compare sev…

2021

Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval

ICASSP 2021accepted

Multimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conduct…

Cited by 0SourceScholar
2021

Continuous Cnn For Nonuniform Time Series

ICASSP 2021accepted

CNN for time series data implicitly assumes that the data are uniformly sampled, whereas many event-based and multi-modal data are nonuniform or have heterogeneous sampling rates. Directly applying regular CNN to nonuniform time series is ungrounded, because it is unable to recognize and extract com…

Cited by 0SourceScholar
2021

Global Prosody Style Transfer Without Text Transcriptions

ICML 2021oral

Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody information. Two major components of prosody are pitch and rhythm. Disentangling the prosody information, particularly t…

Cited by 42SourcePDFScholar
2021

How Phonotactics Affect Multilingual and Zero-Shot ASR Performance

ICASSP 2021accepted

The idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcription…

Cited by 0SourceScholar
2021

Interpretable Visual Reasoning via Induced Symbolic Space

ICCV 2021poster

We study the problem of concept induction in visual reasoning, i.e., identifying concepts and their hierarchical relationships from question-answer pairs associated with images; and achieve an interpretable model via working on the induced symbolic concept space. To this end, we first design a new f…

Cited by 22PDFcodeScholar
2021

Multi-Decoder Dprnn: Source Separation for Variable Number of Speakers

ICASSP 2021accepted

We propose an end-to-end trainable approach to single-channel speech separation with unknown number of speakers. Our approach extends the MulCat source separation backbone with additional output heads: a count-head to infer the number of speakers, and decoder-heads for reconstructing the original si…

Cited by 0SourceScholar
2021

Show and Speak: Directly Synthesize Spoken Description of Images

ICASSP 2021accepted

This paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the need for any text or phonemes. The basic structure of SAS is an encoder-decoder architecture that takes an image as inp…

Cited by 0SourceScholar
2021

Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded Vocabulary

ICASSP 2021accepted

Dysarthria is a condition where people experience a reduction in speech intelligibility due to a neuromotor disorder. Previous works in dysarthric speech recognition have focused on accurate recognition of words encountered in training data. Due to the rarity of dysarthria in the general population,…

Cited by 0SourceScholar
2021

Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question Answering

NAACL 2021long

Although Question-Answering has long been of research interest, its accessibility to users through a speech interface and its support to multiple languages have not been addressed in prior studies. Towards these ends, we present a new task and a synthetically-generated dataset to do Fact-based Visua…

2020

F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder

ICASSP 2020accepted

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networks (GANs) and variational autoencoders (VAEs) have been proposed. Recently, AutoVC, a conditional autoencoders (CAEs) ba…

Cited by 0SourceScholar
2020

Training Spoken Language Understanding Systems with Non-Parallel Speech and Text

ICASSP 2020accepted

End-to-end spoken language understanding (SLU) systems are typically trained on large amounts of data. In many practical scenarios, the amount of labeled speech is often limited as opposed to text. In this study, we investigate the use of non-parallel speech and text to improve the performance of di…

Cited by 0SourceScholar
2020

Unsupervised Speech Decomposition via Triple Information Bottleneck

ICML 2020poster

Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to sp…

2019

AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss

ICML 2019oral

Despite the progress in voice conversion, many-to-many voice conversion trained on non-parallel data, as well as zero-shot voice conversion, remains under-explored. Deep style transfer algorithms, generative adversarial networks (GAN) in particular, are being applied as new solutions in this field.…

2019

Dimensional Analysis of Laughter in Female Conversational Speech

ICASSP 2019accepted

How do people hear laughter in expressive, unprompted speech? What is the range of expressivity and function of laughter in this speech, and how can laughter inform the recognition of higher-level expressive dimensions in a corpus? This paper presents a scalable method for collecting natural human d…

Cited by 0SourceScholar
2019

Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast News

ICASSP 2019accepted

In this work, we investigate pre-training of neural network based speaker embeddings for low-latency speaker change detection. Our proposed system takes two speech segments, generates embeddings using shared Siamese layers and then classifies the concatenated embeddings depending on whether they are…

Cited by 0SourceScholar
2019

When CTC Training Meets Acoustic Landmarks

ICASSP 2019accepted

Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in resource-constrained scenarios. In this paper, the convergence properties of C…

Cited by 0SourceScholar
2018

Bayesian Models for Unit Discovery on a Very Low Resource Language

ICASSP 2018accepted

Developing speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to u…

Cited by 0SourceScholar
2018

Deep Learning Based Speech Beamforming

ICASSP 2018accepted

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would otherwise be too complicated. On the other hand, deep learn…

Cited by 0SourceScholar
2018

Image Restoration with Deep Generative Models

ICASSP 2018accepted

Many image restoration problems are ill-posed in nature, hence, beyond the input image, most existing methods rely on a carefully engineered image prior, which enforces some local image consistency in the recovered image. How tightly the prior assumptions are fulfilled has a big impact on the result…

Cited by 0SourceScholar
2018

Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition

ICASSP 2018accepted

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several ac…

Cited by 0SourceScholar
2018

Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

ICASSP 2018accepted

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translate…

Cited by 0SourceScholar
2018

Recognizing Zero-Resourced Languages Based on Mismatched Machine Transcriptions

ICASSP 2018accepted

Mismatched crowdsourcing based probabilistic human transcription has been proposed recently for training and adapting acoustic models for zero-resourced languages where we do not have any native transcriptions. This paper describes a machine transcription based phone recognition system for recognizi…

Cited by 0SourceScholar
2018

Time-Frequency Networks for Audio Super-Resolution

ICASSP 2018accepted

Audio super-resolution (a.k.a. bandwidth extension) is the challenging task of increasing the temporal resolution of audio signals. Recent deep networks approaches achieved promising results by modeling the task as a regression problem in either time or frequency domain. In this paper, we introduced…

Cited by 0SourceScholar
2017

Discovering dimensions of perceived vocal expression in semi-structured, unscripted oral history accounts

ICASSP 2017accepted

What do people hear in expressive, unprompted speech? And how can their descriptions be transformed into a representative set of dimensions of vocal expression? This paper presents a methodology for collecting user description of vocal expression, transforms the user descriptions into a set of measu…

Cited by 0SourceScholar
2017

Semantic Image Inpainting With Deep Generative Models

CVPR 2017poster

Semantic image inpainting is a challenging task where large missing regions have to be filled based on the available visual data. Existing methods which extract information from only a single image generally produce unsatisfactory results due to the lack of high level context. In this paper, we pro…

Cited by 1485PDFScholar
2016

Adapting ASR for under-resourced languages using mismatched transcriptions

ICASSP 2016accepted

Mismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages,…

Cited by 0SourceScholar
2016

Landmark of Mandarin nasal codas and its application in pronunciation error detection

ICASSP 2016accepted

L2 learners of Mandarin have difficulty learning native-like pronunciation of nasal codas. In order to help them learn native-like pronunciation, we propose to develop targeted classifiers for automatic pronunciation error detection. In this paper, perceptual experiments with modified speech are des…

Cited by 0SourceScholar
2015

Multichannel transient acoustic signal classification using task-driven dictionary with joint sparsity and beamforming

ICASSP 2015accepted

We are interested in a multichannel transient acoustic signal classification task which suffers from additive/convolutionary noise corruption. To address this problem, we propose a double-scheme classifier that takes the advantage of multichannel data to improve noise robustness. Both schemes adopt…

Cited by 0SourceScholar