← Search

Nima Mesgarani

31 accepted papers

2026

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

AAAI 2026technical

Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric optimization for speech generation components, but duration predict

Cited by 0SourcePDFScholar
2026

Decoding Inner Speech with an End-to-End Brain-to-Text Neural Interface

ICLR 2026poster

Speech brain–computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stag…

Cited by 0SourceScholar
2026

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

ICASSP 2026poster

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio…

Cited by 0SourcePDFScholar
2025

AAD-LLM: Neural Attention-Driven Auditory Scene Understanding

ACL 2025long

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific speakers while ignoring others in complex auditory scenes. Existi…

2025

Decoding the Unintelligible: Neural Speech Tracking in Low Signal-to-Noise Ratios

ICASSP 2025accepted

Understanding speech in noisy environments is challenging for both human listeners and speech technologies, with significant implications for hearing aid design and communication systems. Auditory attention decoding (AAD) aims to decode the attended talker from neural signals to enhance their speech…

Cited by 0SourceScholar
2025

Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation

ICASSP 2025accepted

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in computation and memory. Recent models incorporate new layers and modules along wit…

Cited by 0SourceScholar
2025

Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement

NeurIPS 2025poster

Understanding how the human brain progresses from processing simple linguistic inputs to performing high-level reasoning is a fundamental challenge in neuroscience. While modern large language models (LLMs) are increasingly used to model neural responses to language, their internal representations a…

Cited by 0SourceScholar
2025

Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence

ACL 2025finding

This study investigates the linguistic understanding of Large Language Models (LLMs) regarding signifier (form) and signified (meaning) by distinguishing two LLM assessment paradigms: psycholinguistic and neurolinguistic. Traditional psycholinguistic evaluations often reflect statistical rules that…

Cited by 0SourcePDFScholar
2025

Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations

EMNLP 2025

Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode shallow acoustic and phonetic features, the extent to which SLMs encode nuanced syntactic and conceptual features remains

Cited by 0SourcePDFScholar
2025

Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis

ICASSP 2025accepted

It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks. To reach this conclusion, we propose and evaluate three models for three tasks: Mamba-TasNe…

Cited by 0SourceScholar
2025

StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion

NAACL 2025long

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex pre-trained neural codec representations, and difficulties i…

2025

ZeroSep: Separate Anything in Audio with Zero Training

NeurIPS 2025poster

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive, task-specific labeled data and struggle to generalize to the…

Cited by 0SourceScholar
2024

Exploring Self-supervised Contrastive Learning of Spatial Sound Event Representation

ICASSP 2024accepted

In this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and spatial representations from unlabeled spatial audios, thereby enhancing both event classification and sound localizatio…

Cited by 0SourceScholar
2023

Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions

ICASSP 2023accepted

Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are usually word-level or sup-phoneme-level and jointly trained with phonemes, maki…

Cited by 0SourceScholar
2023

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

NeurIPS 2023poster

In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis. StyleTTS 2 differs from its predecessor by modeling styles as a latent random variable through dif…

2021

Rethinking The Separation Layers In Speech Separation Networks

ICASSP 2021accepted

Modules in all existing speech separation networks can be categorized into single-input-multi-output (SIMO) modules and single-input-single-output (SISO) modules. SIMO modules generate more outputs than input, and SISO modules keep the numbers of input and output the same. While the majority of sepa…

Cited by 0SourceScholar
2021

Understanding Adaptive, Multiscale Temporal Integration In Deep Speech Recognition Systems

NeurIPS 2021poster

Natural signals such as speech are hierarchically structured across many different timescales, spanning tens (e.g., phonemes) to hundreds (e.g., words) of milliseconds, each of which is highly variable and context-dependent. While deep neural networks (DNNs) excel at recognizing complex patterns fro…

2020

End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation

ICASSP 2020accepted

An important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones. The former requires the system to be invariant to different indexing of the microphones with the same locations, while the latter requi…

Cited by 0SourceScholar
2018

Lip2Audspec: Speech Reconstruction from Silent Lip Movements Video

ICASSP 2018accepted

In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generation method resulting in a more natural sounding reconstructed speech. Our propos…

Cited by 0SourceScholar
2017

Deep clustering and conventional networks for music separation: Stronger together

ICASSP 2017accepted

Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging si…

Cited by 0SourceScholar
2017

NAPLib: An open source toolbox for real-time and offline Neural Acoustic Processing

ICASSP 2017accepted

In this paper, we introduce the Neural Acoustic Processing Library (NAPLib), a toolbox containing novel processing methods for real-time and offline analysis of neural activity in response to speech. Our method divides the speech signal and resultant neural activity into segmental units (e.g., phone…

Cited by 0SourceScholar
2017

Understanding the Representation and Computation of Multilayer Perceptrons: A Case Study in Speech Recognition

ICML 2017accepted

Despite the recent success of deep learning, the nature of the transformations they apply to the input features remains poorly understood. This study provides an empirical framework to study the encoding properties of node activations in various layers of the network, and to construct the exact func…

Cited by 40SourcePDFScholar
2016

Synaptic depression in deep neural networks for speech processing

ICASSP 2016accepted

A characteristic property of biological neurons is their ability to dynamically change the synaptic efficacy in response to variable input conditions. This mechanism, known as synaptic depression, significantly contributes to the formation of normalized representation of speech features. Synaptic de…

Cited by 0SourceScholar