← Search

Nicolas Obin

6 accepted papers

2026

GELINA: UNIFIED SPEECH AND GESTURE SYNTHESIS VIA INTERLEAVED TOKEN PREDICTION

ICASSP 2026oral

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-spee…

Cited by 0SourcePDFScholar
2024

Auditory Cortex-Inspired Spectral Attention Modulation for Binaural Sound Localization in HRTF Mismatch

ICASSP 2024accepted

In applications like noise cancellation and virtual reality, precise sound source localization is crucial. Existing data-driven binaural systems offer high performance in adverse conditions such as noise and reverberation but face limitations with real-time operation and performance degradation in H…

Cited by 0SourceScholar
2024

BWSNET: Automatic Perceptual Assessment of Audio Signals

ICASSP 2024accepted

This paper introduces BWSNet, a model that can be trained from raw human judgements obtained through a Best-Worst scaling (BWS) experiment. It maps sound samples into an embedded space that represents the perception of a studied attribute. To this end, we propose a set of cost functions and constrai…

Cited by 0SourceScholar
2016

A source/filter model with adaptive constraints for NMF-based speech separation

ICASSP 2016accepted

This paper introduces a constrained source/filter model for semi-supervised speech separation based on non-negative matrix factorization (NMF). The objective is to inform NMF with prior knowledge about speech, providing a physically meaningful speech separation. To do so, a source/filter model (indi…

Cited by 0SourceScholar
2015

The role of glottal source parameters for high-quality transformation of perceptual age

ICASSP 2015accepted

The intuitive control of voice transformation (e.g., age/sex, emotions) is useful to extend the expressive repertoire of a voice. This paper explores the role of glottal source parameters for the control of voice transformation. First, the SVLN speech synthesizer (Separation of the Vocal-tract with…

Cited by 0SourceScholar