← Search

Aren Jansen

20 accepted papers

2025

Long-Form Speech Generation with Spoken Language Models

ICML 2025oral

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past tens of seconds, due to high temporal resolution of speech tok…

2025

MELODI: Exploring Memory Compression for Long Contexts

ICLR 2025poster

We present MELODI, a novel memory architecture designed to efficiently process long documents using short context windows. The key principle behind MELODI is to represent short-term and long-term memory as a hierarchical compression scheme across both transformer layers and context windows. Specific…

Cited by 3SourcePDFScholar
2024

A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation

NeurIPS 2024poster

Training diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training a separate model for each task which is expensive. Here, we…

2024

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

AAAI 2024technical

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the exploration of video-acoustic signatures has been confined to sp…

Cited by 12SourcePDFScholar
2023

Dataset Balancing Can Hurt Model Performance

ICASSP 2023accepted

Machine learning from training data with a skewed distribution of examples per class can lead to models that favor performance on common classes at the expense of performance on rare ones. AudioSet has a very wide range of priors over its 527 sound event classes. Classification performance on AudioS…

Cited by 0SourceScholar
2022

Universal Paralinguistic Speech Representations Using self-Supervised Conformers

ICASSP 2022accepted

Many speech applications require understanding aspects beyond the words being spoken, such as recognizing emotion, detecting whether the speaker is wearing a mask, or distinguishing real from synthetic speech. In this work, we introduce a new state-of-the-art paralinguistic representation derived fr…

Cited by 0SourceScholar
2021

Attention Bottlenecks for Multimodal Fusion

NeurIPS 2021poster

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for unimodal benchmarks. A common approach for building multimodal m…

2021

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

ICLR 2021poster

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work, we present AudioScope, a novel audio-visual sound separation framework that can…

Cited by 86SourcePDFScholar
2021

The Benefit of Temporally-Strong Labels in Audio Event Classification

ICASSP 2021accepted

To reveal the importance of temporal precision in ground truth audio event labels, we collected precise (∼0.1 sec resolution) "strong" labels for a portion of the AudioSet dataset. We devised a temporally-strong evaluation set (including explicit negatives of varying difficulty) and a small strong-l…

Cited by 160SourceScholar
2020

Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision

ICASSP 2020accepted

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on multimodal unsupervised learning (as infants) and active learnin…

Cited by 0SourceScholar
2020

Improving Universal Sound Separation Using Sound Classification

ICASSP 2020accepted

Deep learning approaches have recently achieved impressive performance on both audio source separation and sound classification. Most audio source separation approaches focus only on separating sources belonging to a restricted domain of source classes, such as speech and music. However, recent work…

Cited by 0SourceScholar
2020

Large-Scale Weakly-Supervised Content Embeddings for Music Recommendation and Tagging

ICASSP 2020accepted

We explore content-based representation learning strategies tailored for large-scale, uncurated music collections that afford only weak supervision through unstructured natural language metadata and co-listen statistics. At the core is a hybrid training scheme that uses classification and metric lea…

Cited by 0SourceScholar
2018

Unsupervised Learning of Semantic Audio Representations

ICASSP 2018accepted

Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We consider several class-agnostic semantic constraints that apply to unlabeled nonspeech audio: (i) noise and translations in…

Cited by 0SourceScholar
2017

Audio Set: An ontology and human-labeled dataset for audio events

ICASSP 2017accepted

Audio event recognition, the human-like ability to identify and relate sounds from audio, is a nascent problem in machine perception. Comparable problems such as object detection in images have reaped enormous benefits from comprehensive datasets - principally ImageNet. This paper describes the crea…

Cited by 0SourceScholar
2017

CNN architectures for large-scale audio classification

ICASSP 2017accepted

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Ne…

Cited by 3037SourceScholar
2017

Large-scale audio event discovery in one million YouTube videos

ICASSP 2017accepted

Internet videos provide a virtually boundless source of audio with a conspicuous lack of localized annotations, presenting an ideal setting for unsupervised methods. With this motivation, we perform an unprecedented exploration into the large-scale discovery of recurring audio events in a diverse co…

Cited by 0SourceScholar
2016

Context-dependent point process models for keyword search and detection-based ASR

ICASSP 2016accepted

The point process model (PPM) for keyword search (KWS) is a whole-word parametric approach that characterizes each query type by the timing of phonetic events observed during its production. In this paper, we first extend the PPM modeling framework to operate on context-dependent phonetic event patt…

Cited by 0SourceScholar
2015

Content-based recommender systems for spoken documents

ICASSP 2015accepted

Content-based recommender systems use preference ratings and features that characterize media to model users' interests or information needs for making future recommendations. While previously developed in the music and text domains, we present an initial exploration of content-based recommendation…

Cited by 16SourceScholar
2015

Unsupervised neural network based feature extraction using weak top-down constraints

ICASSP 2015accepted

Deep neural networks (DNNs) have become a standard component in supervised ASR, used in both data-driven feature extraction and acoustic modelling. Supervision is typically obtained from a forced alignment that provides phone class targets, requiring transcriptions and pronunciations. We propose a n…

Cited by 0SourceScholar