← Search

Yuki Mitsufuji

93 accepted papers

2026

3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation

ICLR 2026poster

We present 3DScenePrompt, a framework for camera-controllable video generation that maintains scene consistency when extending arbitrary-length input videos along user-specified trajectories. Unlike existing video generative methods limited to conditioning on a single image or just a few frames, we…

Cited by 0SourcecodeScholar
2026

Automatic Music Mixing using a Generative Model of Effect Embeddings

ICASSP 2026oral

Music mixing involves combining individual tracks into a cohesive mixture, a task characterized by subjectivity where multiple valid solutions exist for the same input. Existing automatic mixing systems treat this task as a deterministic regression problem, thus ignoring this multiplicity of solutio…

Cited by 0SourcePDFScholar
2026

CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow-Map Models

ICLR 2026poster

Flow map models such as Consistency Models (CM) and Mean Flow (MF) enable few-step generation by learning the long jump of the ODE solution of diffusion models, yet training remains unstable, sensitive to hyperparameters, and costly. Initializing from a pre-trained diffusion model helps, but still r…

Cited by 0SourcecodeScholar
2026

Concept-TRAK: Understanding how diffusion models learn concepts through concept attribution

ICLR 2026poster

While diffusion models excel at image generation, their growing adoption raises critical concerns about copyright issues and model transparency. Existing attribution methods identify training examples influencing an entire image, but fall short in isolating contributions to specific elements, such a…

Cited by 0SourcecodeScholar
2026

Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion

ICML 2026poster

Masked diffusion models have shown promising performance in generating high-quality samples in a wide range of domains, but accelerating their sampling process remains relatively underexplored. To investigate efficient samplers for masked diffusion, this paper theoretically analyzes the MaskGIT samp…

Cited by 0SourceScholar
2026

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

CVPR 2026

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling challenge in multimodal-to-audio generation, examining whether models trained on sho

Cited by 0SourcecodeScholar
2026

FOLEYBENCH: A BENCHMARK FOR VIDEO-TO-AUDIO MODELS

ICASSP 2026oral

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires generating audio that is both semantically aligned with visible event…

Cited by 0SourcePDFScholar
2026

G2D2: Gradient-Guided Discrete Diffusion for Inverse Problem Solving

ICML 2026poster

Recent literature has effectively leveraged diffusion models trained on continuous variables as priors for solving inverse problems. Notably, discrete diffusion models with discrete latent codes have shown strong performance, particularly in modalities suited for discrete compressed representations,…

Cited by 0SourceScholar
2026

GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning

ICML 2026poster

Training-data attribution for vision generative models aims to identify which training data influenced a given output. While most methods score individual examples, practitioners often need group-level answers (e.g., artistic styles or object classes). Group-wise attribution is counterfactual: how w…

Cited by 0SourceScholar
2026

Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment

ICLR 2026poster

Slot Attention (SA) with pretrained diffusion models has recently shown promise for object-centric learning (OCL), but suffers from slot entanglement and weak alignment between object slots and image content. We propose Contrastive Object-centric Diffusion Alignment (CODA), a simple extension that (…

Cited by 0SourcecodeScholar
2026

Improving Classifier-Free Guidance in Masked Diffusion: Low-Dim Theoretical Insights with High-Dim Impact

ICLR 2026poster

Classifier-Free Guidance (CFG) is a widely used technique for conditional generation and improving sample quality in continuous diffusion models, and its extensions to discrete diffusion has recently started to be investigated. In order to improve the algorithms in a principled way, this paper start…

Cited by 0SourceScholar
2026

LEVERAGING WHISPER EMBEDDINGS FOR AUDIO-BASED LYRICS MATCHING

ICASSP 2026poster

Audio-based lyrics matching can be an appealing alternative to other content-based retrieval approaches, but existing methods often suffer from limited reproducibility and inconsistent baselines. In this work, we introduce WEALY, a fully reproducible pipeline that leverages Whisper decoder embedding…

Cited by 0SourcePDFScholar
2026

LLM2Fx-Tools: Tool Calling for Music Post-Production

ICLR 2026poster

This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses a large language model (LLM) to understand audio inputs, select audio effects types, determine their order, and estimat…

Cited by 0SourcecodeScholar
2026

Learning Compact 3D Representations from Feed-Forward Novel View Synthesis

CVPR 2026

Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3D Gaussian Splatting for reconstruction, followed by a 2D-to-3D feature lifting stage for scene understanding. However,

Cited by 0SourcecodeScholar
2026

Learning to Route Languages for Multilingual Preference Optimization

ICML 2026poster

Large language models (LLMs) are trained on heterogeneous multilingual corpora, yet existing preference optimization methods often implicitly restrict each training question to a single response language or rely on a fixed dominant language for supervision. We propose language-routed preference opti…

Cited by 0SourceScholar
2026

SAVGBENCH: BENCHMARKING SPATIALLY ALIGNED AUDIO-VIDEO GENERATION

ICASSP 2026poster

This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook the spatial alignment between audio and visuals, which is e…

Cited by 0SourcePDFScholar
2026

SONA: Learning Conditional, Unconditional, and Matching-Aware Discriminator

ICLR 2026poster

Deep generative models have made significant advances in generating complex content, yet conditional generation remains a fundamental challenge. Existing conditional generative adversarial networks often struggle to balance the dual objectives of assessing authenticity and conditional alignment of i…

Cited by 0SourcecodeScholar
2026

Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

ICML 2026poster

Multimodal LLMs lack a systematic understanding of visual dynamics in complex human world activities, which requires the model to predict or simulate multiple levels of dynamic constituents, such as the general progression of actions and the associated changes of low-level details in the world. To a…

Cited by 0SourceScholar
2026

SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing

AAAI 2026technical

Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods ofte

Cited by 0SourcePDFScholar
2026

Summary of The Inaugural Music Source Restoration Challenge

ICASSP 2026poster

Music Source Restoration (MSR) aims to recover original, unprocessed instrument stems from professionally mixed and degraded audio, requiring the reversal of both production effects and real-world degradations. We present the inaugural MSR Challenge, which features objective evaluation on studio-pro…

Cited by 0SourcePDFScholar
2026

TOWARDS BLIND DATA CLEANING: A CASE STUDY IN MUSIC SOURCE SEPARATION

ICASSP 2026poster

The performance of deep learning models for music source separation heavily depends on training data quality. However, datasets are often corrupted by difficult-to-detect artifacts such as audio bleeding and label noise. Since the type and extent of contamination are typically unknown, cleaning meth…

Cited by 0SourcePDFScholar
2026

VIRTUE: Visual-Interactive Text-Image Universal Embedder

ICLR 2026poster

Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embedding models with instruction-following capabilities. However, existing embedding models lack visual-interactive capabil…

Cited by 0SourcecodeScholar
2026

Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

AAAI 2026technical

We introduce a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view video data for training. Traditional reconstruction methods struggle with

Cited by 0SourcePDFScholar
2025

30+ Years of Source Separation Research: Achievements and Future Challenges

ICASSP 2025accepted

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> anniversary, we review the major contribut…

Cited by 23SourceScholar
2025

CARE: Multilingual Human Preference Learning for Cultural Awareness

EMNLP 2025

Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied. In this paper, we systematically analyze how native human cultural preferences can be incorpora

2025

Classifier-Free Guidance Inside the Attraction Basin May Cause Memorization

CVPR 2025poster

Diffusion models are prone to exactly reproduce images from the training data. This exact reproduction of the training data is concerning as it can lead to copyright infringement and/or leakage of privacy-sensitive information. In this paper, we present a novel perspective on the memorization phenom…

2025

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

EMNLP 2025

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model’s ability to analyze and interpret various musical elements. These improvements primarily focused on integrating both music and text inputs. However, the potential

2025

Distillation of Discrete Diffusion through Dimensional Correlations

ICML 2025poster

Diffusion models have demonstrated exceptional performances in various fields of generative modeling, but suffer from slow sampling speed due to their iterative nature. While this issue is being addressed in continuous domains, discrete diffusion models face unique challenges, particularly in captur…

2025

Enhancing 3D Reconstruction for Dynamic Scenes

NeurIPS 2025poster

In this work, we address the task of 3D reconstruction in dynamic scenes, where object motions frequently degrade the quality of previous 3D pointmap regression methods, such as DUSt3R, that are originally designed for static 3D scene reconstruction. Although these methods provide an elegant and pow…

Cited by 0SourceScholar
2025

HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning

ICLR 2025poster

Controllable generation through Stable Diffusion (SD) fine-tuning aims to improve fidelity, safety, and alignment with human guidance. Existing reinforcement learning from human feedback methods usually rely on predefined heuristic reward functions or pretrained reward models built on large-scale da…

2025

Jump Your Steps: Optimizing Sampling Schedule of Discrete Diffusion Models

ICLR 2025poster

Diffusion models have seen notable success in continuous domains, leading to the development of discrete diffusion models (DDMs) for discrete variables. Despite recent advances, DDMs face the challenge of slow sampling speeds. While parallel sampling methods like $\tau$-leaping accelerate this proce…

Cited by 3SourcePDFScholar
2025

Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer

ICASSP 2025accepted

Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired mono…

Cited by 0SourceScholar
2025

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

CVPR 2025poster

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework (MMAudio). In contrast to single-modality training conditioned on (limited) video data only, MMAudio is jointly trained with larger-scale, readily…

2025

MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation

ICLR 2025poster

This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides single-modal models to cooperatively generate well-aligned samples across…

2025

Mining your own secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models

ICLR 2025poster

Personalized text-to-image diffusion models have grown popular for their ability to efficiently acquire a new concept from user-defined text descriptions and a few images. However, in the real world, a user may wish to personalize a model on multiple concepts but one at a time, with no access to the…

Cited by 1SourcePDFScholar
2025

SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

ICLR 2025poster

Sound content creation, essential for multimedia works such as video games and films, often involves extensive trial-and-error, enabling creators to semantically reflect their artistic ideas and inspirations, which evolve throughout the creation process, into the sound. Recent high-quality diffusion…

2025

Supervised Contrastive Learning from Weakly-Labeled Audio Segments for Musical Version Matching

ICML 2025poster

Detecting musical versions (different renditions of the same piece) is a challenging task with important applications. Because of the ground truth nature, existing approaches match musical versions at the track level (e.g., whole song). However, most applications require to match them at the segment…

Cited by 0SourcePDFScholar
2025

TITAN-Guide: Taming Inference-Time Alignment for Guided Text-to-Video Diffusion Models

ICCV 2025poster

In the recent development of conditional diffusion models still require heavy supervised fine-tuning for performing control on a category of tasks. Training-free conditioning via guidance with off-the-shelf models is a favorable alternative to avoid further fine-tuning on the base model. However, th…

2025

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

NeurIPS 2025poster

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos w…

Cited by 0SourceScholar
2025

Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models

ICCV 2025poster

Parameter-Efficient Fine-Tuning (PEFT) of text-to-image models has become an increasingly popular technique with many applications. Among the various PEFT methods, Low-Rank Adaptation (LoRA) and its variants have gained significant attention due to their effectiveness, enabling users to fine-tune mo…

2025

Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges

ICASSP 2025accepted

In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and A…

Cited by 1SourceScholar
2025

VCT: Training Consistency Models with Variational Noise Coupling

ICML 2025poster

Consistency Training (CT) has recently emerged as a strong alternative to diffusion models for image generation. However, non-distillation CT often suffers from high variance and instability, motivating ongoing research into its training dynamics. We propose Variational Consistency Training (VCT), a…

2025

Variable Bitrate Residual Vector Quantization for Audio Coding

ICASSP 2025accepted

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple…

Cited by 12SourceScholar
2025

VinaBench: Benchmark for Faithful and Consistent Visual Narratives

CVPR 2025poster

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge const…

Cited by 1SourcePDFScholar
2025

Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric

ICLR 2025spotlight

In typical multimodal contrastive learning, such as CLIP, encoders produce one point in the latent representation space for each input. However, one-point representation has difficulty in capturing the relationship and the similarity structure of a huge amount of instances in the real world. For ric…

Cited by 0SourcePDFScholar
2024

BIGVSAN: Enhancing Gan-Based Neural Vocoders with Slicing Adversarial Network

ICASSP 2024accepted

Generative adversarial network (GAN)-based vocoders have been intensively studied because they can synthesize high-fidelity audio waveforms faster than real-time. However, it has been reported that most GANs fail to obtain the optimal projection for discriminating between real and fake data in the f…

Cited by 0SourceScholar
2024

Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

ICLR 2024poster

Consistency Models (CM) (Song et al., 2023) accelerate score-based diffusion model sampling at the cost of sample quality but lack a natural way to trade-off quality for speed. To address this limitation, we propose Consistency Trajectory Model (CTM), a generalization encompassing CM and score-based…

2024

DiffuCOMET: Contextual Commonsense Knowledge Diffusion

ACL 2024long

Inferring contextually-relevant and diverse commonsense to understand narratives remains challenging for knowledge models. In this work, we develop a series of knowledge models, DiffuCOMET, that leverage diffusion to learn to reconstruct the implicit semantic connections between narrative contexts a…

2024

Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

ICASSP 2024accepted

Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimated by a predictive SE system. However, the pipeline structure currently does not…

Cited by 0SourceScholar
2024

Enhancing Semantic Communication with Deep Generative Models: An Overview

ICASSP 2024accepted

Semantic communication is poised to play a pivotal role in shaping the landscape of future AI-driven communication systems. Its challenge of extracting semantic information from the original complex content and regenerating semantically consistent data at the receiver, possibly being robust to chann…

Cited by 0SourceScholar
2024

GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping

NeurIPS 2024poster

Generating novel views from a single image remains a challenging task due to the complexity of 3D scenes and the limited diversity in the existing multi-view datasets to train a model on. Recent research combining large-scale text-to-image (T2I) models with monocular depth estimation (MDE) has shown…

2024

Manifold Preserving Guided Diffusion

ICLR 2024poster

Despite the recent advancements, conditional image generation still faces challenges of cost, generalizability, and the need for task-specific training. In this paper, we propose Manifold Preserving Guided Diffusion (MPGD), a training-free conditional generation framework that leverages pretrained d…

Cited by 50SourcePDFScholar
2024

MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models

IJCAI 2024poster

Recent advances in text-to-music generation models have opened new avenues in musical creativity. However, the task of editing these generated music remains a significant challenge. This paper introduces a novel approach to edit music generated by such models, enabling the modification of specific a…

2024

On the Language Encoder of Contrastive Cross-modal Models

ACL 2024findings

Contrastive cross-modal models such as CLIP and CLAP aid various vision-language (VL) and audio-language (AL) tasks. However, there has been limited investigation of and improvement in their language encoder – the central component of encoding natural language descriptions of image/audio into vector…

Cited by 0SourcePDFScholar
2024

PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher

NeurIPS 2024poster

The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the training costs through three stages: training diffusion on…

2024

SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer

ICLR 2024poster

Generative adversarial networks (GANs) learn a target probability distribution by optimizing a generator and a discriminator with minimax objectives. This paper addresses the question of whether such optimization actually provides the generator with gradients that make its distribution close to the…

2024

Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription

ICASSP 2024accepted

In recent years, research on music transcription has focused mainly on architecture design and instrument-specific data acquisition. With the lack of availability of diverse datasets, progress is often limited to solo-instrument tasks such as piano transcription. Several works have explored multi-in…

Cited by 0SourceScholar
2024

VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple Guidance

ICASSP 2024accepted

Restoring degraded music signals is essential to enhance audio quality for downstream music manipulation. Recent diffusion-based music restoration methods have demonstrated impressive performance, and among them, diffusion posterior sampling (DPS) stands out given its intrinsic properties, making it…

Cited by 0SourceScholar
2024

Zero- and Few-Shot Sound Event Localization and Detection

ICASSP 2024accepted

Sound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems have performed well in various sets of target classes, but they only output the DOA and temporal activation of preset c…

Cited by 0SourceScholar
2023

An Attention-Based Approach to Hierarchical Multi-Label Music Instrument Classification

ICASSP 2023accepted

Although music is typically multi-label, many works have studied hierarchical music tagging with simplified settings such as single-label data. Moreover, there lacks a framework to describe various joint training methods under the multi-label setting. In order to discuss the above topics, we introdu…

Cited by 0SourceScholar
2023

CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos

ICLR 2023poster

Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a target sound out of an audio mixture given a text query. Such text-queried sound…

2023

Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining Capability

ICASSP 2023accepted

In this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our m…

Cited by 0SourceScholar
2023

FP-Diffusion: Improving Score-based Diffusion Models by Enforcing the Underlying Score Fokker-Planck Equation

ICML 2023poster

Score-based generative models (SGMs) learn a family of noise-conditional score functions corresponding to the data density perturbed with increasingly large amounts of noise. These perturbed data densities are linked together by the Fokker-Planck equation (FPE), a partial differential equation (PDE)…

2023

GibbsDDRM: A Partially Collapsed Gibbs Sampler for Solving Blind Inverse Problems with Denoising Diffusion Restoration

ICML 2023oral

Pre-trained diffusion models have been successfully used as priors in a variety of linear inverse problems, where the goal is to reconstruct a signal from noisy linear measurements. However, existing approaches require knowledge of the linear operator. In this paper, we propose GibbsDDRM, an extensi…

2023

Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects

ICASSP 2023accepted

We propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording.…

Cited by 0SourceScholar
2023

PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives

ACL 2023long

Sustaining coherent and engaging narratives requires dialogue or storytelling agents to understandhow the personas of speakers or listeners ground the narrative. Specifically, these agents must infer personas of their listeners to produce statements that cater to their interests. They must also lear…

2023

STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events

NeurIPS 2023poster

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sou…

2023

Unsupervised Vocal Dereverberation with Diffusion-Based Generative Models

ICASSP 2023accepted

Removing reverb from reverberant music is a necessary technique to clean up audio for downstream music manipulations. Reverberation of music contains two categories, natural reverb, and artificial reverb. Artificial reverb has a wider diversity than natural reverb due to its various parameter setups…

Cited by 0SourceScholar
2022

Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks

ICASSP 2022accepted

A central task of a Disc Jockey (DJ) is to create a mixset of music with seamless transitions between adjacent tracks. In this paper, we explore a data-driven approach that uses a generative adversarial network to create the song transition by learning from real-world DJ mixes. The generator uses tw…

Cited by 0SourceScholar
2022

ComFact: A Benchmark for Linking Contextual Commonsense Knowledge

EMNLP 2022finding

Understanding rich narratives, such as dialogues and stories, often requires natural language processing systems to access relevant knowledge from commonsense knowledge graphs. However, these systems typically retrieve facts from KGs using simple heuristics that disregard the complex challenges of i…

2022

Multi-ACCDOA: Localizing And Detecting Overlapping Sounds From The Same Class With Auxiliary Duplicating Permutation Invariant Training

ICASSP 2022accepted

Sound event localization and detection (SELD) involves identifying the direction-of-arrival (DOA) and the event class. The SELD methods with a class-wise output format make the model predict activities of all sound event classes and corresponding locations. The class-wise methods can output activity…

Cited by 111SourceScholar
2022

Music Source Separation With Deep Equilibrium Models

ICASSP 2022accepted

While deep neural network-based music source separation (MSS) is very effective and achieves high performance, its model size is often a problem for practical deployment. Deep implicit architectures such as deep equilibrium models (DEQ) were recently proposed, which can achieve higher performance th…

Cited by 0SourceScholar
2022

SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic Quantization

ICML 2022spotlight

One noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuri…

2022

Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection

ICASSP 2022accepted

Recording and annotating real sound events for a sound event localization and detection (SELD) task is time consuming, and data augmentation techniques are often favored when the amount of data is limited. However, how to augment the spatial information in a dataset, including unlabeled directional…

Cited by 0SourceScholar
2022

Spatial Mixup: Directional Loudness Modification as Data Augmentation for Sound Event Localization and Detection

ICASSP 2022accepted

Data augmentation methods have shown great importance in diverse supervised learning problems where labeled data is scarce or costly to obtain. For sound event localization and detection (SELD) tasks several augmentation methods have been proposed, with most borrowing ideas from other domains such a…

Cited by 0SourceScholar
2021

Accdoa: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization And Detection

ICASSP 2021accepted

Neural-network (NN)-based methods show high performance in sound event localization and detection (SELD). Conventional NN-based methods use two branches for a sound event detection (SED) target and a direction-of-arrival (DOA) target. The two-branch representation with a single network has to decide…

Cited by 0SourceScholar
2021

All For One And One For All: Improving Music Separation By Bridging Networks

ICASSP 2021accepted

This paper proposes several improvements for music separation with deep neural networks (DNNs), namely a multi-domain loss (MDL) and two combination schemes. First, by using MDL we take advantage of the frequency and time domain representation of audio signals. Next, we utilize the relationship amon…

Cited by 0SourceScholar
2020

Array-Geometry-Aware Spatial Active Noise Control Based on Direction-of-Arrival Weighting

ICASSP 2020accepted

Active noise control (ANC) over a sizeable space ideally requires uniformly distributed sensors and secondary sources, which limits the feasibility of practically realizing such systems. In this paper, we propose a direction of arrival (DOA) weighting algorithm for the adaptive filter update, which…

Cited by 0SourceScholar
2020

Improving Voice Separation by Incorporating End-To-End Speech Recognition

ICASSP 2020accepted

Despite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic nature of speech by taking a transfer learning approach using an…

Cited by 0SourceScholar
2019

Global and Local Mode-domain Adaptive Algorithms for Spatial Active Noise Control Using Higher-order Sources

ICASSP 2019accepted

The aim of spatial active noise control (ANC) is to attenuate noise over a certain space. Although a large-scale system is required to achieve spatial ANC, mode-domain signal processing makes it possible to reduce the computational cost and improve the performance. A higher-order source (HOS) has an…

Cited by 11SourceScholar
2018

Mode Domain Spatial Active Noise Control Using Sparse Signal Representation

ICASSP 2018accepted

Active noise control (ANC) over a sizeable space requires a large number of reference and error microphones to satisfy the spatial Nyquist sampling criterion, which limits the feasibility of practical realization of such systems. This paper proposes a mode-domain feedforward ANC method to attenuate…

Cited by 0SourceScholar
2017

Improving music source separation based on deep neural networks through data augmentation and network blending

ICASSP 2017accepted

This paper deals with the separation of music into individual instrument tracks which is known to be a challenging problem. We describe two different deep neural network architectures for this task, a feed-forward and a recurrent one, and show that each of them yields themselves state-of-the art res…

Cited by 0SourceScholar
2016

Multichannel blind source separation based on non-negative tensor factorization in wavenumber domain

ICASSP 2016accepted

Multichannel non-negative matrix factorization based on a spatial covariance model is one of the most promising techniques for blind source separation. However, this approach is not tractable for a large number of microphones, M, because the computational cost is of order O(M <sup xmlns:mml="http://…

Cited by 0SourceScholar
2015

NMF-based blind source separation using a linear predictive coding error clustering criterion

ICASSP 2015accepted

Non-negative matrix factorization (NMF) based sound source separation involves two phases: First, the signal spectrum is decomposed into components which, in a second step, are clustered in order to obtain estimates of the source signal spectra. The major challenge with this approach is the accuracy…

Cited by 0SourceScholar