← Search

W. Bastiaan Kleijn

34 accepted papers

2025

Kolmogorov-Arnold Networks Still Catastrophically Forget but Differently from MLP

AAAI 2025technical

Catastrophic forgetting is when a neural network loses previously learnt information after learning a new task sequentially. Avoiding catastrophic forgetting could reduce the resources necessary to update neural networks. Recently, Kolmogorov–Arnold Networks (KAN) gained the community's attention as…

2025

On Exact Bit-level Reversible Transformers Without Changing Architecture

ICML 2025poster

In this work we present the BDIA-transformer, which is an exact bit-level reversible transformer that uses an unchanged standard architecture for inference. The basic idea is to first treat each transformer block as the Euler integration approximation for solving an ordinary differential equation (O…

2025

Revisiting 1-peer exponential graph for enhancing decentralized learning efficiency

NeurIPS 2025poster

For communication-efficient decentralized learning, it is essential to employ dynamic graphs designed to improve the expected spectral gap by reducing deviations from global averaging. The $1$-peer exponential graph demonstrates its finite-time convergence property--achieved by maximizing the expect…

Cited by 0SourceScholar
2024

A Practical Online Multichannel Dereverberation Approach with Data-Reuse Technique

ICASSP 2024accepted

One of the most effective online dereverberation algorithms is the weighted prediction error (WPE) method and its improved version, switching WPE (SwWPE). This paper proposes a practical online dereverberation approach to improve SwWPE by introducing a datareuse technique, DR-SwWPE; we then show ana…

Cited by 0SourceScholar
2024

Directed Diffusion: Direct Control of Object Placement through Attention Guidance

AAAI 2024technical

Text-guided diffusion models such as DALLE-2, Imagen, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In many cases the images are of very high quality. However, these models often struggle to com…

2024

Exact Diffusion Inversion via Bidirectional Integration Approximation

ECCV 2024oral

"Recently, various methods have been proposed to address the inconsistency issue of DDIM inversion to enable image editing, such as EDICT [?] and Null-text inversion [?]. However, the above methods introduce considerable computational overhead. In this paper, we propose a new technique, named bidire…

2024

On Accelerating Diffusion-Based Sampling Processes via Improved Integration Approximation

ICLR 2024poster

A popular approach to sample a diffusion-based generative model is to solve an ordinary differential equation (ODE). In existing samplers, the coefficients of the ODE solvers are pre-determined by the ODE formulation, the reverse discrete timesteps, and the employed ODE methods. In this paper, we co…

Cited by 6SourcePDFScholar
2023

LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models

ICASSP 2023accepted

We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language…

Cited by 0SourceScholar
2023

Lookahead Diffusion Probabilistic Models for Refining Mean Estimation

CVPR 2023poster

We propose lookahead diffusion probabilistic models (LA-DPMs) to exploit the correlation in the outputs of the deep neural networks (DNNs) over subsequent timesteps in diffusion probabilistic models (DPMs) to refine the mean estimation of the conditional Gaussian distributions in the backward proces…

2023

Neural Optimization Of Geometry And Fixed Beamformer For Linear Microphone Arrays

ICASSP 2023accepted

Fixed beamforming based on uniform linear microphone arrays often suffers from non-optimal performance for broadband signals. This paper addresses the issue by jointly optimizing the array geometry and spatial filters through a neural network based model. The model, composed of two feed forward neur…

Cited by 6SourceScholar
2022

Wave-Domain Approach for Cancelling Noise Entering Open Windows

ICASSP 2022accepted

Active control of noise propagating through apertures is commonly realized with closed-loop LMS algorithms. However, these algorithms require a large number of error microphones and provide only local attenuation. Slow convergence and high computational effort are additional disadvantages. We propos…

Cited by 0SourceScholar
2021

Asynchronous Decentralized Optimization With Implicit Stochastic Variance Reduction

ICML 2021spotlight

A novel asynchronous decentralized optimization method that follows Stochastic Variance Reduction (SVR) is proposed. Average consensus algorithms, such as Decentralized Stochastic Gradient Descent (DSGD), facilitate distributed training of machine learning models. However, the gradient will drift wi…

2021

Generative Speech Coding with Predictive Variance Regularization

ICASSP 2021accepted

The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of generative models deteriorates significantly with the distortions present in real-world input signals. We argue that thi…

Cited by 0SourceScholar
2020

A Linear-time Independence Criterion Based on a Finite Basis Approximation

AISTATS 2020poster

Detection of statistical dependence between random variables is an essential component in many machine learning algorithms. We propose a novel independence criterion for two random variables with linear-time complexity. We establish that our independence criterion is an upper bound of the Hirschfeld…

Cited by 4SourcePDFScholar
2020

Projected Weight Regularization to Improve Neural Network Generalization

ICASSP 2020accepted

Generalization of a deep neural network (DNN) is one major concern when employing the deep learning approach for solving practical problems. In this paper we propose a new technique, named projected weight regularization (PWR), to improve the generalization capacity of a DNN model. Consider a weight…

Cited by 0SourceScholar
2020

Robust Low Rate Speech Coding Based on Cloned Networks and Wavenet

ICASSP 2020accepted

Rapid advances in machine-learning based generative modeling of speech make its use in speech coding attractive. However, the current performance of such models drops rapidly with noise contamination of the input, preventing use in practical applications. We present a new speech-coding scheme that i…

Cited by 0SourceScholar
2018

On the Comparison of Two Room Compensation / Dereverberation Methods Employing Active Acoustic Boundary Absorption

ICASSP 2018accepted

In this paper, we compare the performance of two active dereverberation techniques using a planar array of microphones and loudspeakers. The two techniques are based on a solution to the Kirchhoff-Helmholtz Integral Equation (KHIE). We adapt a Wave Field Synthesis (WFS) based method to the applicati…

Cited by 0SourceScholar
2018

Wavenet Based Low Rate Speech Coding

ICASSP 2018accepted

Traditional parametric coding of speech facilitates low rate but provides poor reconstruction quality because of the inadequacy of the model used. We describe how a WaveNet generative speech model can be used to generate high quality speech from the bit stream of a standard parametric coder operatin…

Cited by 155SourceScholar
2017

Active speech control using wave-domain processing with a linear wall of dipole secondary sources

ICASSP 2017accepted

In this paper, we investigate the effects of compensating for wave-domain filtering delay in an active speech control system. An active control system utilising wave-domain processed basis functions is evaluated for a linear array of dipole secondary sources. The target control soundfield is matched…

Cited by 5SourceScholar
2017

Machine learning based non-intrusive quality estimation with an augmented feature set

ICASSP 2017accepted

We present a method that improves the objective quality estimation of a speech utterance. We show that including raw features that are presumably redundant reduces the effect of input noise and improves the performance of linear regressors. To exploit this effect we propose the novel idea to augment…

Cited by 0SourceScholar
2016

Distributed linear blind source separation over wireless sensor networks with arbitrary connectivity patterns

ICASSP 2016accepted

Broad areal coverage and low cost make wireless sensor networks natural platforms for blind source separation (BSS). In this context, distributed processing is attractive because of low power requirements and scalability. However, existing distributed BSS algorithms either require a fully connected…

Cited by 0SourceScholar
2016

Distributed sparse MVDR beamforming using the bi-alternating direction method of multipliers

ICASSP 2016accepted

Until now, distributed acoustic beamforming has focused on optimizing for a beamformer over an entire network, with each node contributing to the beamformer output. We present a novel approach that introduces sparsity to this beamformer computation, where we attempt to optimize for a subset of nodes…

Cited by 0SourceScholar
2016

Globally optimized least-squares post-filtering for microphone array speech enhancement

ICASSP 2016accepted

Existing post-filtering techniques for microphone array speech enhancement have two common deficiencies. First, they assume that the noise is either white or diffuse and cannot deal with point inter-ferers. Second, they estimate the post-filter coefficients using only two microphones at a time and t…

Cited by 21SourceScholar
2016

Jointly optimal near-end and far-end multi-microphone speech intelligibility enhancement based on mutual information

ICASSP 2016accepted

The processing required for the global maximization of the intelligibility of speech acquired by multiple microphones and rendered by a single loudspeaker, is considered in this paper. The intelligibility is quantized, based on the mutual information rate between the message spoken by the talker and…

Cited by 0SourceScholar
2015

Domain Generalization for Object Recognition With Multi-Task Autoencoders

ICCV 2015poster

The problem of domain generalization is to take knowledge acquired from a number of related domains, where training data is available, and to then successfully apply it to previously unseen domains. We propose a new feature learning algorithm, Multi-Task Autoencoder (MTAE), that provides good genera…

Cited by 833PDFScholar
2015

Sparse HMM-based speech enhancement method for stationary and non-stationary noise environments

ICASSP 2015accepted

We propose a sparse hidden Markov model (HMM)-based single-channel speech enhancement method that models the speech and noise gains accurately in both stationary and nonstationary environments. The objective function is augmented with an lp regularization term resulting in a sparse autoregressive HM…

Cited by 0SourceScholar