← Search

Philip J. B. Jackson

12 accepted papers

2025

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

ICLR 2025poster

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle…

2024

Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection

ICASSP 2024accepted

Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included. Few audio-visual (AV)-SELD works have been published and mo…

Cited by 0SourceScholar
2024

Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event Classification

ICASSP 2024accepted

In the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structure…

Cited by 0SourceScholar
2019

Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge

ICASSP 2019accepted

Humans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the crea…

Cited by 0SourceScholar
2019

Robust Full-sphere Binaural Sound Source Localization Using Interaural and Spectral Cues

ICASSP 2019accepted

A binaural sound source localization method is proposed that uses interaural and spectral cues for localization of sound sources with any direction of arrival on the full-sphere. The method is designed to be robust to the presence of reverberation, additive noise and different types of sounds. The m…

Cited by 0SourceScholar
2018

Acoustic Reflector Localization and Classification

ICASSP 2018accepted

The process of understanding acoustic properties of environments is important for several applications, such as spatial audio, augmented reality and source separation. In this paper, multichannel room impulse responses are recorded and transformed into their direction of arrival (DOA)-time domain, b…

Cited by 0SourceScholar
2018

Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech Separation

ICASSP 2018accepted

In this paper, we propose an iterative deep neural network (DNN)-based binaural source separation scheme, for recovering two concurrent speech signals in a room environment. Besides the commonly-used spectral features, the DNN also takes non-linearly wrapped binaural spatial features as input, which…

Cited by 0SourceScholar
2018

Synthesis of Images by Two-Stage Generative Adversarial Networks

ICASSP 2018accepted

In this paper, we propose a divide-and-conquer approach using two generative adversarial networks (GANs) to explore how a machine can draw colorful pictures (bird) using a small amount of training data. In our work, we simulate the procedure of an artist drawing a picture, where one begins with draw…

Cited by 0SourceScholar
2017

Fast tagging of natural sounds using marginal co-regularization

ICASSP 2017accepted

Automatic and fast tagging of natural sounds in audio collections is a very challenging task due to wide acoustic variations, the large number of possible tags, the incomplete and ambiguous tags provided by different labellers. To handle these problems, we use a co-regularization approach to learn a…

Cited by 0SourceScholar
2015

IVA algorithms using a multivariate Student's t source prior for speech source separation in real room environments

ICASSP 2015accepted

The independent vector analysis (IVA) algorithm employs a multivariate source prior to retain the dependency between different frequency bins of each source and thereby avoids the permutation problem that is inherent to blind source separation (BSS). In this paper, a multivariate Student's t distrib…

Cited by 0SourceScholar