← Search

Craig T. Jin

12 accepted papers

2025

TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances

AAAI 2025technical

The concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene.…

2024

Active Noise Control Over 3D Space with A Dynamic Noise Source

ICASSP 2024accepted

Spatial Active noise control (ANC) systems are proposed to minimize the noise over a spatial region of interest around people’s heads by generating an anti-noise field with multiple microphones and loudspeakers. Recently, a realistic microphone geometry was designed to allow the system to monitor th…

Cited by 4SourceScholar
2024

Addressing Data Scarcity in Voice Disorder Detection with Self-Supervised Models

ICASSP 2024accepted

Machine learning (ML) has shown promising results in the field of voice disorder detection over the past decade. However, the diversity of recording conditions, audio content, languages, and the scarcity of examples for each of these combinations pose a challenge in building ML models that can relia…

Cited by 0SourceScholar
2024

From RIR to BRIR: A Sparse Recovery Beamforming Approach for Virtual Binaural Sound Rendering

ICASSP 2024accepted

The creation of a spatial sound scene through binaural rendering draws increasing research interest, given the rising application of virtual reality and augmented reality. High-fidelity augmented reality audio requires accurate room acoustic simulation. Typically, this is achieved with binaural rend…

Cited by 0SourceScholar
2019

Improved Multipath Time Delay Estimation Using Cepstrum Subtraction

ICASSP 2019accepted

When a motor-powered vessel travels past a fixed hydrophone in a multipath environment, a Lloyd's mirror constructive/destructive interference pattern is observed in the output spectrogram. The power cepstrum detects the periodic structure of the Lloyd's mirror pattern by generating a sequence of pu…

Cited by 0SourceScholar
2018

Considerations Regarding Individualization of Head-Related Transfer Functions

ICASSP 2018accepted

This paper provides some considerations regarding using individualized head-related transfer functions for rendering binaural spatial audio over headphones. It briefly considers the degree of benefit that individualization may provide. It then examines the degree of variation existing within the ear…

Cited by 0SourceScholar
2018

Sound Source Localization in a Multipath Environment Using Convolutional Neural Networks

ICASSP 2018accepted

The propagation of sound in a shallow water environment is characterized by boundary reflections from the sea surface and sea floor. These reflections result in multiple (indirect) sound propagation paths, which can degrade the performance of passive sound source localization methods. This paper pro…

Cited by 0SourceScholar
2017

Convolutional neural networks for passive monitoring of a shallow water environment using a single sensor

ICASSP 2017accepted

A cost effective approach to remote monitoring of protected areas such as marine reserves and restricted naval waters is to use passive sonar to detect, classify, localize, and track marine vessel activity (including small boats and autonomous underwater vehicles). Cepstral analysis of underwater ac…

Cited by 0SourceScholar
2017

Kernel principal component analysis of the ear morphology

ICASSP 2017accepted

This paper describes features in the ear shape that change across a population of ears and explores the corresponding changes in ear acoustics. The statistical analysis conducted over the space of ear shapes uses a kernel principal component analysis (KPCA). Further, it utilizes the framework of lar…

Cited by 0SourceScholar
2016

Generating a morphable model of ears

ICASSP 2016accepted

This paper describes the generation of a morphable model for external ear shapes. The aim for the morphable model is to characterize an ear shape using only a few parameters in order to assist the study of morphoacoustics. The model is derived from a statistical analysis of a population of 58 ears f…

Cited by 0SourceScholar
2015

Super-resolution acoustic imaging using sparse recovery with spatial priming

ICASSP 2015accepted

In this paper, we propose a new strategy to obtain superresolution maps of the sound field recorded by a spherical microphone array. In recent works, we have demonstrated that sparse recovery (SR) algorithms based on the minimisation of the l <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns…

Cited by 0SourceScholar