← Search

Qingming Tang

18 accepted papers

2025

Effective Techniques for Scaling Audio Encoder Pretraining

ICASSP 2025accepted

This work presents advancements in audio pretraining objectives designed to generate semantically rich embeddings, capable of addressing a wide range of audio-related tasks. Despite significant progress in the field, current methods often emphasize full fine-tuning in downstream applications, which…

Cited by 0SourceScholar
2025

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

ICML 2025poster

Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous natu…

Cited by 0SourcePDFScholar
2025

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling

ICML 2025poster

Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio generation. Despite achieving high audio fidelity, they incur significant inference…

2024

Cross-Triggering Issue in Audio Event Detection and Mitigation

ICASSP 2024accepted

Cross-triggering is a critical problem for applications of audio event detection (AED), particularly in low-resource settings. However, not much attention (if not none) has been paid to this problem in the AED research community. In this work, we tackle this problem via a regularization approach. We…

Cited by 0SourceScholar
2024

Learning for Transductive Threshold Calibration in Open-World Recognition

CVPR 2024poster

In deep metric learning for visual recognition the calibration of distance thresholds is crucial for achieving desired model performance in the true positive rates (TPR) or true negative rates (TNR). However calibrating this thresh- old presents challenges in open-world scenarios where the test clas…

Cited by 0SourcePDFScholar
2024

On-Device Constrained Self-Supervised Learning for Keyword Spotting via Quantization Aware Pre-Training and Fine-Tuning

ICASSP 2024accepted

Large self-supervised models have excelled in various speech processing tasks, but their deployment on resource-limited devices is often impractical due to their substantial memory footprint. Previous studies have demonstrated the effectiveness of self-supervised pre-training for keyword spotting, e…

Cited by 0SourceScholar
2024

Threshold-Consistent Margin Loss for Open-World Deep Metric Learning

ICLR 2024poster

Existing losses used in deep metric learning (DML) for image retrieval often lead to highly non-uniform intra-class and inter-class representation structures across test classes and data distributions. When combined with the common practice of using a fixed threshold to declare a match, this gives r…

Cited by 4SourcePDFScholar
2023

FedRPO: Federated Relaxed Pareto Optimization for Acoustic Event Classification

ICASSP 2023accepted

Performance and robustness of real-world Acoustic Event Classification (AEC) solutions depend on ability to train on diverse data from wide range of end-point devices and acoustic environments. Federated Learning (FL) provides a framework to leverage annotated and non-annotated AEC data from servers…

Cited by 0SourceScholar
2023

Weight-Sharing Supernet for Searching Specialized Acoustic Event Classification Networks Across Device Constraints

ICASSP 2023accepted

Acoustic Event Classification (AEC) has been widely used in devices such as smart speakers and mobile phones for home safety or accessibility support [1]. As AEC models run on more and more devices with diverse computation resource constraints, it became increasingly expensive to develop models that…

Cited by 0SourceScholar
2022

Federated Self-Supervised Learning for Acoustic Event Classification

ICASSP 2022accepted

Standard acoustic event classification (AEC) solutions require large-scale collection of data from client devices for model optimization. Federated learning (FL) is a compelling frame- work that decouples data collection and model training to enhance customer privacy. In this work, we investigate th…

Cited by 0SourceScholar
2022

Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology

ICASSP 2022accepted

Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is…

Cited by 0SourceScholar
2022

Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event Classification

ICASSP 2022accepted

Acoustic event classification (AEC) is the task of determining whether certain events occur in an audio clip. Inspired by previous research [1], [2], [3] that embeddings from event labels can be leveraged to facilitate the learning of new detectors with no or limited audio samples, we introduce Wiki…

Cited by 0SourceScholar
2021

Multi-Task Self-Supervised Pre-Training for Music Classification

ICASSP 2021accepted

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. B…

Cited by 0SourceScholar
2020

Unsupervised Pre-Training of Bidirectional Speech Encoders via Masked Reconstruction

ICASSP 2020accepted

We propose an approach for pre-training speech representations via a masked reconstruction loss. Our pre-trained encoder networks are bidirectional and can therefore be used directly in typical bidirectional speech recognition models. The pre-trained networks can then be fine-tuned on a smaller amou…

Cited by 0SourceScholar
2019

Hierarchical Residual-pyramidal Model for Large Context Based Media Presence Detection

ICASSP 2019accepted

We study media presence detection, that is, learning to recognize if a sound segment (typically lasting for a few seconds) of a long recorded stream contains media (TV) sound. This problem is difficult because non-media sound sources can be quite diverse (e.g. human voicing, non-vocal sounds and non…

Cited by 0SourceScholar
2018

Dependency-aware Attention Control for Unconstrained Face Recognition with Image Sets

ECCV 2018poster

This paper targets the problem of image set-based face verification and identification. Unlike traditional single media (an image or video) setting, we encounter a set of heterogeneous contents containing orderless images and videos. The importance of each image is usually considered either equal or…

Cited by 53SourcePDFScholar