← Search

Heinrich Dinkel

14 accepted papers

2026

ACAVCAPS: ENABLING LARGE-SCALE TRAINING FOR FINE-GRAINED AND DIVERSE AUDIO UNDERSTANDING

ICASSP 2026poster

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the scale and descriptive granularity required to train truly ve…

Cited by 0SourcePDFScholar
2026

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

ICML 2026poster

While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation metrics, fail to reliably distinguish between generic and highl…

Cited by 0SourceScholar
2024

CED: Consistent Ensemble Distillation for Audio Tagging

ICASSP 2024accepted

Augmentation and knowledge distillation (KD) are well-established techniques employed in audio classification tasks, aimed at enhancing performance and reducing model sizes on the widely recognized Audioset (AS) benchmark. Although both techniques are effective individually, their combined use, call…

Cited by 0SourceScholar
2023

Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction

ICASSP 2023accepted

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio an…

Cited by 0SourceScholar
2023

Unified Keyword Spotting and Audio Tagging on Mobile Devices with Transformers

ICASSP 2023accepted

Keyword spotting (KWS) is a core human-machine-interaction front-end task for most modern intelligent assistants. Recently, a unified (UniKW-AT) framework has been proposed that adds additional capabilities in the form of audio tagging (AT) to a KWS model. However, previous work did not consider the…

Cited by 0SourceScholar
2022

Category-Adapted Sound Event Enhancement with Weakly Labeled Data

ICASSP 2022accepted

Previous audio enhancement training usually requires clean signals with additive noises; hence commonly focuses on speech enhancement, where clean speech is easy to access. This paper goes beyond a broader sound event enhancement by using a weakly supervised approach via sound event detection (SED)…

Cited by 0SourceScholar
2022

Pseudo Strong Labels for Large Scale Weakly Supervised Audio Tagging

ICASSP 2022accepted

Large-scale audio tagging datasets inevitably contain imperfect labels, such as clip-wise annotated (temporally weak) tags with no exact on- and offsets, due to a high manual labeling cost. This work proposes pseudo strong labels (PSL), a simple label augmentation framework that enhances the supervi…

Cited by 0SourceScholar
2021

Investigating Local and Global Information for Automated Audio Captioning with Transfer Learning

ICASSP 2021accepted

Automated audio captioning (AAC) aims at generating summarizing descriptions for audio clips. Multitudinous concepts are described in an audio caption, ranging from local information such as sound events to global information like acoustic scenery. Currently, the mainstream paradigm for AAC is the e…

Cited by 0SourceScholar
2021

Text-to-Audio Grounding: Building Correspondence Between Captions and Sound Events

ICASSP 2021accepted

Automated Audio Captioning is a cross-modal task, generating natural language descriptions to summarize the audio clips’ sound events. However, grounding the actual sound events in the given audio based on its corresponding caption has not been investigated. This paper contributes an Audio-Grounding…

Cited by 0SourceScholar
2020

Multiple Sound Sources Localization from Coarse to Fine

ECCV 2020poster

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework that disentangles audio and visual representations of different…