← Search

Juncheng B. Li

5 accepted papers

2024

Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio Generation

ICASSP 2024accepted

Despite recent progress, machine learning (ML) models for open-domain audio generation need to catch up to generative models for image, text, speech, and music. The lack of massive open-domain audio datasets is the main reason for this performance gap; we overcome this challenge through a novel data…

Cited by 0SourceScholar
2022

Masked Autoencoders that Listen

NeurIPS 2022accept

This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only th…

2022

On Adversarial Robustness Of Large-Scale Audio Visual Learning

ICASSP 2022accepted

As audio-visual systems are being deployed for safety-critical tasks such as surveillance and malicious content filtering, their robustness remains an under-studied area. Existing published work on robustness either does not scale to large-scale dataset, or does not deal with multiple modalities. Th…

Cited by 0SourceScholar
2021

Audio-Visual Event Recognition Through the Lens of Adversary

ICASSP 2021accepted

As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving the accuracy. This work aims to study several key questions related to multimodal learning through the lens of adversarial…

Cited by 0SourceScholar