← Search

Arda Senocak

14 accepted papers

2026

How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?

CVPR 2026

We present the first scalable framework for training sound source localization (SSL) models using synthetic data from text-to-X models. Although SSL has made notable progress, existing models remain constrained by limited-scale, uncurated real-world datasets that often suffer from semantic misalignm

Cited by 0SourceScholar
2026

Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

CVPR 2026

We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The

Cited by 0SourcecodeScholar
2025

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

ICLR 2025poster

Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understandi…

2025

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

NeurIPS 2025poster

We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guid…

Cited by 0SourcecodeScholar
2025

Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes

CVPR 2025poster

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently…

Cited by 0SourcePDFScholar
2024

From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers

ICASSP 2024accepted

Transformers have become central to recent advances in audio classification. However, training an audio spectrogram transformer, e.g. AST, from scratch can be resource and time-intensive. Furthermore, the complexity of transformers heavily depends on the input audio spectrogram size. In this work, w…

Cited by 0SourceScholar
2024

Speech Guided Masked Image Modeling for Visually Grounded Speech

ICASSP 2024accepted

The objective of this study is to investigate the learning process of Visually Grounded Speech (VGS) models through joint learning that combines contrastive learning and masked image modeling. Typically, VGS models ahn to establish audio-visual alignment between images and then spoken captions withi…

Cited by 0SourceScholar
2023

Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples

ICASSP 2023accepted

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both languages. However, in reality, there can be an imbalance among the languages for…

Cited by 0SourceScholar
2023

Sound Source Localization is All about Cross-Modal Alignment

ICCV 2023poster

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective. However, prior arts and existing benchmarks do not account for…

Cited by 18PDFScholar
2023

Sound to Visual Scene Generation by Audio-to-Visual Latent Alignment

CVPR 2023poster

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We design a model that works by scheduling the learning procedur…

2022

Learning Sound Localization Better from Semantically Similar Samples

ICASSP 2022accepted

The objective of this work is to localize the sound sources in visual scenes. Existing audio-visual works employ contrastive learning by assigning corresponding audio-visual pairs from the same source as positives while randomly mismatched pairs as negatives. However, these negative pairs may contai…

Cited by 0SourceScholar
2018

Learning to Localize Sound Source in Visual Scenes

CVPR 2018poster

Visual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene pairs like human? In this paper, we propose a novel unsupervis…

Cited by 397SourcePDFScholar