← Search

Saksham Singh Kushwaha

5 accepted papers

2026

Object-WIPER: Training-Free Object and Associated Effect Removal in Videos

CVPR 2026

In this paper, we introduce Object-WIPER, a training-free framework for removing dynamic objects and their associated visual effects from videos, and inpainting them with semantically consistent and temporally coherent content. Our approach leverages a pre-trained text-to-video diffusion transformer

Cited by 0SourceScholar
2026

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

CVPR 2026

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical instruments, and human speech). Prior video-conditioned audio genera

Cited by 0SourcecodeScholar
2025

$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time

NeurIPS 2025poster

While recent audio-visual models have demonstrated impressive performance, their robustness to distributional shifts at test-time remains not fully understood. Existing robustness benchmarks mainly focus on single modalities, making them insufficient for thoroughly assessing the robustness of audio-…

Cited by 0SourcecodeScholar
2025

Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

ICASSP 2025accepted

Spatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalability, and assume independence between semantic and spatial information. To address these issues, we explore end-to-end spat…

Cited by 0SourceScholar