← Search

Gaurav Bharaj

13 accepted papers

2026

FoeGlass: When Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

ICML 2026poster

Audio deepfake detection (ADD) models are critical for countering the malicious use of text-to-speech (TTS) models. Evaluating and strengthening ADD models requires developing datasets that span the space of generated audio and highlight high-error regions. Existing dataset development strategies fa…

Cited by 0SourceScholar
2025

Common Sense Bias Modeling for Classification Tasks

AAAI 2025technical

Machine learning model bias can arise from dataset composition: correlated sensitive features can distort the downstream classification model's decision boundary and lead to performance differences along these features. Existing de-biasing works tackle the most prominent bias features, such as color…

Cited by 0SourcePDFScholar
2025

PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image Detectors

NeurIPS 2025poster

Synthetic image detectors (SIDs) are a key defense against the risks posed by the growing realism of images from text-to-image (T2I) models. Red teaming improves SID’s effectiveness by identifying and exploiting their failure modes via misclassified synthetic images. However, existing red-teaming so…

Cited by 0SourceScholar
2025

What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain

ICASSP 2025accepted

Adding explanations to audio deepfake detection (ADD) models will enable insights on the decision making process and thus boost their real-world application. In this paper, we propose a relevancy-based explainable AI (XAI) method to analyze the predictions of transformer-based ADD models. We compare…

Cited by 0SourceScholar
2024

AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection

CVPR 2024poster

With the rapid growth in deepfake video content we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the audio and visual modalities. While the former disregards the au…

Cited by 18SourcePDFScholar
2024

SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection

NeurIPS 2024poster

Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized by generative AI models. Existing ADD models suffer from generalization issues to unseen attacks, with a large performance discrepancy between in-domain and out-of-domain data. Moreover, the black-box nature of exis…

Cited by 10SourcePDFScholar
2023

Few-Shot Geometry-Aware Keypoint Localization

CVPR 2023poster

Supervised keypoint localization methods rely on large manually labeled image datasets, where objects can deform, articulate, or occlude. However, creating such large keypoint labels is time-consuming and costly, and is often error-prone due to inconsistent labeling. Thus, we desire an approach that…

2023

Implicit Neural Head Synthesis via Controllable Local Deformation Fields

CVPR 2023poster

High-quality reconstruction of controllable 3D head avatars from 2D videos is highly desirable for virtual human applications in movies, games, and telepresence. Neural implicit fields provide a powerful representation to model 3D head avatars with personalized shape, expressions, and facial parts,…

Cited by 12SourcePDFScholar
2023

Unsupervised Facial Performance Editing via Vector-Quantized StyleGAN Representations

ICCV 2023poster

High-fidelity virtual human avatar applications create a need for photorealistic video face synthesis with controllable semantic editing over facial features. While recent generative neural methods have shown significant progress in portrait video synthesis, intuitive facial control, e.g., of mouth…

Cited by 1PDFcodeScholar
2020

StyleRig: Rigging StyleGAN for 3D Control Over Portrait Images

CVPR 2020oral

StyleGAN generates photorealistic portrait images of faces with eyes, teeth, hair and context (neck, shoulders, background), but lacks a rig-like control over semantic face parameters that are interpretable in 3D, such as face pose, expressions, and scene illumination. Three-dimensional morphable fa…

Cited by 473PDFScholar
2019

FML: Face Model Learning From Videos

CVPR 2019oral

Monocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on data-driven priors that are built from limited 3D face scans. In…

Cited by 179PDFScholar