← Search

Mayank Vatsa

25 accepted papers

2026

ACID Test: A Benchmark for Cultural Safety and Alignment in LALMs

AAAI 2026technical

Large Audio Language Models (LALMs) are transforming AI by processing and generating human language directly from audio. As these models proliferate in real-world applications, it becomes critical to evaluate their performance to ensure equitable and safe use across diverse linguistic and cultural c

Cited by 0SourcePDFScholar
2026

NutriScreener: Retrieval Augmented Multi-Pose Graph Attention Network for Malnourishment Screening

AAAI 2026technical

Child malnutrition remains a global crisis, yet existing screening methods are laborious and poorly scalable, hindering early intervention. In this work, we present NutriScreener, a retrieval-augmented, multi-pose graph attention network that combines CLIP-based visual embeddings, class-boosted know

Cited by 0SourcePDFScholar
2026

Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation

AAAI 2026technical

The architectural blueprint of today’s leading text-to-image models contains a fundamental flaw: an inability to handle logical composition. This survey investigates this breakdown across three core primitives—negation, counting, and spatial relations. Our analysis reveals a dramatic performance col

Cited by 0SourcePDFScholar
2025

AQUAFace: Age-Invariant Quality Adaptive Face Recognition for Unconstrained Selfie vs ID Verification

AAAI 2025technical

Face recognition in the presence of age and quality variations poses a formidable challenge. While recent margin-based loss functions have shown promise in addressing these variations individually, real-world scenarios such as selfie versus ID face matching often involve simultaneous variations of b…

Cited by 0SourcePDFScholar
2025

Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?

ICASSP 2025accepted

The proliferation of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), yet their development has largely overlooked low-resource languages. This paper addresses this disparity by evaluating three prominent Large Audio Language Models (LALMs) – LTU-AS, GAMA, and Pengi –…

Cited by 0SourceScholar
2025

Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models

CVPR 2025poster

Existing unlearning algorithms in text-to-image generative models often fail to preserve the knowledge of semantically related concepts when removing specific target concepts--a challenge known as adjacency. To address this, we propose FADE (Fine-grained Attenuation for Diffusion Erasure), introduci…

Cited by 3SourcePDFScholar
2025

Harmonizing Geometry and Uncertainty: Diffusion with Hyperspheres

ICML 2025poster

Do contemporary diffusion models preserve the class geometry of hyperspherical data? Standard diffusion models rely on isotropic Gaussian noise in the forward process, inherently favoring Euclidean spaces. However, many real-world problems involve non-Euclidean distributions, such as hyperspherical…

Cited by 0SourcePDFScholar
2025

ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake Dataset

ICLR 2025poster

The proliferation of deepfakes and AI-generated content has led to a surge in media forgeries and misinformation, necessitating robust detection systems. However, current datasets lack diversity across modalities, languages, and real-world scenarios. To address this gap, we present ILLUSION (Integra…

Cited by 0SourcePDFScholar
2025

SHIELD: A Self-supervised, Silicosis-focused Hierarchical Imaging Framework for Occupational Lung Disease Diagnosis

IJCAI 2025

Silicosis is an irreversible lung disease caused by silica dust exposure in industrial settings. Early detection is crucial, but automatic diagnostic methods are hindered by limited data availability. We propose SHIELD - a self-supervised, Silicosis-focused Hierarchical Imaging framework for early o

Cited by 0SourcePDFScholar
2024

BirdCollect: A Comprehensive Benchmark for Analyzing Dense Bird Flock Attributes

AAAI 2024technical

Automatic recognition of bird behavior from long-term, un controlled outdoor imagery can contribute to conservation efforts by enabling large-scale monitoring of bird populations. Current techniques in AI-based wildlife monitoring have focused on short-term tracking and monitoring birds individually…

Cited by 2SourcePDFScholar
2024

ToonerGAN: Reinforcing GANs for Obfuscating Automated Facial Indexing

CVPR 2024poster

The rapid evolution of automatic facial indexing tech- nologies increases the risk of compromising personal and sensitive information. To address the issue we propose cre- ating cartoon avatars or 'toon avatars' designed to effec- tively obscure identity features. The primary objective is to deceive…

Cited by 2SourcePDFScholar
2023

DF-Platter: Multi-Face Heterogeneous Deepfake Dataset

CVPR 2023poster

Deepfake detection is gaining significant importance in the research community. While most of the research efforts are focused around high-quality images and videos, deepfake generation algorithms today have the capability to generate low-resolution videos, occluded deepfakes, and multiple-subject d…

Cited by 60SourcePDFScholar
2023

IdProv: Identity-Based Provenance for Synthetic Image Generation (Student Abstract)

AAAI 2023technical

Recent advancements in Generative Adversarial Networks (GANs) have made it possible to obtain high-quality face images of synthetic identities. These networks see large amounts of real faces in order to learn to generate realistic looking synthetic images. However, the concept of a synthetic identit…

Cited by 1SourcePDFScholar
2023

Long-term Monitoring of Bird Flocks in the Wild

IJCAI 2023poster

Monitoring and analysis of wildlife are key to conservation planning and conflict management. The widespread use of camera traps coupled with AI-based analysis tools serves as an excellent example of successful and non-invasive use of technology for design, planning, and evaluation of conservation p…

Cited by 1SourcePDFScholar
2023

NutriAI: AI-Powered Child Malnutrition Assessment in Low-Resource Environments

IJCAI 2023poster

Malnutrition among infants and young children is a pervasive public health concern, particularly in developing countries where resources are limited. Millions of children globally suffer from malnourishment and its complications1. Despite the best efforts of governments and organizations, malnourish…

Cited by 3SourcePDFScholar
2023

On AI-Assisted Pneumoconiosis Detection from Chest X-rays

IJCAI 2023poster

According to theWorld Health Organization, Pneumoconiosis affects millions of workers globally, with an estimated 260,000 deaths annually. The burden of Pneumoconiosis is particularly high in low-income countries, where occupational safety standards are often inadequate, and the prevalence of…

Cited by 2SourcePDFScholar
2023

Uncovering the Deceptions: An Analysis on Audio Spoofing Detection and Future Prospects

IJCAI 2023poster

Audio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications including person authentication to banking to virtual assistants. Research has shown that these syst…

Cited by 7SourcePDFScholar
2022

Mannet: A Large-Scale Manipulated Image Detection Dataset And Baseline Evaluations

ICASSP 2022accepted

The sharing of fake content on social media platforms has become a major concern. In many cases, the same content with small variations is shared multiple times on different social media platforms. This leads to the circulation of manipulated content on the web. With the rapid advancement in deep le…

Cited by 0SourceScholar
2019

Dual Directed Capsule Network for Very Low Resolution Image Recognition

ICCV 2019poster

Very low resolution (VLR) image recognition corresponds to classifying images with resolution 16x16 or less. Though it has widespread applicability when objects are captured at a very large stand-off distance (e.g. surveillance scenario) or from wide angle mobile cameras, it has received limited att…

Cited by 81PDFScholar
2018

Learning Structure and Strength of CNN Filters for Small Sample Size Training

CVPR 2018poster

Convolutional Neural Networks have provided state-of-the-art results in several computer vision problems. However, due to a large number of parameters in CNNs, they require a large number of training samples which is a limiting factor for small sample size problems. To address this limitation, in th…

Cited by 127SourcePDFScholar
2017

Face Sketch Matching via Coupled Deep Transform Learning

ICCV 2017poster

Face sketch to digital image matching is an important challenge of face recognition that involves matching across different domains. Current research efforts have primarily focused on extracting domain invariant representations or learning a mapping from one domain to the other. In this research, we…

Cited by 30PDFScholar