← Search

Rishabh Ranjan

19 accepted papers

2026

ACID Test: A Benchmark for Cultural Safety and Alignment in LALMs

AAAI 2026technical

Large Audio Language Models (LALMs) are transforming AI by processing and generating human language directly from audio. As these models proliferate in real-world applications, it becomes critical to evaluate their performance to ensure equitable and safe use across diverse linguistic and cultural c

Cited by 0SourcePDFScholar
2026

PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models

ICML 2026poster

Relational Foundation Models (RFMs) facilitate data-driven decision-making by learning from complex multi-table databases. However, the diverse relational databases needed to train such models are rarely public due to privacy constraints. While there are methods to generate synthetic tabular data of…

Cited by 0SourceScholar
2026

Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data

ICLR 2026poster

Pretrained transformers readily adapt to new sequence modeling tasks via zero-shot prompting, but relational domains still lack architectures that transfer across datasets and tasks. The core challenge is the diversity of relational data, with varying heterogeneous schemas, graph structures, and fun…

Cited by 0SourcecodeScholar
2025

Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?

ICASSP 2025accepted

The proliferation of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), yet their development has largely overlooked low-resource languages. This paper addresses this disparity by evaluating three prominent Large Audio Language Models (LALMs) – LTU-AS, GAMA, and Pengi –…

Cited by 0SourceScholar
2025

ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake Dataset

ICLR 2025poster

The proliferation of deepfakes and AI-generated content has led to a surge in media forgeries and misinformation, necessitating robust detection systems. However, current datasets lack diversity across modalities, languages, and real-world scenarios. To address this gap, we present ILLUSION (Integra…

Cited by 0SourcePDFScholar
2025

SHIELD: A Self-supervised, Silicosis-focused Hierarchical Imaging Framework for Occupational Lung Disease Diagnosis

IJCAI 2025

Silicosis is an irreversible lung disease caused by silica dust exposure in industrial settings. Early detection is crucial, but automatic diagnostic methods are hindered by limited data availability. We propose SHIELD - a self-supervised, Silicosis-focused Hierarchical Imaging framework for early o

Cited by 0SourcePDFScholar
2024

Position: Relational Deep Learning - Graph Representation Learning on Relational Databases

ICML 2024poster

Much of the world's most valued data is stored in relational databases and data warehouses, where the data is organized into tables connected by primary-foreign key relations. However, building machine learning models using this data is both challenging and time consuming because no ML algorithm can…

Cited by 12SourcePDFScholar
2024

Post-Hoc Reversal: Are We Selecting Models Prematurely?

NeurIPS 2024poster

Trained models are often composed with post-hoc transforms such as temperature scaling (TS), ensembling and stochastic weight averaging (SWA) to improve performance, robustness, uncertainty estimation, etc. However, such transforms are typically applied only after the base models have already been f…

2024

RelBench: A Benchmark for Deep Learning on Relational Databases

NeurIPS 2024poster

We present RelBench, a public benchmark for solving predictive tasks in relational databases with deep learning. RelBench provides databases and tasks spanning diverse domains, scales, and database dimensions, and is intended to be a foundational infrastructure for future research in this direction…

Cited by 11SourcePDFScholar
2024

SelfVC: Voice Conversion With Iterative Refinement using Self Transformations

ICML 2024poster

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that separately encode speaker characteristics and linguistic content.…

Cited by 7SourcePDFScholar
2023

On AI-Assisted Pneumoconiosis Detection from Chest X-rays

IJCAI 2023poster

According to theWorld Health Organization, Pneumoconiosis affects millions of workers globally, with an estimated 260,000 deaths annually. The burden of Pneumoconiosis is particularly high in low-income countries, where occupational safety standards are often inadequate, and the prevalence of…

Cited by 2SourcePDFScholar
2023

Uncovering the Deceptions: An Analysis on Audio Spoofing Detection and Future Prospects

IJCAI 2023poster

Audio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications including person authentication to banking to virtual assistants. Research has shown that these syst…

Cited by 7SourcePDFScholar
2022

A Solver-free Framework for Scalable Learning in Neural ILP Architectures

NeurIPS 2022accept

There is a recent focus on designing architectures that have an Integer Linear Programming (ILP) layer within a neural model (referred to as \emph{Neural ILP} in this paper). Neural ILP architectures are suitable for pure reasoning tasks that require data-driven constraint learning or for tasks requ…

2022

GREED: A Neural Framework for Learning Graph Distance Functions

NeurIPS 2022accept

Similarity search in graph databases is one of the most fundamental operations in graph analytics. Among various distance functions, graph and subgraph edit distances (GED and SED respectively) are two of the most popular and expressive measures. Unfortunately, exact computations for both are NP-har…

Cited by 58SourcePDFScholar
2021

Multi-Scale Residual Network for Covid-19 Diagnosis Using Ct-Scans

ICASSP 2021accepted

To mitigate the outbreak of highly contagious COVID-19, we need a sensitive, robust automated diagnostic tool. This paper proposes a three-level approach to separate the cases of COVID-19, pneumonia from normal patients using chest CT scans. At the first level, we fine tune a multi-scale ResNet50 mo…

Cited by 0SourceScholar
2019

Parametric Hear through Equalization for Augmented Reality Audio

ICASSP 2019accepted

Augmented Reality (AR) audio applications require headphones to be acoustically transparent so that real sounds can pass through unaltered for natural fusion with virtual sounds. In this paper, we consider a multiple source scenario for hear through (HT) equalization (EQ) using closed-back circumaur…

Cited by 0SourceScholar
2017

Fast HRFT measurement system with unconstrained head movements for 3D audio in virtual and augmented reality applications

ICASSP 2017accepted

Binaural audio plays an indispensable role in virtual reality (VR) and augmented reality (AR). Binaural audio recreates the sensation of the three dimensional auditory experience using Head- Related Transfer Functions (HRTFs). HRTFs are as unique as our fingerprint. To achieve an immersive audio exp…

Cited by 7SourceScholar
2016

Fast continuous HRTF acquisition with unconstrained movements of human subjects

ICASSP 2016accepted

Head related transfer function (HRTF) is widely used in 3D audio reproduction, especially over headphones. Conventionally, HRTF database is acquired at discrete directions and the acquisition process is time-consuming. Recent works have been proposed to improve HRTF acquisition efficiency via contin…

Cited by 0SourceScholar