← Search

Vaibhav Singh

7 accepted papers

2026

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

ICML 2026poster

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We introduce DiffuMamba, a masked diffusion language model built on a b…

Cited by 0SourceScholar
2026

Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away

ICML 2026poster

Reinforcement learning (RL) based post-training for explicit chain-of-thought (e.g., GRPO) improves the reasoning ability of multimodal large-scale reasoning models (MLRMs). But recent evidence shows that it can simultaneously degrade safety alignment and increase jailbreak success rates. We propose…

Cited by 0SourceScholar
2026

VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion

ICML 2026poster

Group Emotion Recognition (GER) is crucial for understanding social dynamics, ranging from interpreting intimate conversations to evaluating crowd behavior in large-scale surveillance scenarios. While current AI models can analyze these scenes, they often act as black boxes that take shortcuts. Inst…

Cited by 0SourceScholar
2025

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification

NAACL 2025findings

We propose ARISE, a framework that iteratively induces rules and generates synthetic data for text classification. We combine synthetic data generation and automatic rule induction, via bootstrapping, to iteratively filter the generated rules and data. We induce rules via inductive generalisation of…

Cited by 0SourcePDFScholar
2025

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

CVPR 2025poster

With the widespread deployment of Multimodal Large Language Models (MLLMs) for visual-reasoning tasks, improving their safety has become crucial. Recent research indicates that despite training-time safety alignment, these models remain vulnerable to jailbreak attacks--carefully crafted image-prompt…

Cited by 3SourcePDFScholar
2025

RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples

EMNLP 2025

Reward models are essential for aligning large language models (LLMs) with human preferences. However, most open-source multilingual reward models are primarily trained on preference datasets in high-resource languages, resulting in unreliable reward signals for low-resource Indic languages. Collect

Cited by 0SourcePDFScholar
2021

Speech Recognition Using RFID Tattoos (Extended Abstract)

IJCAI 2021poster

This paper presents a radio-frequency (RF) based assistive technology for voice impairments (i.e., dysphonia), which occurs in an estimated 1% of the global population. We specifically focus on acquired voice disorders where users continue to be able to make facial and lip gestures associated with s…

Cited by 0SourcePDFScholar