← Search

Sadbhavana Babar

3 accepted papers

2026

Graph-Theoretic Intrinsic Reward: Guiding RL with Effective Resistance

ICLR 2026poster

Exploration of dynamic environments with sparse rewards is a significant challenge in Reinforcement Learning, often leading to inefficient exploration and brittle policies. To address this, we introduce a novel graph-based intrinsic reward using Effective Resistance, a metric from spectral graph the…

Cited by 0SourceScholar
2025

Beyond Mere Token Analysis: A Hypergraph Metric Space Framework for Defending Against Socially Engineered LLM Attacks

ICLR 2025poster

Recent jailbreak attempts on Large Language Models (LLMs) have shifted from algorithm-focused to human-like social engineering attacks, with persuasion-based techniques emerging as a particularly effective subset. These attacks evolve rapidly, demonstrate high creativity, and boast superior attack s…

Cited by 0SourcePDFScholar
2025

SafeQuant: LLM Safety Analysis via Quantized Gradient Inspection

NAACL 2025long

Contemporary jailbreak attacks on Large Language Models (LLMs) employ sophisticated techniques with obfuscated content to bypass safety guardrails. Existing defenses either use computationally intensive LLM verification or require adversarial fine-tuning, leaving models vulnerable to advanced attack…

Cited by 0SourcePDFScholar