← Search

Shivam Bhardwaj

2 accepted papers

2026

Graph-Theoretic Intrinsic Reward: Guiding RL with Effective Resistance

ICLR 2026poster

Exploration of dynamic environments with sparse rewards is a significant challenge in Reinforcement Learning, often leading to inefficient exploration and brittle policies. To address this, we introduce a novel graph-based intrinsic reward using Effective Resistance, a metric from spectral graph the…

Cited by 0SourceScholar
2025

EFFICIENT JAILBREAK ATTACK SEQUENCES ON LARGE LANGUAGE MODELS VIA MULTI-ARMED BANDIT-BASED CONTEXT SWITCHING

ICLR 2025poster

Content warning: This paper contains examples of harmful language and content. Recent advances in large language models (LLMs) have made them increasingly vulnerable to jailbreaking attempts, where malicious users manipulate models into generating harmful content. While existing approaches rely on e…

Cited by 0SourcePDFScholar