← Search

Shreyas Chaudhari

10 accepted papers

2025

Peer-to-Peer Learning Dynamics of Wide Neural Networks

ICASSP 2025accepted

Peer-to-peer learning is an increasingly popular framework that enables beyond-5G distributed edge devices to collaboratively train deep neural networks in a privacy-preserving manner without the aid of a central server. Neural network training algorithms for emerging environments, e.g., smart citie…

Cited by 0SourceScholar
2025

PersonaGym: Evaluating Persona Agents and LLMs

EMNLP 2025

Persona agents, which are LLM agents conditioned to act according to an assigned persona, enable contextually rich and user-aligned interactions across domains like education and healthcare.However, evaluating how faithfully these agents adhere to their personas remains a significant challenge, part

Cited by 0SourcePDFScholar
2024

Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation

NeurIPS 2024poster

Evaluating policies using off-policy data is crucial for applying reinforcement learning to real-world problems such as healthcare and autonomous driving. Previous methods for *off-policy evaluation* (OPE) generally suffer from high variance or irreducible bias, leading to unacceptably high predicti…

2024

Distributional Off-Policy Evaluation for Slate Recommendations

AAAI 2024technical

Recommendation strategies are typically evaluated by using previously logged data, employing off-policy evaluation methods to estimate their expected performance. However, for strategies that present users with slates of multiple items, the resulting combinatorial action space renders many of these…

2024

From Past to Future: Rethinking Eligibility Traces

AAAI 2024technical

In this paper, we introduce a fresh perspective on the challenges of credit assignment and policy evaluation. First, we delve into the nuances of eligibility traces and explore instances where their updates may result in unexpected credit assignment to preceding states. From this investigation emerg…

Cited by 1SourcePDFScholar
2023

Learning Gradients of Convex Functions with Monotone Gradient Networks

ICASSP 2023accepted

While much effort has been devoted to deriving and analyzing effective convex formulations of signal processing problems, the gradients of convex functions also have critical applications ranging from gradient-based optimization to optimal transport. Recent works have explored data-driven methods fo…

Cited by 0SourceScholar
2021

A Unified Approach to Translate Classical Bandit Algorithms to Structured Bandits

ICASSP 2021accepted

We consider a finite-armed structured bandit problem in which mean rewards of different arms are known functions of a common hidden parameter θ <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">*</sup> . This problem setting subsumes several previously st…

Cited by 0SourceScholar
2021

Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers

ICLR 2021poster

We propose a simple, practical, and intuitive approach for domain adaptation in reinforcement learning. Our approach stems from the idea that the agent's experience in the source domain should look similar to its experience in the target domain. Building off of a probabilistic view of RL, we achieve…

Cited by 105SourcePDFScholar
2021

Unsupervised Clustering of Time Series Signals Using Neuromorphic Energy-Efficient Temporal Neural Networks

ICASSP 2021accepted

Unsupervised time series clustering is a challenging problem with diverse industrial applications such as anomaly detection, bio-wearables, etc. These applications typically involve small, low-power devices on the edge that collect and process real-time sensory signals. State-of-the-art time-series…

Cited by 0SourceScholar