← Search

Gaurav Srivastava

9 accepted papers

2026

A Diagnostic Study of Multi-Agent LLMs for Real-World Debates

ICML 2026poster

Multi-agent LLM debates are increasingly deployed in domains such as policy analysis and city planning, where no objective ground truth exists. Despite this, debate quality is typically evaluated using outcome-based proxies such as LLM-as-judge scores that provide little insight into whether meaning…

Cited by 0SourceScholar
2026

BeyondBench: Benchmark-Free Evaluation of Reasoning in Language Models

ICLR 2026poster

Evaluating language models fairly is becoming harder as static benchmarks available on the internet risk contamination by training data. This makes it unclear whether models are truly reasoning or just recalling answers. In this paper, we introduce $\textbf{BeyondBench}$, an evaluation framework tha…

Cited by 0SourcecodeScholar
2026

JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation

AAAI 2026technical

While small language models (SLMs) have shown promise on various reasoning tasks, their ability to judge the correctness of answers remains unclear compared to large language models (LLMs). Prior work on LLM-as-a-judge frameworks typically relies on comparing candidate answers against ground-truth l

Cited by 0SourcePDFScholar
2026

Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs

CVPR 2026

While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images," i.e., to reason through multi-step visual interactions, remains limited. We introduce VISTA-Gym, a scalable training environment for incentivizing tool-integrated visual reasoning

Cited by 0SourcecodeScholar
2026

effGen: Enabling Small Language Models as Capable Autonomous Agents

ICML 2026poster

Most existing language model agentic systems today are built and optimized for large language models (e.g., GPT, Claude, Gemini) via API calls. While powerful, this approach faces several limitations including high token costs and privacy concerns for sensitive applications. We introduce $\textbf{ef…

Cited by 0SourceScholar
2025

DEBATE, TRAIN, EVOLVE: Self‐Evolution of Language Model Reasoning

EMNLP 2025

Large language models (LLMs) have improved significantly in their reasoning through extensive training on massive datasets. However, relying solely on additional data for improvement is becoming increasingly impractical, highlighting the need for models to autonomously enhance their reasoning withou

2024

Sample-Efficient Personalization: Modeling User Parameters as Low Rank Plus Sparse Components

AISTATS 2024poster

Personalization of machine learning (ML) predictions for individual users/domains/enterprises is critical for practical recommendation systems. Standard personalization approaches involve learning a user/domain specific \emph{embedding} that is fed into a fixed global model which can be limiting. On…

Cited by 1SourcePDFScholar
2019

Joint Optimization of Quantization and Structured Sparsity for Compressed Deep Neural Networks

ICASSP 2019accepted

The usage of Deep Neural Networks (DNN) on resource-constrained edge devices has been limited due to their high computation and large memory requirement. In this work, we propose an algorithm to compress DNNs by jointly optimizing structured sparsity and quantization constraints in a single DNN trai…

Cited by 0SourceScholar