← Search

Anmol Kabra

5 accepted papers

2026

Learning from Synthetic Data Improves Multi-hop Reasoning

ICLR 2026poster

Reinforcement Learning (RL) has been shown to significantly boost reasoning capabilities of large language models (LLMs) in math, coding, and multi-hop reasoning tasks. However, RL fine-tuning requires abundant high-quality verifiable data, often obtained through human-annotated datasets and LLM-as-…

Cited by 0SourcecodeScholar
2025

PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation

ICML 2025poster

High-quality benchmarks are essential for evaluating reasoning and retrieval capabilities of large language models (LLMs). However, curating datasets for this purpose is not a permanent solution as they are prone to data leakage and inflated performance results. To address these challenges, we prop…

2022

Exponential Family Model-Based Reinforcement Learning via Score Matching

NeurIPS 2022accept

We propose an optimistic model-based algorithm, dubbed SMRL, for finite-horizon episodic reinforcement learning (RL) when the transition model is specified by exponential family distributions with $d$ parameters and the reward is bounded and known. SMRL uses score matching, an unnormalized density e…

2021

Characterizing the Loss Landscape in Non-Negative Matrix Factorization

AAAI 2021technical

Non-negative matrix factorization (NMF) is a highly celebrated algorithm for matrix decomposition that guarantees non-negative factors. The underlying optimization problem is computationally intractable, yet in practice, gradient-descent-based methods often find good solutions. In this paper, we rev…

Cited by 8SourcePDFScholar