← Search

Shital Shah

6 accepted papers

2026

Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Creative Writing

ICML 2026poster

Reinforcement learning with verifiable rewards (RLVR) on foundation models has led to significant improvements in math and code generation. Extending these gains to open-ended domains remains challenging: ground-truth verification is unavailable, human annotation is expensive, and learnt reward mode…

Cited by 0SourceScholar
2026

h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning

ICML 2026spotlight

Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or step-level supervision, neither of which scales easily. In this work, we introduce a scalable method to …

Cited by 0SourceScholar
2022

LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models

NeurIPS 2022accept

The Transformer architecture is ubiquitously used as the building block of largescale autoregressive language models. However, finding architectures with the optimal trade-off between task performance (perplexity) and hardware constraints like peak memory utilization and latency is non-trivial. This…

2021

Understanding Failures of Deep Networks via Robust Feature Extraction

CVPR 2021poster

Traditional evaluation metrics for learned models that report aggregate scores over a test set are insufficient for surfacing important and informative patterns of failure over features and instances. We introduce and study a method aimed at characterizing and explaining failures by identifying visu…

Cited by 87PDFcodeScholar
2020

Safe Reinforcement Learning via Curriculum Induction

NeurIPS 2020spotlight

In safety-critical applications, autonomous agents may need to learn in an environment where mistakes can be very costly. In such settings, the agent needs to behave safely not only after but also while learning. To achieve this, existing safe reinforcement learning methods make an agent rely on pri…

2017

Submodular Trajectory Optimization for Aerial 3D Scanning

ICCV 2017poster

Drones equipped with cameras are emerging as a powerful tool for large-scale aerial 3D scanning, but existing automatic flight planners do not exploit all available information about the scene, and can therefore produce inaccurate and incomplete 3D models. We present an automatic method to generate…

Cited by 177PDFScholar