← Search

Chetan Bansal

10 accepted papers

2026

ICPO: Provable and Practical In-Context Policy Optimization for Test-Time Scaling

ICLR 2026poster

We study test-time scaling, where a model improves its answer through multi-round self-reflection at inference. We introduce In-Context Policy Optimization (ICPO), in which an agent optimizes its response in context using self-assessed or externally observed rewards without modifying its parameters.…

Cited by 0SourceScholar
2026

Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

ICML 2026poster

Agent memory systems must accommodate continuously growing information while supporting efficient, context-aware retrieval for downstream tasks. Abstraction is essential for scaling agent memory, yet it often comes at the cost of specificity, obscuring the fine-grained details required for effective…

Cited by 0SourceScholar
2025

AMPO: Active Multi Preference Optimization for Self-play Preference Selection

ICML 2025poster

Multi-preference optimization enriches language-model alignment beyond pairwise preferences by contrasting entire sets of helpful and undesired responses, enabling richer training signals for large language models. During self-play alignment, these models often produce numerous candidate answers per…

Cited by 0SourcePDFScholar
2025

Anyprefer: An Agentic Framework for Preference Data Synthesis

ICLR 2025poster

High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding approach, where the target model generates and annotates its…

Cited by 0SourcePDFScholar
2025

CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling

ACL 2025finding

Reward modeling in large language models is known to be susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In RLHF, and more generally during post-training, flawed reward signals often lead to outpu…

Cited by 0SourcePDFScholar
2025

CREAM: Consistency Regularized Self-Rewarding Language Models

ICLR 2025poster

Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations for preference data. These methods commonly utilize the same LLM to act as both the policy model (which generates response…

2025

Generative Caching for Structurally Similar Prompts and Responses

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly being used to plan, reason, and execute tasks across diverse scenarios. In use cases like repeatable workflows and agentic settings, prompts are often reused with minor variations while having a similar structure for recurring tasks. This opens up opport…

Cited by 0SourceScholar
2025

Synergistic Weak-Strong Collaboration by Aligning Preferences

ACL 2025long

Current Large Language Models excel in general reasoning yet struggle with specialized tasks requiring proprietary or domain-specific knowledge. Fine-tuning large models for every niche application is often infeasible due to black-box constraints and high computational overhead. To address this, we…

Cited by 0SourcePDFScholar
2025

Verifiable Format Control for Large Language Model Generations

NAACL 2025findings

Recent Large Language Models (LLMs) have demonstrated satisfying general instruction following ability. However, small LLMs with about 7B parameters still struggle fine-grained format following (e.g., JSON format), which seriously hinder the advancements of their applications. Most existing methods…

2024

CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models

NeurIPS 2024poster

Artificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing s…