← Search

Taneesh Gupta

4 accepted papers

2025

AMPO: Active Multi Preference Optimization for Self-play Preference Selection

ICML 2025poster

Multi-preference optimization enriches language-model alignment beyond pairwise preferences by contrasting entire sets of helpful and undesired responses, enabling richer training signals for large language models. During self-play alignment, these models often produce numerous candidate answers per…

Cited by 0SourcePDFScholar
2025

CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling

ACL 2025finding

Reward modeling in large language models is known to be susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In RLHF, and more generally during post-training, flawed reward signals often lead to outpu…

Cited by 0SourcePDFScholar
2024

Map It Anywhere: Empowering BEV Map Prediction using Large-scale Public Datasets

NeurIPS 2024poster

Top-down Bird's Eye View (BEV) maps are a popular perception representation for ground robot navigation due to their richness and flexibility for downstream tasks. While recent methods have shown promise for predicting BEV maps from First-Person View (FPV) images, their generalizability is limited t…

2023

On Designing Light-Weight Object Trackers Through Network Pruning: Use CNNS or Transformers?

ICASSP 2023accepted

Object trackers deployed on low-power devices need to be light-weight, however, most of the current state-of-the-art (SOTA) methods rely on using compute-heavy backbones built using CNNs or Transformers. Large sizes of such models do not allow their deployment in low-power conditions and designing c…

Cited by 0SourceScholar