← Search

Hiteshi Sharma

5 accepted papers

2024

Enhancing Language Model Alignment: A Confidence-Based Approach to Label Smoothing

EMNLP 2024main

In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains. Within the training pipeline of LLMs, the Reinforcement Learning with Human Feedback (RLHF) phase is crucial for aligning LLMs with human preferences and values. Label smoothing, a techniq…

Cited by 0SourcePDFScholar
2024

Language Models can be Deductive Solvers

NAACL 2024findings

Logical reasoning is a fundamental aspect of human intelligence and a key component of tasks like problem-solving and decision-making. Recent advancements have enabled Large Language Models (LLMs) to potentially exhibit reasoning capabilities, but complex logical reasoning remains a challenge. The s…

2023

Evaluating Cognitive Maps and Planning in Large Language Models with CogEval

NeurIPS 2023poster

Recently an influx of studies claims emergent cognitive abilities in large language models (LLMs). Yet, most rely on anecdotes, overlook contamination of training sets, or lack systematic Evaluation involving multiple tasks, control conditions, multiple iterations, and statistical robustness tests.…

Cited by 64SourcePDFScholar
2020

Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes

ICML 2020poster

Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduced for learning infinite-horizon average-reward Markov Decision Processes (MDPs). The first algorithm reduces the problem…

Cited by 135SourcePDFScholar
2019

Approximate Relative Value Learning for Average-reward Continuous State MDPs

UAI 2019poster

In this paper, we propose an approximate relative value learning (ARVL) algorithm for non- parametric MDPs with continuous state space and finite actions and average reward criterion. It is a sampling based algorithm combined with kernel density estimation and function approximation via nearest neig…

Cited by 17SourcePDFScholar