← Search

Ramit Debnath

2 accepted papers

2026

The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

ICML 2026poster

Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral tendencies resembling human traits like agreeableness and self-reg…

Cited by 0SourceScholar
2025

Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes

ACL 2025long

Large Language Models (LLMs) are often used as automated judges to evaluate text, but their effectiveness can be hindered by various unintentional biases. We propose using linear classifying probes, trained by leveraging differences between contrasting pairs of prompts, to directly access LLMs’ late…