← Search

Krishna Mandal

2 accepted papers

2025

VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models

ICLR 2025poster

Large language models (LLMs) often exhibit subtle yet distinctive characteristics in their outputs that users intuitively recognize, but struggle to quantify. These "vibes" -- such as tone, formatting, or writing style -- influence user preferences, yet traditional evaluations focus primarily on the…

2025

VisionArena: 230k Real World User-VLM Conversations with Preference Labels

CVPR 2025poster

The growing adoption and capabilities of vision-language models (VLMs) demand benchmarks that reflect real-world user interactions. We introduce VisionArena, the largest existing dataset of crowdsourced real-world conversations between users and VLMs. While most visual question-answering datasets fo…