← Search

Anisha Gunjal

4 accepted papers

2026

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than bi…

Cited by 0SourceScholar
2024

Detecting and Preventing Hallucinations in Large Vision Language Models

AAAI 2024technical

Instruction tuned Large Vision Language Models (LVLMs) have significantly advanced in generalizing across a diverse set of multi-modal tasks, especially for Visual Question Answering (VQA). However, generating detailed responses that are visually grounded is still a challenging task for these models…

2024

Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification

EMNLP 2024finding

Automatic factuality verification of large language model (LLM) generations is becoming more and more widely used to combat hallucinations. A major point of tension in the literature is the granularity of this fact-checking: larger chunks of text are hard to fact-check, but more atomic facts like pr…