← Search

Ankit Aich

7 accepted papers

2026

Reliable Weak-to-Strong Monitoring of LLM Agents

ICLR 2026oral

We stress test monitoring systems for detecting covert misbehavior in LLM agents (e.g., secretly exfiltrating data). We propose a monitor red teaming (MRT) workflow that varies agent and monitor awareness, adversarial evasion strategies, and evaluation across tool-calling (SHADE-Arena) and computer-…

Cited by 0SourcecodeScholar
2026

ResearchRubrics: A Benchmark of Prompts and Rubrics For Deep Research Agents

ICLR 2026poster

Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities, including multi-step reasoning, cross-document synthesis, and the generation of evidence-backed, long-form answers. Eval…

Cited by 0SourceScholar
2025

The Illusion of Empathy: How AI Chatbots Shape Conversation Perception

AAAI 2025technical

As AI chatbots increasingly incorporate empathy, understanding user-centered perceptions of chatbot empathy and its impact on conversation quality remains essential yet under-explored. This study examines how chatbot identity and perceived empathy influence users' overall conversation experience. An…

2024

Modeling Human Subjectivity in LLMs Using Explicit and Implicit Human Factors in Personas

EMNLP 2024finding

Large language models (LLMs) are increasingly being used in human-centered social scientific tasks, such as data annotation, synthetic data creation, and engaging in dialog. However, these tasks are highly subjective and dependent on human factors, such as one’s environment, attitudes, beliefs, and…

Cited by 4SourcePDFScholar
2024

Towards Comprehensive Language Analysis for Clinically Enriched Spontaneous Dialogue

COLING 2024main

Contemporary NLP has rapidly progressed from feature-based classification to fine-tuning and prompt-based techniques leveraging large language models. Many of these techniques remain understudied in the context of real-world, clinically enriched spontaneous dialogue. We fill this gap by systematical…

2022

Demystifying Neural Fake News via Linguistic Feature-Based Interpretation

COLING 2022main

The spread of fake news can have devastating ramifications, and recent advancements to neural fake news generators have made it challenging to understand how misinformation generated by these models may best be confronted. We conduct a feature-based study to gain an interpretative understanding of t…

Cited by 21SourcePDFScholar
2022

Towards Intelligent Clinically-Informed Language Analyses of People with Bipolar Disorder and Schizophrenia

EMNLP 2022finding

NLP offers a myriad of opportunities to support mental health research. However, prior work has almost exclusively focused on social media data, for which diagnoses are difficult or impossible to validate. We present a first-of-its-kind dataset of manually transcribed interactions with people clinic…