← Search

Kaivalya Hariharan

2 accepted papers

2024

Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

NeurIPS 2024poster

AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model”. This raises questions. Do such models "know'' that they are LLMs and reliably act on this knowledge? Are they "aware" of their current circumstances, such as being deployed to the public? We refer…

2023

Red Teaming Deep Neural Networks with Feature Synthesis Tools

NeurIPS 2023poster

Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts. Despite the attention this area of study receives, there are comparatively few cases where these tools have identified previously unknown bugs in models. We argue that this i…