← Search

Julia Kiseleva

5 accepted papers

2026

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation

ICML 2026poster

We present \textbf{ExCyTIn-Bench}, the first benchmark to \textbf{E}valuate an LLM agent \textbf{X} on the task of \textbf{Cy}ber \textbf{T}hreat \textbf{In}vestigation through security questions derived from investigation graphs. Real‑world security analysts must sift through a large number of hete…

Cited by 0SourceScholar
2024

Assessing and Verifying Task Utility in LLM-Powered Applications

EMNLP 2024main

The rapid development of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents, assisting humans in their daily tasks. However, a significant gap remains in assessing to what extent LLM-powered applications genuinely enhance user experien…

2022

What Makes a Good and Useful Summary? Incorporating Users in Automatic Summarization Research

NAACL 2022long

Automatic text summarization has enjoyed great progress over the years and is used in numerous applications, impacting the lives of many. Despite this development, there is little research that meaningfully investigates how the current research focus in automatic summarization aligns with users’ nee…

2021

Building and Evaluating Open-Domain Dialogue Corpora with Clarifying Questions

EMNLP 2021main

Enabling open-domain dialogue systems to ask clarifying questions when appropriate is an important direction for improving the quality of the system response. Namely, for cases when a user request is not specific enough for a conversation system to provide an answer right away, it is desirable to as…

2021

Learning to Decompose and Organize Complex Tasks

NAACL 2021long

People rely on digital task management tools, such as email or to-do apps, to manage their tasks. Some of these tasks are large and complex, leading to action paralysis and feelings of being overwhelmed on the part of the user. The micro-productivity literature has shown that such tasks could benefi…