← Search

Hanan Salam

9 accepted papers

2026

AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

ICML 2026poster

Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I models, leaving the prompting proficiency of this upstream compo…

Cited by 0SourceScholar
2026

CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding

ICML 2026poster

LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to capture this shift, excluding problems that require both human reasoning to guide solutions and AI efficiency for implementati…

Cited by 7SourceScholar
2026

Knowledge Synthesis in Dynamic Human-Swarm Interactions Using LLMs

ICRA 2026poster

Collectively exploring and understanding an environment is an open challenge, particularly in dynamic settings where agents must rely on limited information that may only be intermittently available. In this paper, we focus on how agents can maximize information capture in these contexts. As agents …

Cited by 0Scholar
2025

AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents

NeurIPS 2025poster

Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators often miss dangers in agents' step-by-step actions, overlook subtle meanings, fail to see how small issues compound, an…

Cited by 0SourcecodeScholar
2025

Agentic-ToM: Cognition-Inspired Agentic Processing For Enhancing Theory of Mind Reasoning

EMNLP 2025

The capacity to attribute mental states like beliefs, desires, and intentions to oneself and others, known as Theory of Mind (ToM), is fundamental to human social intelligence. As Large Language Models (LLMs) are increasingly integrated into complex interactive systems, developing their ToM capabili

Cited by 0SourcePDFScholar
2025

Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition

COLING 2025main

Theory of Mind (ToM) is the ability to understand and reflect on the mental states of others. Although this capability is crucial for human interaction, testing on Large Language Models (LLMs) reveals that they possess only a rudimentary understanding of it. Although the most capable closed-source L…

2025

DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognition

EMNLP 2025

The advancements of Large Language Models (LLMs) have spurred a growing interest in their application to Named Entity Recognition (NER) methods. However, existing datasets are primarily designed for traditional machine learning methods and are inadequate for LLM-based methods, in terms of corpus sel