← Search

Denis Peskoff

7 accepted papers

2026

ResearchRubrics: A Benchmark of Prompts and Rubrics For Deep Research Agents

ICLR 2026poster

Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities, including multi-step reasoning, cross-document synthesis, and the generation of evidence-backed, long-form answers. Eval…

Cited by 0SourceScholar
2025

Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?

EMNLP 2025

The social impact of Natural Language Processing (NLP) is increasingly important, with a rising community focus on initiatives related to NLP for Social Good (NLP4SG). Indeed, in recent years, almost 20% of all papers in the ACL Anthology address topics related to social good as defined by the UN Su

2025

Personalized Help for Optimizing Low-Skilled Users’ Strategy

NAACL 2025short

AIs can beat humans in game environments; however, how helpful those agents are to human remains understudied. We augment Cicero, a natural language agent that demonstrates superhuman performance in Diplomacy, to generate both move and message advice based on player intentions. A dozen Diplomacy gam…

Cited by 0SourcePDFScholar
2025

Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL

ACL 2025finding

An increasingly common socio-technical problem is people being taken in by offers that sound “too good to be true”, where persuasion and trust shape decision-making. This paper investigates how AI can help detect these deceptive scenarios. We analyze how humans strategically deceive each other in Di…

Cited by 0SourcePDFScholar
2024

More Victories, Less Cooperation: Assessing Cicero’s Diplomacy Play

ACL 2024long

The boardgame Diplomacy is a challenging setting for communicative and cooperative artificial intelligence. The most prominent communicative Diplomacy AI, Cicero, has excellent strategic abilities, exceeding human players. However, the best Diplomacy players master communication, not just tactics, w…

2023

GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and Doves

EMNLP 2023short findings

Markets and policymakers around the world hang on the consequential monetary policy decisions made by the Federal Open Market Committee (FOMC). Publicly available textual documentation of their meetings provides insight into members’ attitudes about the economy. We use GPT-4 to quantify dissent amon…

Cited by 0SourcecodeScholar