← Search

Matthew Lyle Olson

4 accepted papers

2026

Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy

AAAI 2026technical

We present the first evaluation harness that enables any out-of-the-box, local, Large Language Models (LLMs) to play full-press Diplomacy without fine-tuning or specialized training. Previous work required frontier LLMs, or fine-tuning, due to the high complexity and information density of Diplomacy

Cited by 0SourcePDFScholar
2026

LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

AAAI 2026technical

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbo

Cited by 0SourcePDFScholar
2025

Probing Semantic Routing in Large Mixture-of-Expert Models

EMNLP 2025

In the past year, large ( >100 B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expe

Cited by 0SourcePDFScholar
2024

Why do LLaVA Vision-Language Models Reply to Images in English?

EMNLP 2024finding

We uncover a surprising multilingual bias occurring in a popular class of multimodal vision-language models (VLMs). Including an image in the query to a LLaVA-style VLM significantly increases the likelihood of the model returning an English response, regardless of the language of the query. This pa…

Cited by 4SourcePDFScholar