← Search

Neale Ratzlaff

6 accepted papers

2026

LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

AAAI 2026technical

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbo

Cited by 0SourcePDFScholar
2025

Probing Semantic Routing in Large Mixture-of-Expert Models

EMNLP 2025

In the past year, large ( >100 B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of efficiency, prior work has also explored functional differentiation through routing behavior. We investigate whether expe

Cited by 0SourcePDFScholar
2021

Generative Particle Variational Inference via Estimation of Functional Gradients

ICML 2021spotlight

Recently, particle-based variational inference (ParVI) methods have gained interest because they can avoid arbitrary parametric assumptions that are common in variational inference. However, many ParVI approaches do not allow arbitrary sampling from the posterior, and the few that do allow such samp…

Cited by 1SourcePDFScholar