← Search

Robert Ness

4 accepted papers

2026

Multiverse Mechanica: A Testbed for Learning Game Mechanics via Counterfactual Worlds

ICLR 2026poster

We study how generative world models trained on video games can go beyond mere reproduction of gameplay visuals to learning game mechanics—the modular rules that causally govern gameplay. We introduce a formalization of the concept of game mechanics that operationalizes mechanic-learning as a causal…

Cited by 0SourceScholar
2025

Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations

ICLR 2025spotlight

Large language models (LLMs) are capable of generating *plausible* explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be *unfaithful*. This, in turn, can lead to over-trust and misuse. We introduce…

2023

Evaluating Cognitive Maps and Planning in Large Language Models with CogEval

NeurIPS 2023poster

Recently an influx of studies claims emergent cognitive abilities in large language models (LLMs). Yet, most rely on anecdotes, overlook contamination of training sets, or lack systematic Evaluation involving multiple tasks, control conditions, multiple iterations, and statistical robustness tests.…

Cited by 64SourcePDFScholar
2019

Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systems

NeurIPS 2019poster

This manuscript contributes a general and practical framework for casting a Markov process model of a system at equilibrium as a structural causal model, and carrying out counterfactual inference. Markov processes mathematically describe the mechanisms in the system, and predict the system’s equilib…