← Search

Andreas Bulling

20 accepted papers

2026

MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning

ICML 2026poster

Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. Despite the motivation of lifelong learning, most continual RL papers consider only 3–10 sequential tasks, as CPU-bound environments make longer sequences impracti…

Cited by 0SourceScholar
2026

RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo

ICLR 2026poster

Standard benchmarks for optical flow, scene flow, and stereo vision algorithms generally focus on model accuracy rather than robustness to image corruptions like noise or rain. Hence, the resilience of models to such real-world perturbations is largely unquantified. To address this, we present Robus…

Cited by 0SourceScholar
2026

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

ICML 2026oral

We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD generates training partners on-the-fly and selects them adaptively based on a learnability criterion, removing the need for pre-trained partner populations or m…

Cited by 0SourceScholar
2025

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

EMNLP 2025

Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is critical - not only to move beyond surface-level performance, but also for mode

2025

ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions

EMNLP 2025

Most existing Theory of Mind (ToM) benchmarks for foundation models rely on variations of the Sally-Anne test, offering only a very limited perspective on ToM and neglecting the complexity of human social interactions. To address this gap, we propose ToM-SSI: a new benchmark specifically designed to

Cited by 0SourcePDFScholar
2025

V^2Dial: Unification of Video and Visual Dialog via Multimodal Experts

CVPR 2025poster

We present V2Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler tasks (e.g., VQA, VideoQA, video-text retrieval) and often neglect the more challeng…

Cited by 0SourcePDFScholar
2024

InteRead: An Eye Tracking Dataset of Interrupted Reading

COLING 2024main

Eye movements during reading offer a window into cognitive processes and language comprehension, but the scarcity of reading data with interruptions – which learners frequently encounter in their everyday learning environments – hampers advances in the development of intelligent learning technologie…

Cited by 4SourcePDFScholar
2024

Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition

ACL 2024long

Recent work on dialogue-based collaborative plan acquisition (CPA) has suggested that Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. Although ToM was claimed to be important for effective collaboration, its real impact on…

Cited by 6SourcePDFScholar
2024

OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog

COLING 2024main

We present the Object Language Video Transformer (OLViT) – a novel model for video dialog operating over a multi-modal attention-based dialog state tracker. Existing video dialog models struggle with questions requiring both spatial and temporal localization within videos, long-term temporal reasoni…

Cited by 2SourcePDFScholar
2021

Neural Photofit: Gaze-Based Mental Image Reconstruction

ICCV 2021poster

We propose a novel method that leverages human fixations to visually decode the image a person has in mind into a photofit (facial composite). Our method combines three neural networks: An encoder, a scoring network, and a decoder. The encoder extracts image features and predicts a neural activation…

Cited by 14PDFScholar
2020

Improving Natural Language Processing Tasks with Human Gaze-Guided Neural Attention

NeurIPS 2020poster

A lack of corpora has so far limited advances in integrating human gaze data as a supervisory signal in neural attention mechanisms for natural language processing (NLP). We propose a novel hybrid text saliency model (TSM) that, for the first time, combines a cognitive model of reading with explicit…

Cited by 85SourcePDFScholar
2015

Prediction of Search Targets From Fixations in Open-World Settings

CVPR 2015poster

Previous work on predicting the target of visual search from human fixations only considered closed-world settings in which training labels are available and predictions are performed for a known set of potential targets. In this work we go beyond the state of the art by studying search target predi…

Cited by 75SourcePDFScholar
2015

Rendering of Eyes for Eye-Shape Registration and Gaze Estimation

ICCV 2015poster

Images of the eye are key in several computer vision problems, such as shape registration and gaze estimation. Recent large-scale supervised methods for these problems require time-consuming data collection and manual annotation, which can be unreliable. We propose synthesizing perfectly labelled ph…

Cited by 422PDFScholar