← Search

Elia Bruni

9 accepted papers

2025

Agents generalize to novel levels of abstraction by using adaptive linguistic strategies

ACL 2025finding

We study abstraction in an emergent communication paradigm. In emergent communication, two artificial neural network agents develop a language while solving a communicative task. In this study, the agents play a concept-level reference game. This means that the speaker agent has to describe a concep…

Cited by 0SourcePDFScholar
2025

Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs

EMNLP 2025

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general knowledge. SPLICE includes 3,381 human-filtered videos spanning 12 ca

Cited by 0SourcePDFScholar
2025

iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs

EMNLP 2025

Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. iVISPAR is based on a varia

2024

Context Shapes Emergent Communication about Concepts at Different Levels of Abstraction

COLING 2024main

We study the communication of concepts at different levels of abstraction and in different contexts in an agent-based, interactive reference game. While playing the concept-level reference game, the neural network agents develop a communication system from scratch. We use a novel symbolic dataset th…

Cited by 4SourcePDFScholar
2024

GRASP: A Novel Benchmark for Evaluating Language GRounding and Situated Physics Understanding in Multimodal Language Models

IJCAI 2024poster

This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-tier approach leveraging Unity simulations. The first level tests for language gro…

2022

Emergence of Hierarchical Reference Systems in Multi-agent Communication

COLING 2022main

In natural language, referencing objects at different levels of specificity is a fundamental pragmatic mechanism for efficient communication in context. We develop a novel communication game, the hierarchical reference game, to study the emergence of such reference systems in artificial agents. We c…

2022

The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study

ACL 2022long

Obtaining human-like performance in NLP is often argued to require compositional generalisation. Whether neural networks exhibit this ability is usually studied by training models on highly compositional synthetic data. However, compositionality in natural language is much more complex than the rigi…

2020

Compositionality Decomposed: How do Neural Networks Generalise? (Extended Abstract)

IJCAI 2020poster

Despite a multitude of empirical studies, little consensus exists on whether neural networks are able to generalise compositionally. As a response to this controversy, we present a set of tests that provide a bridge between, on the one hand, the vast amount of linguistic and philosophical theory abo…

Cited by 0SourcePDFScholar