← Search

Vanya Cohen

5 accepted papers

2026

MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models

ICML 2026poster

Entity state tracking is a necessary component of world modeling that requires maintaining coherent representations of entities over time. Previous work has benchmarked entity tracking performance in purely text-based tasks. We introduce MET-Bench, a multimodal entity tracking benchmark designed to …

Cited by 0SourceScholar
2024

A Survey of Robotic Language Grounding: Tradeoffs between Symbols and Embeddings

IJCAI 2024poster

With large language models, robots can understand language more flexibly and more capable than ever before. This survey reviews and situates recent literature into a spectrum with two poles: 1) mapping between language and some manually defined formal representation of meaning, and 2) mapping betwee…

Cited by 11SourcePDFScholar
2024

CAPE: Corrective Actions from Precondition Errors using Large Language Models

ICRA 2024poster

Extracting knowledge and reasoning from large language models (LLMs) offers a path to designing intelligent robots. Common approaches that leverage LLMs for planning are unable to recover when actions fail and resort to retrying failed actions without resolving the underlying cause. We propose a nov…

Cited by 34SourceScholar
2024

CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans

EMNLP 2024main

Understanding the abilities of LLMs to reason about natural language plans, such as instructional text and recipes, is critical to reliably using them in decision-making systems. A fundamental aspect of plans is the temporal order in which their steps need to be executed, which reflects the underlyi…

2019

Grounding Language Attributes to Objects using Bayesian Eigenobjects

IROS 2019poster

We develop a system to disambiguate object instances within the same class based on simple physical descriptions. The system takes as input a natural language phrase and a depth image containing a segmented object and predicts how similar the observed object is to the object described by the phrase.…

Cited by 23SourceScholar