← Search

Yonatan Bisk

63 accepted papers

2026

Affordance RAG: Hierarchical Multimodal Retrieval With Affordance-Aware Embodied Memory for Mobile Manipulation

RA-L 2026

In this study, we address the problem of openvocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on freeform natural language instructions. This task is challenging, as it involves understanding visual semantics and the affordance of manipu

Cited by 2SourceScholar
2026

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

ICRA 2026poster

In this study, we address the problem of open-vocabulary mobile manipulation, where a robot is required to carry a wide range of objects to receptacles based on free-form natural language instructions. This task is challenging, as it involves understanding visual semantics and the affordance of mani…

2026

MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning

CVPR 2026

Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal language models (MLLMs) for this purpose, but their substantial computational cost hinders practical application.This limi

Cited by 0SourcecodeScholar
2026

Position: Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment

ICML 2026poster

Proper accounting of the energy requirements and environmental impact of artificial intelligence (AI) systems is necessary for researchers, developers, policy makers, and users to assess the barriers to building systems at scale. With the growing complexity of pipelines and underlying infrastructure…

Cited by 0SourceScholar
2026

RAVEN: Resilient Aerial Navigation Via Open-Set Semantic Memory and Behavior Adaptation

ICRA 2026poster

Aerial outdoor semantic navigation requires robots to explore large, unstructured environments to locate target objects. Recent advances in semantic navigation have demonstrated open-set object-goal navigation in indoor settings, but these methods remain limited by constrained spatial ranges and str…

2026

RobotArena $\infty$: Unlimited Robot Benchmarking via Real-to-Sim Translation

ICLR 2026poster

The pursuit of robot generalists—instructable agents capable of performing diverse tasks across diverse environments—demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to re…

Cited by 0SourcecodeScholar
2025

CASPER: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models

CoRL 2025poster

Assistive teleoperation, where control is shared between a human and a robot, enables efficient and intuitive human-robot collaboration in diverse and unstructured environments. A central challenge in real-world assistive teleoperation is for the robot to infer a wide range of human intentions from…

Cited by 0SourcecodeScholar
2025

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

NAACL 2025long

Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for open-ended conversations,…

2025

Energy Considerations of Large Language Model Inference and Efficiency Optimizations

ACL 2025long

As large language models (LLMs) scale in size and adoption, their computational and environmental costs continue to rise. Prior benchmarking efforts have primarily focused on latency reduction in idealized settings, often overlooking the diverse real-world inference workloads that shape energy use.…

Cited by 0SourcePDFScholar
2025

Explore Theory of Mind: program-guided adversarial data generation for theory of mind reasoning

ICLR 2025poster

Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence. However, all rely on limited datasets with simple patterns that can potentially lead to probl…

Cited by 5SourcePDFScholar
2025

MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction

EMNLP 2025

Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors. This raises important concerns about their robustness and reliability in scientific applications. To suppo

2025

Re-thinking Temporal Search for Long-Form Video Understanding

CVPR 2025poster

Efficient understanding of long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding, studying a fundamental issue pertaining to all state-of-the-art (SOTA) long-context vision-language models (VLMs). In…

2024

ANAVI: Audio Noise Awareness using Visual of Indoor environments for NAVIgation

CoRL 2024poster

We propose Audio Noise Awareness using Visuals of Indoors for NAVIgation for quieter robot path planning. While humans are naturally aware of the noise they make and its impact on those around them, robots currently lack this awareness. A key challenge in achieving audio awareness for robots is es…

Cited by 0SourcecodeScholar
2024

Diffusion PID: Interpreting Diffusion via Partial Information Decomposition

NeurIPS 2024poster

Text-to-image diffusion models have made significant progress in generating naturalistic images from textual inputs, and demonstrate the capacity to learn and represent complex visual-semantic relationships. While these diffusion models have achieved remarkable success, the underlying mechanisms dri…

2024

Gradient Localization Improves Lifelong Pretraining of Language Models

EMNLP 2024finding

Large Language Models (LLMs) trained on web-scale text corpora have been shown to capture world knowledge in their parameters. However, the mechanism by which language models store different types of knowledge is poorly understood. In this work, we examine two types of knowledge relating to temporal…

Cited by 0SourcePDFScholar
2024

How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models

EMNLP 2024finding

Given the growing influx of misinformation across news and social media, there is a critical need for systems that can provide effective real-time verification of news claims. Large language or multimodal model based verification has been proposed to scale up online policing mechanisms for mitigatin…

Cited by 0SourcePDFScholar
2024

MOSAIC: Learning Unified Multi-Sensory Object Property Representations for Robot Learning via Interactive Perception

ICRA 2024poster

A holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science studies that emphasize the significance of multi-sensory int…

Cited by 2SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

OpenEQA: Embodied Question Answering in the Era of Foundation Models

CVPR 2024poster

We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory exemplified by agents on smart glasses or b…

Cited by 118SourcePDFScholar
2024

SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents

ACL 2024long

Humans learn social skills through both imitation and social interaction. This social learning process is largely understudied by existing research on building language agents. Motivated by this gap, we propose an interactive learning method, SOTOPIA-π, that improves the social intelligence of langu…

2024

SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

ICLR 2024spotlight

*Humans are social beings*; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and…

Cited by 148SourcePDFScholar
2024

Situated Instruction Following

ECCV 2024poster

"Language is never spoken in a vacuum. It is expressed, comprehended, and contextualized within the holistic backdrop of the speaker’s history, actions, and environment. Since humans are used to communicating efficiently with situated language, the practicality of robotic assistants hinge on their a…

2024

Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation

ECCV 2024poster

"The literature on text-to-image generation is plagued by issues of faithfully composing entities with relations. But there lacks a formal understanding of how entity-relation compositions can be effectively learned. Moreover, the underlying phenomenon space that meaningfully reflects the problem st…

2024

VISREAS: Complex Visual Reasoning with Unanswerable Questions

ACL 2024findings

Verifying a question’s validity before answering is crucial in real-world applications, where users may provide imperfect instructions. In this scenario, an ideal model should address the discrepancies in the query and convey them to the users rather than generating the best possible answer. Address…

Cited by 3SourcePDFScholar
2024

Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

RSS 2024poster

Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the robot must understand a person's intent and perform the inferred task despite differences in the embodiments and environme…

2024

WebArena: A Realistic Web Environment for Building Autonomous Agents

ICLR 2024poster

With advances in generative AI, there is now potential for autonomous agents to manage daily tasks via natural language commands. However, current agents are primarily created and tested in simplified synthetic environments, leading to a disconnect with real-world scenarios. In this paper, we build…

2023

Computational Language Acquisition with Theory of Mind

ICLR 2023poster

Unlike current state-of-the-art language models, young children actively acquire language through interactions with their surrounding environment and caretakers. One mechanism that has been argued to be critical to language learning is the ability to infer the mental states of other agents in social…

2023

EXCALIBUR: Encouraging and Evaluating Embodied Exploration

CVPR 2023poster

Experience precedes understanding. Humans constantly explore and learn about their environment out of curiosity, gather information, and update their models of the world. On the other hand, machines are either trained to learn passively from static and fixed datasets, or taught to complete specific…

Cited by 17SourcePDFScholar
2023

HomeRobot: Open-Vocabulary Mobile Manipulation

CoRL 2023poster

HomeRobot (noun): An affordable compliant robot that navigates homes and manipulates a wide range of objects in order to complete everyday tasks. Open-Vocabulary Mobile Manipulation (OVMM) is the problem of picking any object in any unseen environment, and placing it in a commanded location. This i…

Cited by 98SourcecodeScholar
2023

SLAP: Spatial-Language Attention Policies

CoRL 2023poster

Despite great strides in language-guided manipulation, existing work has been constrained to table-top settings. Table-tops allow for perfect and consistent camera angles, properties are that do not hold in mobile manipulation. Task plans that involve moving around the environment must be robust to…

Cited by 8SourcecodeScholar
2023

SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs

NeurIPS 2023spotlight

In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM…

Cited by 59SourcePDFScholar
2023

SPRING: Studying Papers and Reasoning to play Games

NeurIPS 2023poster

Open-world survival games pose significant challenges for AI algorithms due to their multi-tasking, deep exploration, and goal prioritization requirements. Despite reinforcement learning (RL) being popular for solving games, its high sample complexity limits its effectiveness in complex open-world g…

2023

Self-Supervised Object Goal Navigation with In-Situ Finetuning

IROS 2023poster

A household robot should be able to navigate to target objects without requiring users to first annotate everything in their home. Most current approaches to object navigation do not test on real robots and rely solely on reconstructed scans of houses and their expensively labeled semantic 3D meshes…

Cited by 7SourceScholar
2023

The Framework Tax: Disparities Between Inference Efficiency in NLP Research and Deployment

EMNLP 2023long main

Increased focus on the computational efficiency of systems in natural language processing has motivated the design of efficient model architectures and improvements to underlying hardware accelerators. However, the resulting increases in computational throughput and reductions in floating point ope…

Cited by 0SourcecodeScholar
2022

A Framework for Learning to Request Rich and Contextually Useful Information from Humans

ICML 2022spotlight

When deployed, AI agents will encounter problems that are beyond their autonomous problem-solving capabilities. Leveraging human assistance can help agents overcome their inherent limitations and robustly cope with unfamiliar situations. We present a general interactive framework that enables an age…

2022

Don’t Copy the Teacher: Data and Model Challenges in Embodied Dialogue

EMNLP 2022main

Embodied dialogue instruction following requires an agent to complete a complex sequence of tasks from a natural language exchange. The recent introduction of benchmarks raises the question of how best to train and evaluate models for this multi-turn, multi-agent, long-horizon task. This paper contr…

2022

EvEntS ReaLM: Event Reasoning of Entity States via Language Models

EMNLP 2022main

This paper investigates models of event implications. Specifically, how well models predict entity state-changes, by targeting their understanding of physical attributes. Nominally, Large Language models (LLM) have been exposed to procedural knowledge about how objects interact, yet our benchmarking…

2022

FILM: Following Instructions in Language with Modular Methods

ICLR 2022poster

Recent methods for embodied instruction following are typically trained end-to-end using imitation learning. This often requires the use of expert trajectories and low-level language instructions. Such approaches assume that neural states will integrate multimodal semantics to perform state tracking…

2022

KAT: A Knowledge Augmented Transformer for Vision-and-Language

NAACL 2022long

The primary focus of recent work with large-scale transformers has been on optimizing the amount of information packed into the model’s parameters. In this work, we ask a complementary question: Can multimodal transformers leverage explicit knowledge in their reasoning? Existing, primarily unimodal,…

2022

On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization

EMNLP 2022finding

Combining the visual modality with pretrained language models has been surprisingly effective for simple descriptive tasks such as image captioning. More general text generation however remains elusive. We take a step back and ask: How do these models work for more complex generative tasks, i.e. con…

Cited by 3SourcePDFScholar
2021

ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

ICLR 2021poster

Given a simple request like Put a washed apple in the kitchen fridge, humans can reason in purely abstract terms by imagining action sequences and scoring their likelihood of success, prototypicality, and efficiency, all without moving a muscle. Once we see the kitchen in question, we can update our…

2021

Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question Answering

AAAI 2021technical

Recent developments in pre-trained neural language modeling have led to leaps in accuracy on common-sense question-answering benchmarks. However, there is increasing concern that models overfit to specific tasks, without learning to utilize external knowledge or perform general semantic reasoning.…

2020

ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks

CVPR 2020poster

We present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks. ALFRED includes long, compositional tasks with non-reversible state changes to shri…

Cited by 922PDFcodeScholar
2020

Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games

COLING 2020main

In visual guessing games, a Guesser has to identify a target object in a scene by asking questions to an Oracle. An effective strategy for the players is to learn conceptual representations of objects that are both discriminative and expressive enough to ask questions and guess correctly. However, a…

2019

Defending Against Neural Fake News

NeurIPS 2019poster

Recent progress in natural language generation has raised dual-use concerns. While applications like summarization and translation are positive, the underlying technology also might enable adversaries to generate neural fake news: targeted propaganda that closely mimics the style of real news.

2019

EARLY FUSION for Goal Directed Robotic Vision

IROS 2019poster

Building perceptual systems for robotics which perform well under tight computational budgets requires novel architectures which rethink the traditional computer vision pipeline. Modern vision architectures require the agent to build a summary representation of the entire scene, even if most of the…

Cited by 10SourceScholar
2019

Improving Robot Success Detection using Static Object Data

IROS 2019poster

We use static object data to improve success detection for stacking objects on and nesting objects in one another. Such actions are necessary for certain robotics tasks, e.g., clearing a dining table or packing a warehouse bin. However, using an RGB-D camera to detect success can be insufficient: sa…

Cited by 14SourcecodeScholar
2019

Prospection: Interpretable plans from language by predicting the future

ICRA 2019poster

High-level human instructions often correspond to behaviors with multiple implicit steps. In order for robots to be useful in the real world, they must be able to to reason over both motions and intermediate goals implied by human instructions. In this work, we propose a framework for learning repre…

Cited by 58SourceScholar
2019

Tactical Rewind: Self-Correction via Backtracking in Vision-And-Language Navigation

CVPR 2019oral

We present the Frontier Aware Search with backTracking (FAST) Navigator, a general framework for action decoding, that achieves state-of-the-art results on the 2018 Room-to-Room (R2R) Vision-and-Language navigation challenge. Given a natural language instruction and photo-realistic image views of a…

Cited by 194PDFcodeScholar