← Search

Guy Rosman

48 accepted papers

2026

Probing Multimodal LLMs As World Models for Driving

ICRA 2026poster

We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o, their performance in complex driving environments remains l…

2025

Computational Teaching for Driving via Multi-Task Imitation Learning

ICRA 2025

Learning motor skills for sports or performance driving is often done with professional instruction from expert human teachers, whose availability is limited. Our goal is to enable automated teaching via a learned model that interacts with the student similar to a human teacher. However, training su

Cited by 4SourceScholar
2025

Estimating cognitive biases with attention-aware inverse planning

NeurIPS 2025spotlight

People's goal-directed behaviors are influenced by their cognitive biases, and autonomous systems that interact with people should be aware of this. For example, people's attention to objects in their environment will be biased in a way that systematically affects how they perform everyday tasks suc…

Cited by 0SourceScholar
2025

Generating Out-of-Distribution Scenarios Using Language Models

ICRA 2025

The deployment of autonomous vehicles controlled by machine learning techniques requires extensive testing in diverse real-world environments, robust handling of edge cases and out-of-distribution scenarios, and comprehensive safety validation to ensure that these systems can navigate safely and eff

Cited by 10SourceScholar
2025

Hypergraph-Transformer (HGT) for Interaction Event Prediction in Laparoscopic and Robotic Surgery

ICRA 2025

Understanding and anticipating events and actions is critical for intraoperative assistance and decision-making during minimally invasive surgery. We propose a predictive neural network that is capable of understanding and predicting critical interaction aspects of surgical workflow based on endosco

Cited by 6SourceScholar
2025

Probing Multimodal LLMs as World Models for Driving

RA-L 2025

We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o, their performance in complex driving environments remains l

Cited by 21SourceScholar
2025

ReGen: Generative Robot Simulation via Inverse Design

ICLR 2025poster

Simulation plays a key role in scaling robot learning and validating policies, but constructing simulations remains labor-intensive. In this paper, we introduce ReGen, a generative simulation framework that automates this process using inverse design. Given an agent's behavior (such as a motion traj…

Cited by 0SourcePDFScholar
2025

Safety with Agency: Human-Centered Safety Filter with Application to AI-Assisted Motorsports

RSS 2025poster

Recent advances in safe autonomy open new opportunities in assisting humans in safety-critical and time-sensitive tasks such as motorsports. However, existing safe control algorithms predominantly focus on fully automated settings and often undermine key requirements in human–AI shared control domai…

Cited by 0PDFScholar
2025

Think Deep and Fast: Learning Neural Nonlinear Opinion Dynamics from Inverse Dynamic Games for Split-Second Interactions

ICRA 2025

Non-cooperative interactions commonly occur in multi-agent scenarios such as car racing, where an ego vehicle can choose to overtake the rival, or stay behind it until a safe overtaking “corridor” opens. While an expert human can do well at making such time-sensitive decisions, autonomous agents are

Cited by 9SourceScholar
2024

Blending Data-Driven Priors in Dynamic Games

RSS 2024poster

As intelligent robots like autonomous vehicles become increasingly deployed in the presence of people, the extent to which these systems should leverage model-based game-theoretic planners versus data-driven policies for safe, interaction-aware motion planning remains an open question. Existing dyna…

2024

Dreaming to Assist: Learning to Align with Human Objectives for Shared Control in High-Speed Racing

CoRL 2024poster

Tight coordination is required for effective human-robot teams in domains involving fast dynamics and tactical decisions, such as multi-car racing. In such settings, robot teammates must react to cues of a human teammate's tactical objective to assist in a way that is consistent with the objective…

Cited by 2SourceScholar
2024

Drive Anywhere: Generalizable End-to-end Autonomous Driving with Multi-modal Foundation Models

ICRA 2024poster

As autonomous driving technology matures, end-to-end methodologies have emerged as a leading strategy, promising seamless integration from perception to control via deep learning. However, existing systems grapple with challenges such as unexpected open set environments and the complexity of black-b…

Cited by 31SourceScholar
2024

Online Adaptation of Learned Vehicle Dynamics Model with Meta-Learning Approach

IROS 2024poster

We represent a vehicle dynamics model for autonomous driving near the limits of handling via a multilayer neural network. Online adaptation is desirable in order to address unseen environments. However, the model needs to adapt to new environments without forgetting previously encountered ones. In t…

Cited by 2SourceScholar
2023

Dynamic Multi-Team Racing: Competitive Driving on 1/10-th Scale Vehicles via Learning in Simulation

CoRL 2023poster

Autonomous racing is a challenging task that requires vehicle handling at the dynamic limits of friction. While single-agent scenarios like Time Trials are solved competitively with classical model-based or model-free feedback control, multi-agent wheel-to-wheel racing poses several challenges inclu…

Cited by 6SourceScholar
2023

MPOGames: Efficient Multimodal Partially Observable Dynamic Games

ICRA 2023poster

Game theoretic methods have become popular for planning and prediction in situations involving rich multi-agent interactions. However, these methods often assume the existence of a single local Nash equilibria and are hence unable to handle uncertainty in the intentions of different agents. While ma…

Cited by 12SourceScholar
2023

Multi-Abstractive Neural Controller: An Efficient Hierarchical Control Architecture for Interactive Driving

RA-L 2023

As learning-based methods make their way from perception systems to planning/control stacks, robot control systems have started to enjoy the benefits that data-driven methods provide. Because control systems directly affect the motion of the robot, data-driven methods, especially black box approache

Cited by 1SourceScholar
2022

A Deep Concept Graph Network for Interaction-Aware Trajectory Prediction

ICRA 2022poster

Temporal patterns (how vehicles behave in our observed past) underline our reasoning of how people drive on the road, and can explain why we make certain predictions about interactions among road agents. In this paper we propose the ConceptNet trajectory predictor - a novel prediction framework that…

Cited by 12SourceScholar
2022

HYPER: Learned Hybrid Trajectory Prediction via Factored Inference and Adaptive Sampling

ICRA 2022poster

Modeling multi-modal high-level intent is important for ensuring diversity in trajectory prediction. Existing approaches explore the discrete nature of human intent before predicting continuous trajectories, to improve accuracy and support explainability. However, these approaches often assume the i…

Cited by 33SourceScholar
2022

Learning an Explainable Trajectory Generator Using the Automaton Generative Network (AGN)

RA-L 2022

Symbolic reasoning is a key component for enabling practical use of data-driven planners in autonomous driving. In that context, deterministic finite state automata (DFA) are often used to formalize the underlying high-level decision-making process. Manual design of an effective DFA can be tedious.

Cited by 5SourceScholar
2022

Leveraging Smooth Attention Prior for Multi-Agent Trajectory Prediction

ICRA 2022poster

Multi-agent interactions are important to model for forecasting other agents' behaviors and trajectories. At a certain time, to forecast a reasonable future trajectory, each agent needs to pay attention to the interactions with only a small group of most relevant agents instead of unnecessarily payi…

Cited by 11SourceScholar
2022

SUPR-GAN: SUrgical PRediction GAN for Event Anticipation in Laparoscopic and Robotic Surgery

RA-L 2022

Comprehension of surgical workflow is the foundation upon which artificial intelligence (AI) and machine learning (ML) holds the potential to assist intraoperative decision making and risk mitigation. In this work, we move beyond mere identification of past surgical phases, into prediction of future

Cited by 19SourceScholar
2022

TIP: Task-Informed Motion Prediction for Intelligent Vehicles

IROS 2022poster

When predicting trajectories of road agents, motion predictors often approximate the future distribution by a limited number of samples. This constraint requires the predictors to generate samples that best support the task given task specifications. However, existing predictors are often optimized…

Cited by 15SourceScholar
2022

Trajectory Prediction with Linguistic Representations

ICRA 2022poster

Language allows humans to build mental models that interpret what is happening around them resulting in more accurate long-term predictions. We present a novel trajectory prediction model that uses linguistic intermediate representations to forecast trajectories, and is trained using trajectory samp…

Cited by 22SourceScholar
2021

Aggregating Long-Term Context for Learning Laparoscopic and Robot-Assisted Surgical Workflows

ICRA 2021poster

Analyzing surgical workflow is crucial for surgical assistance robots to understand surgeries. With the understanding of the complete surgical workflow, the robots are able to assist the surgeons in intra-operative events, such as by giving a warning when the surgeon is entering specific keys or hig…

Cited by 22SourceScholar
2021

CARPAL: Confidence-Aware Intent Recognition for Parallel Autonomy

RA-L 2021

Predicting driver intentions is a difficult and crucial task for advanced driver assistance systems. Traditional confidence measures on predictions often ignore the way predicted trajectories affect downstream decisions for safe driving. In this letter, we propose a novel multi-task intent recogniti

Cited by 7SourceScholar
2021

Vehicle Trajectory Prediction Using Generative Adversarial Network With Temporal Logic Syntax Tree Features

RA-L 2021

In this work, we propose a novel approach for integrating rules into traffic agent trajectory prediction. Consideration of rules is important for understanding how people behave-yet, it cannot be assumed that rules are always followed. To address this challenge, we evaluate different approaches of i

Cited by 53SourceScholar
2020

Behaviorally Diverse Traffic Simulation via Reinforcement Learning

IROS 2020poster

Traffic simulators are important tools in autonomous driving development. While continuous progress has been made to provide developers more options for modeling various traffic participants, tuning these models to increase their behavioral diversity while maintaining quality is often very challengi…

Cited by 0SourceScholar
2020

Deep Context Maps: Agent Trajectory Prediction Using Location-Specific Latent Maps

RA-L 2020

In this letter, we propose a novel approach for agent motion prediction in cluttered environments. One of the main challenges in predicting agent motion is accounting for location and context-specific information. Our main contribution is the concept of learning context maps to improve the predictio

Cited by 8SourceScholar
2020

Differentiable Logic Layer for Rule Guided Trajectory Prediction

CoRL 2020

In this work, we propose a method for integration of temporal logic formulas into a neural network. Our main contribution is a new logic optimization layer that uses differentiable optimization on the formulas’ robustness function. This allows incorporating traffic rules into deep learning based tra

Cited by 0SourcePDFScholar
2020

DiversityGAN: Diversity-Aware Vehicle Motion Prediction via Latent Semantic Sampling

RA-L 2020

Vehicle trajectory prediction is crucial for autonomous driving and advanced driver assistant systems. While existing approaches may sample from a predicted distribution of vehicle trajectories, they lack the ability to explore it - a key ability for evaluating safety from a planning and verificatio

Cited by 80SourceScholar
2020

Driving Through Ghosts: Behavioral Cloning with False Positives

IROS 2020poster

Safe autonomous driving requires robust detection of other traffic participants. However, robust does not mean perfect, and safe systems typically minimize missed detections at the expense of a higher false positive rate. This results in conservative and yet potentially dangerous behavior such as av…

Cited by 24SourceScholar
2020

MATS: An Interpretable Trajectory Forecasting Representation for Planning and Control

CoRL 2020

Reasoning about human motion is a core component of modern human-robot interactive systems. In particular, one of the main uses of behavior prediction in autonomous systems is to inform robot motion planning and control. However, a majority of planning and control algorithms reason about system dyna

2020

Reinforcement Learning based Control of Imitative Policies for Near-Accident Driving

RSS 2020poster

Autonomous driving has achieved significant progress in recent years, but autonomous cars are still unable to tackle high-risk situations where a potential accident is likely. In such near-accident scenarios, even a minor change in the vehicle's actions may result in drastically different consequenc…

2019

Infrastructure-free NLoS Obstacle Detection for Autonomous Cars

IROS 2019poster

Current perception systems mostly require direct line of sight to anticipate and ultimately prevent potential collisions at intersections with other road users. We present a fully integrated autonomous system capable of detecting shadows or weak illumination changes on the ground caused by a dynamic…

Cited by 16SourceScholar
2019

Probabilistic Risk Metrics for Navigating Occluded Intersections

RA-L 2019

Among traffic accidents in the USA, 23% of fatal and 32% of non-fatal incidents occurred at intersections. For driver assistance systems, intersection navigation remains a difficult problem that is critically important to increasing driver safety. In this letter, we examine how to navigate an unsign

Cited by 36SourceScholar
2019

Uncertainty-Aware Driver Trajectory Prediction at Urban Intersections

ICRA 2019poster

Predicting the motion of a driver’s vehicle is crucial for advanced driving systems, enabling detection of potential risks towards shared control between the driver and automation systems. In this paper, we propose a variational neural network approach that predicts future driver trajectory distribu…

Cited by 104SourceScholar
2018

Task-Specific Sensor Planning for Robotic Assembly Tasks

ICRA 2018poster

When performing multi-robot tasks, sensory feedback is crucial in reducing uncertainty for correct execution. Yet the utilization of sensors should be planned as an integral part of the task planning, taken into account several factors such as the tolerance of different inferred properties of the sc…

Cited by 13SourceScholar
2018

Variational Autoencoder for End-to-End Control of Autonomous Driving with Novelty Detection and Training De-biasing

IROS 2018poster

This paper introduces a new method for end-to-end training of deep neural networks (DNNs) and evaluates it in the context of autonomous driving. DNN training has been shown to result in high accuracy for perception to action learning given sufficient training data. However, the trained models may fa…

Cited by 108SourceScholar
2017

Duckietown: An open, inexpensive and flexible platform for autonomy education and research

ICRA 2017poster

Duckietown is an open, inexpensive and flexible platform for autonomy education and research. The platform comprises small autonomous vehicles (“Duckiebots”) built from off-the-shelf components, and cities (“Duckietowns”) complete with roads, signage, traffic lights, obstacles, and citizens (duckies…

Cited by 281SourceScholar
2017

Machine learning and coresets for automated real-time video segmentation of laparoscopic and robot-assisted surgery

ICRA 2017poster

Context-aware segmentation of laparoscopic and robot assisted surgical video has been shown to improve performance and perioperative workflow efficiency, and can be used for education and time-critical consultation. Modern pressures on productivity preclude manual video analysis, and hospital polici…

Cited by 80SourceScholar
2017

Persistent surveillance of events with unknown, time-varying statistics

ICRA 2017poster

We consider the problem of monitoring stochastic, time-varying events occurring at discrete locations. Our problem formulation extends prior work in persistent surveillance by considering the objective of maximizing event detections in unknown, dynamic environments where the rates of events are time…

Cited by 11SourceScholar
2016

Real-Time Depth Refinement for Specular Objects

CVPR 2016poster

The introduction of consumer RGB-D scanners set off a major boost in 3D computer vision research. Yet, the precision of existing depth scanners is not accurate enough to recover fine details of a scanned object. While modern shading based depth refinement methods have been proven to work well with L…

Cited by 29PDFScholar
2015

Coresets for visual summarization with applications to loop closure

ICRA 2015poster

In continuously operating robotic systems, efficient representation of the previously seen camera feed is crucial. Using a highly efficient compression coreset method, we formulate a new method for hierarchical retrieval of frames from large video streams collected online by a moving robot. We demon…

Cited by 31SourceScholar
2015

RGBD-Fusion: Real-Time High Precision Depth Recovery

CVPR 2015poster

The popularity of low-cost RGB-D scanners is increasing on a daily basis. Nevertheless, existing scanners often cannot capture subtle details in the environment. We present a novel method to enhance the depth map by fusing the intensity and depth information to create more detailed range profiles. T…

Cited by 134SourcePDFScholar