← Search

Rogerio Bonatti

16 accepted papers

2025

VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

ICLR 2025poster

Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text or static imagery can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focus- ing on text or static image inputs. To bridge this ga…

Cited by 3SourcePDFScholar
2025

Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

ICML 2025poster

Large language models (LLMs) show potential as computer agents, enhancing productivity and software accessibility in multi-modal tasks. However, measuring agent performance in sufficiently realistic and complex environments becomes increasingly challenging as: (i) most benchmarks are limited to sp…

2024

ConBaT: Control Barrier Transformer for Safe Robot Learning from Demonstrations

ICRA 2024poster

Large-scale self-supervised models have recently revolutionized our ability to perform a variety of tasks within the vision and language domains. However, using such models for autonomous systems is challenging because of safety requirements: besides executing correct actions, an autonomous agent mu…

Cited by 1SourceScholar
2023

Is Imitation All You Need? Generalized Decision-Making with Dual-Phase Training

ICCV 2023poster

We introduce DualMind, a generalist agent designed to tackle various decision-making tasks that addresses challenges posed by current methods, such as overfitting behaviors and dependence on task-specific fine-tuning. DualMind uses a novel "Dual-phase" training strategy that emulates how humans lear…

Cited by 17PDFcodeScholar
2023

LATTE: LAnguage Trajectory TransformEr

ICRA 2023poster

Natural language is one of the most intuitive ways to express human intent. However, translating instructions and commands towards robotic motion generation and deployment in the real world is far from being an easy task. The challenge of combining a robot's inherent low-level geometric and kinodyna…

Cited by 79SourcecodeScholar
2023

PACT: Perception-Action Causal Transformer for Autoregressive Robotics Pre-Training

IROS 2023poster

Robotics has long been a field riddled with complex systems architectures whose modules and connections, whether traditional or learning-based, require significant human expertise and prior knowledge. Inspired by large pre-trained language models, this work introduces a paradigm for pretraining a ge…

Cited by 20SourceScholar
2023

SMART: Self-supervised Multi-task pretrAining with contRol Transformers

ICLR 2023top-25%

Self-supervised pretraining has been extensively studied in language and vision domains, where a unified model can be easily adapted to various downstream tasks by pretraining representations without explicit labels. When it comes to sequential decision-making tasks, however, it is difficult to prop…

2022

Reshaping Robot Trajectories Using Natural Language Commands: A Study of Multi-Modal Data Alignment Using Transformers

IROS 2022poster

Natural language is the most intuitive medium for us to interact with other people when expressing commands and instructions. However, using language is seldom an easy task when humans need to express their intent towards robots, since most of the current language interfaces require rigid templates…

Cited by 60SourcecodeScholar
2021

3D Human Reconstruction in the Wild with Collaborative Aerial Cameras

IROS 2021poster

Aerial vehicles are revolutionizing applications that require capturing the 3D structure of dynamic targets in the wild, such as sports, medicine and entertainment. The core challenges in developing a motion-capture system that operates in outdoors environments are: (1) 3D inference requires multipl…

Cited by 23SourceScholar
2021

Batteries, camera, action! Learning a semantic control space for expressive robot cinematography

ICRA 2021poster

Aerial vehicles are revolutionizing the way filmmakers can capture shots of actors by composing novel aerial and dynamic viewpoints. However, despite great advancements in autonomous flight technology, generating expressive camera behaviors is still a challenge and requires non-technical users to ed…

Cited by 24SourceScholar
2021

Do You See What I See? Coordinating Multiple Aerial Cameras for Robot Cinematography

ICRA 2021poster

Aerial cinematography is significantly expanding the capabilities of film-makers. Recent progress in autonomous unmanned aerial vehicles (UAVs) has further increased the potential impact of aerial cameras, with systems that can safely track actors in unstructured cluttered environments. Professional…

Cited by 29SourceScholar
2020

Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal Representations

IROS 2020poster

Machines are a long way from robustly solving open-world perception-control tasks, such as first-person view (FPV) aerial navigation. While recent advances in end-to- end Machine Learning, especially Imitation Learning and Reinforcement appear promising, they are constrained by the need of large amo…

Cited by 61SourcecodeScholar
2019

Can a Robot Become a Movie Director? Learning Artistic Principles for Aerial Cinematography

IROS 2019poster

Aerial filming is constantly gaining importance due to the recent advances in drone technology. It invites many intriguing, unsolved problems at the intersection of aesthetical and scientific challenges. In this work, we propose a deep reinforcement learning agent which supervises motion planning of…

Cited by 72SourceScholar
2019

Improved Generalization of Heading Direction Estimation for Aerial Filming Using Semi-Supervised Regression

ICRA 2019poster

In the task of Autonomous aerial filming of a moving actor (e.g. a person or a vehicle), it is crucial to have a good heading direction estimation for the actor from the visual input. However, the models obtained in other similar tasks, such as pedestrian collision risk analysis and human-robot inte…

Cited by 8SourceScholar
2019

Towards a Robust Aerial Cinematography Platform: Localizing and Tracking Moving Targets in Unstructured Environments

IROS 2019poster

The use of drones for aerial cinematography has revolutionized several applications and industries that require live and dynamic camera viewpoints such as entertainment, sports, and security. However, safely controlling a drone while filming a moving target usually requires multiple expert human ope…

Cited by 108SourceScholar
2018

Integrating kinematics and environment context into deep inverse reinforcement learning for predicting off-road vehicle trajectories

CoRL 2018

Predicting the motion of a mobile agent from a third-person perspective is an important component for many robotics applications, such as autonomous navigation and tracking. With accurate motion prediction of other agents, robots can plan for more intelligent behaviors to achieve specified objective