← Search

Matthew Johnson-Roberson

63 accepted papers

2026

Bi-Manual Joint Camera Calibration and Scene Representation

ICRA 2026poster

Robot manipulation, especially bimanual manipulation, often requires setting up multiple cameras on multiple robot manipulators. Before robot manipulators can generate motion or even build representations of their environments, the cameras rigidly mounted to the robot need to be calibrated. Camera c…

2026

Cross-Modal Instructions for Robot Motion Generation

ICRA 2026poster

Teaching robots novel behaviors typically requires motion demonstrations via teleoperation or kinaesthetic teaching, that is, physically guiding the robot. While recent work has explored using human sketches to specify desired behaviors, data collection remains cumbersome, and demonstration datasets…

2026

DOSE3: Diffusion-Based Unified Out-Of-Distribution Detection on SE(3) Trajectories

ICRA 2026poster

Out-Of-Distribution (OOD) detection, the task of identifying when an input falls outside the distribution seen at training time, is critical for deploying safe and reliable systems. Traditional OOD methods require retraining models whenever the in‐distribution has changed. Recent work introduces uni…

Cited by 0Scholar
2026

DOSE3: Diffusion-Based Unified Out-of-Distribution Detection on $\mathbb{SE}(3)$ Trajectories

RA-L 2026

Out-of-Distribution (OOD) detection, the task of identifying when an input falls outside the distribution seen at training time, is critical for deploying safe and reliable systems. Traditional OOD methods require retraining models whenever the in-distribution has changed. Recent work introduces <it

Cited by 1SourceScholar
2026

DreamSea: Photorealistic 3D Underwater Terrain Generation by Latent Fractal Diffusion Models

ICRA 2026poster

This paper tackles the problem of generating representations of underwater 3D terrain. Off-the-shelf generative models, trained on Internet-scale data but not on specialized underwater images, exhibit downgraded realism, as images of the seafloor are relatively uncommon. To this end, we introduce Dr…

Cited by 0Scholar
2026

Efficient Construction of Implicit Surface Models from a Single Image for Motion Generation

ICRA 2026poster

Implicit representations have been widely applied in robotics for obstacle avoidance and path planning. In this paper, we explore the problem of constructing an implicit distance representation from a single image. Past methods for implicit surface reconstruction, such as NeuS and its variants gener…

2026

Joint Flow Trajectory Optimization for Feasible Robot Motion Generation from Video Demonstrations

ICRA 2026poster

Learning from human video demonstrations offers a scalable alternative to teleoperation or kinesthetic teaching, but poses challenges for robot manipulators due to embodiment differences and joint feasibility constraints. We address this problem by proposing the Joint Flow Trajectory Optimization (J…

2026

Robust Bayesian Scene Reconstruction With Retrieval-Augmented Priors for Precise Grasping and Planning

RA-L 2026

Constructing 3D representations of object geometry is critical for many robotics tasks, particularly manipulation problems. These representations must be built from potentially noisy partial observations. In this work, we focus on the problem of reconstructing a multi-object scene from a single RGBD

Cited by 2SourceScholar
2026

Robust Bayesian Scene Reconstruction with Retrieval-Augmented Priors for Precise Grasping and Planning

ICRA 2026poster

Constructing 3D representations of object geometry is critical for many robotics tasks, particularly manipulation problems. These representations must be built from potentially noisy partial observations. In this work, we focus on the problem of reconstructing a multi-object scene from a single RGBD…

2025

Building 3D Representations and Generating Motions From a Single Image via Video-Generation

NeurIPS 2025poster

Autonomous robots typically need to construct representations of their surroundings and adapt their motions to the geometry of their environment. Here, we tackle the problem of constructing a policy model for collision-free motion generation, consistent with the environment, from a single input RGB…

Cited by 0SourceScholar
2025

RecGS: Removing Water Caustic With Recurrent Gaussian Splatting

RA-L 2025

Water caustics are commonly observed in seafloor imaging data from shallow-water areas. Traditional methods that remove caustic patterns from images often rely on 2D filtering or pre-training on an annotated dataset, hindering the performance when generalizing to real-world seafloor data with 3D str

Cited by 17SourceScholar
2025

Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video

ICLR 2025poster

We aim to redefine robust ego-motion estimation and photorealistic 3D reconstruction by addressing a critical limitation: the reliance on noise-free data in existing models. While such sanitized conditions simplify evaluation, they fail to capture the unpredictable, noisy complexities of real-world…

2024

DarkGS: Learning Neural Illumination and 3D Gaussians Relighting for Robotic Exploration in the Dark

IROS 2024poster

Humans have the remarkable ability to construct consistent mental models of an environment, even under limited or varying levels of illumination. We wish to endow robots with this same capability. In this paper, we tackle the challenge of constructing a photorealistic scene representation under poor…

Cited by 21SourcecodeScholar
2024

Instructing Robots by Sketching: Learning from Demonstration via Probabilistic Diagrammatic Teaching

ICRA 2024poster

Learning from Demonstration (LfD) enables robots to acquire new skills by imitating expert demonstrations, allowing users to communicate their instructions intuitively. Recent progress in LfD often relies on kinesthetic teaching or teleoperation as the medium for users to specify the demonstrations.…

Cited by 10SourceScholar
2024

Simultaneous Geometry and Pose Estimation of Held Objects Via 3D Foundation Models

RA-L 2024

Humans have the remarkable ability to use held objects as tools to interact with their environment. Humans internally estimate how hand movements affect the object's movement. We wish to endow robots with this capability. We contribute methodology to jointly estimate the geometry and pose of objects

Cited by 9SourceScholar
2024

Teaching Robots Where To Go And How To Act With Human Sketches via Spatial Diagrammatic Instructions

IROS 2024poster

This paper introduces Spatial Diagrammatic Instructions (SDIs), an approach for human operators to specify objectives and constraints that are related to spatial regions in the working environment. Human operators are enabled to sketch out regions directly on camera images that correspond to the obj…

Cited by 0SourceScholar
2024

Unifying Representation and Calibration With 3D Foundation Models

RA-L 2024

Representing the environment is a central challenge in robotics, and is essential for effective decision-making. Traditionally, before capturing images with a manipulator-mounted camera, users need to calibrate the camera using a specific external marker, such as a checkerboard or AprilTag. However,

Cited by 12SourceScholar
2024

V-PRISM: Probabilistic Mapping of Unknown Tabletop Scenes

IROS 2024poster

The ability to construct concise scene representations from sensor input is central to the field of robotics. This paper addresses the problem of robustly creating a 3D representation of a tabletop scene from a segmented RGBD image. These representations are then critical for a range of downstream m…

Cited by 8SourcecodeScholar
2023

Beyond NeRF Underwater: Learning Neural Reflectance Fields for True Color Correction of Marine Imagery

RA-L 2023

Underwater imagery often exhibits distorted coloration as a result of light-water interactions, which complicates the study of benthic environments in marine biology and geography. In this research, we propose an algorithm to restore the true color (albedo) in underwater imagery by jointly learning

Cited by 38SourcecodeScholar
2023

CLONeR: Camera-Lidar Fusion for Occupancy Grid-Aided Neural Representations

RA-L 2023

Recent advances in neural radiance fields (NeRFs) achieve state-of-the-art novel view synthesis and facilitate dense estimation of scene properties. However, NeRFs often fail for outdoor, unbounded scenes that are captured under very sparse views with the scene content concentrated far away from the

Cited by 26SourceScholar
2023

Hyperspherical Embedding for Point Cloud Completion

CVPR 2023poster

Most real-world 3D measurements from depth sensors are incomplete, and to address this issue the point cloud completion task aims to predict the complete shapes of objects from partial observations. Previous works often adapt an encoder-decoder architecture, where the encoder is trained to extract e…

2022

Hybrid Visual SLAM for Underwater Vehicle Manipulator Systems

RA-L 2022

This letter presents a novel visual feature based scene mapping method for underwater vehicle manipulator systems (UVMSs), with specific emphasis on robust mapping in natural seafloor environments. Our method uses GPU accelerated SIFT features in a graph optimization framework to build a feature map

Cited by 33SourcecodeScholar
2021

A Kinematic Model for Trajectory Prediction in General Highway Scenarios

RA-L 2021

Highway driving invariably combines high speeds with the need to interact closely with other drivers. Prediction methods enable autonomous vehicles (AVs) to anticipate drivers’ future trajectories and plan accordingly. Kinematic methods for prediction have traditionally ignored the presence of other

Cited by 21SourceScholar
2021

BiTraP: Bi-Directional Pedestrian Trajectory Prediction With Multi-Modal Goal Estimation

RA-L 2021

Pedestrian trajectory prediction is an essential task in robotic applications such as autonomous driving and robot navigation. State-of-the-art trajectory predictors use a conditional variational autoencoder (CVAE) with recurrent neural networks (RNNs) to encode observed trajectories and decode mult

Cited by 185SourcecodeScholar
2021

Coupling Intent and Action for Pedestrian Crossing Behavior Prediction

IJCAI 2021poster

Accurate prediction of pedestrian crossing behaviors by autonomous vehicles can significantly improve traffic safety. Existing approaches often model pedestrian behaviors using trajectories or poses but do not offer a deeper semantic interpretation of a person's actions or how actions influence a pe…

2021

Energy-optimal Path Planning with Active Flow Perception for Autonomous Underwater Vehicles

ICRA 2021poster

Accurate flow predictions are critical for energy-optimal path planning of AUVs with endurance requirements. However, the complex dynamics of ocean currents make it difficult to achieve accurate flow predictions. For an AUV with flow and location sensing capabilities, one can optimize vehicle action…

Cited by 7SourceScholar
2021

Point Set Voting for Partial Point Cloud Analysis

RA-L 2021

The continual improvement of 3D sensors has driven the development of algorithms to perform point cloud analysis. In fact, techniques for point cloud classification and segmentation have in recent years achieved incredible performance driven in part by leveraging large synthetic datasets. Unfortunat

Cited by 42SourceScholar
2020

Leveraging the Template and Anchor Framework for Safe, Online Robotic Gait Design

ICRA 2020poster

Online control design using a high-fidelity, full-order model for a bipedal robot can be challenging due to the size of the state space of the model. A commonly adopted solution to overcome this challenge is to approximate the fullorder model (anchor) with a simplified, reduced-order model (template…

Cited by 16SourcecodeScholar
2020

LiStereo: Generate Dense Depth Maps from LIDAR and Stereo Imagery

ICRA 2020poster

An accurate depth map of the environment is critical to the safe operation of autonomous robots and vehicles. Currently, either light detection and ranging (LIDAR) or stereo matching algorithms are used to acquire such depth information. However, a high-resolution LIDAR is expensive and produces spa…

Cited by 42SourceScholar
2020

Off the Beaten Sidewalk: Pedestrian Prediction in Shared Spaces for Autonomous Vehicles

RA-L 2020

Pedestrians and drivers interact closely in a wide range of environments. Autonomous vehicles (AVs) correspondingly face the need to predict pedestrians' future trajectories in these same environments. Traditional model-based prediction methods have been limited to making predictions in highly struc

Cited by 17SourceScholar
2020

Pedestrian Planar LiDAR Pose (PPLP) Network for Oriented Pedestrian Detection Based on Planar LiDAR and Monocular Images

RA-L 2020

Pedestrian detection is an important task for human-robot interaction and autonomous driving applications. Most previous pedestrian detection methods rely on data collected from three-dimensional (3D) Light Detection and Ranging (LiDAR) sensors in addition to camera imagery, which can be expensive t

Cited by 25SourceScholar
2020

Risk Assessment and Planning with Bidirectional Reachability for Autonomous Driving

ICRA 2020poster

Risk assessment to quantify the danger associated with taking a certain action is critical to navigating safely through crowded urban environments during autonomous driving. Risk assessment and subsequent planning is usually done by first tracking and predicting trajectories of other agents, such as…

Cited by 40SourceScholar
2020

SilhoNet-Fisheye: Adaptation of A ROI Based Object Pose Estimation Network to Monocular Fisheye Images

RA-L 2020

There has been much recent interest in deep learning methods for monocular image based object pose estimation. While object pose estimation is an important problem for autonomous robot interaction with the physical world, and the application space for monocular-based methods is expansive, there has

Cited by 11SourceScholar
2019

A constrained control-planning strategy for redundant manipulators

ICRA 2019poster

This paper presents an interconnected control-planning strategy for redundant manipulators, subject to system and environmental constraints. The method incorporates low-level control characteristics and high-level planning components into a robust strategy for manipulators acting in complex environm…

Cited by 1SourceScholar
2019

Bio-LSTM: A Biomechanically Inspired Recurrent Neural Network for 3-D Pedestrian Pose and Gait Prediction

RA-L 2019

In applications, such as autonomous driving, it is important to understand, infer, and anticipate the intention and future behavior of pedestrians. This ability allows vehicles to avoid collisions and improve ride safety and quality. This letter proposes a biomechanically inspired recurrent neural n

Cited by 80SourceScholar
2019

DispSegNet: Leveraging Semantics for End-to-End Learning of Disparity Estimation From Stereo Imagery

RA-L 2019

Recent work has shown that convolutional neural networks (CNNs) can be applied successfully in disparity estimation, but these methods still suffer from errors in regions of low texture, occlusions, and reflections. Concurrently, deep learning for semantic segmentation has shown great progress in re

Cited by 60SourceScholar
2019

Localization and Tracking of Uncontrollable Underwater Agents: Particle Filter Based Fusion of On-Body IMUs and Stationary Cameras

ICRA 2019poster

Tracking of uncontrollable agents in a controlled environment is an important research question for the coordination of controllable and uncontrollable agents and bio-inspired multi-agent control. This paper presents a framework that approaches the multiagent tracking problem from a localization per…

Cited by 15SourceScholar
2019

Occlusion-Aware Risk Assessment for Autonomous Driving in Urban Environments

RA-L 2019

Navigating safely in urban environments remains a challenging problem for autonomous vehicles. Occlusion and limited sensor range can pose significant challenges to safely navigate among pedestrians and other vehicles in the environment. Enabling vehicles to quantify the risk posed by unseen regions

Cited by 113SourceScholar
2019

PedX: Benchmark Dataset for Metric 3-D Pose Estimation of Pedestrians in Complex Urban Intersections

RA-L 2019

This letter presents a novel dataset titled PedX, a large-scale multimodal collection of pedestrians at complex urban intersections. PedX consists of more than 5 000 pairs of high-resolution (12MP) stereo images and LiDAR data along with providing two-dimensional (2-D) image labels and 3-D labels of

Cited by 73SourceScholar
2019

Sensor Transfer: Learning Optimal Sensor Effect Image Augmentation for Sim-to-Real Domain Adaptation

RA-L 2019

Performance on benchmark datasets has drastically improved with advances in deep learning. Still, cross-dataset generalization performance remains relatively low due to the domain shift that can occur between two different datasets. This domain shift is especially exaggerated between synthetic and r

Cited by 30SourceScholar
2019

Stochastic Sampling Simulation for Pedestrian Trajectory Prediction

IROS 2019poster

Urban environments pose a significant challenge for autonomous vehicles (AVs) as they must safely navigate while in close proximity to many pedestrians. It is crucial for the AV to correctly understand and predict the future trajectories of pedestrians to avoid collision and plan a safe path. Deep n…

Cited by 22SourceScholar
2019

Towards Provably Not-At-Fault Control of Autonomous Robots in Arbitrary Dynamic Environments

RSS 2019poster

As autonomous robots increasingly become part of daily life, they will often encounter dynamic environments while only having limited information about their surroundings. Unfortunately, due to the possible presence of malicious dynamic actors, it is infeasible to develop an algorithm that can guara…

Cited by 62SourcePDFScholar
2019

UWStereoNet: Unsupervised Learning for Depth Estimation and Color Correction of Underwater Stereo Imagery

ICRA 2019poster

Stereo cameras are widely used for sensing and navigation of underwater robotic systems. They can provide high resolution color views of a scene; the constrained camera geometry enables metrically accurate depth estimation; they are also relatively cost-effective. Traditional stereo vision algorithm…

Cited by 53SourceScholar
2018

Failing to Learn: Autonomously Identifying Perception Failures for Self-Driving Cars

RA-L 2018

One of the major open challenges in self-driving cars is the ability to detect cars and pedestrians to safely navigate in the world. Deep learning-based object detector approaches have enabled great advances in using camera imagery to detect and classify objects. But for a safety critical applicatio

Cited by 114SourceScholar
2018

WaterGAN: Unsupervised Generative Network to Enable Real-Time Color Correction of Monocular Underwater Images

RA-L 2018

This letter reports on WaterGAN, a generative adversarial network (GAN) for generating realistic underwater images from in-air image and depth pairings in an unsupervised pipeline used for color correction of monocular underwater images. Cameras onboard autonomous and remotely operated vehicles can

Cited by 848SourcecodeScholar
2017

A framework for enhanced localization of marine mammals using auto-detected video and wearable sensor data fusion

IROS 2017poster

Accurate biological agent localization offers the opportunity for both researchers and institutions to gain new knowledge about individual and group behaviors of biosystems. This paper presents a sensor-fusion approach for tracking biological agents, combining the data from automated video logging w…

Cited by 12SourceScholar
2017

Automatic color correction for 3D reconstruction of underwater scenes

ICRA 2017poster

Mapping of underwater environments is a critical task for a range of activities from monitoring coral reef habitats to surveying submerged archaeological sites. While recent advances in methods for terrestrial mapping can achieve dense 3D reconstructions of scenes in real-time, there remains the cha…

Cited by 26SourceScholar
2017

Driving in the Matrix: Can virtual worlds replace human-generated annotations for real world tasks?

ICRA 2017poster

Deep learning has rapidly transformed the state of the art algorithms used to address a variety of problems in computer vision and robotics. These breakthroughs have relied upon massive amounts of human annotated training data. This time consuming process has begun impeding the progress of these dee…

Cited by 844SourcecodeScholar
2017

Real-Time Certified Probabilistic Pedestrian Forecasting

RA-L 2017

The success of autonomous systems will depend upon their ability to safely navigate human-centric environments. This motivates the need for a real-time, probabilistic forecasting algorithm for pedestrians, cyclists, and other agents since these predictions will form a necessary step in assessing the

Cited by 18SourceScholar
2016

Utilizing high-dimensional features for real-time robotic applications: Reducing the curse of dimensionality for recursive Bayesian estimation

IROS 2016poster

Feature learning has become popular in robotics due to recent advances in machine learning. In this paper, we propose a novel method to utilize the high-dimensional features from these techniques as observations in Bayesian estimation problems in a real-time manner. We develop an approach that: 1) p…

Cited by 25SourceScholar
2015

Building 3D mosaics from an Autonomous Underwater Vehicle, Doppler velocity log, and 2D imaging sonar

ICRA 2015poster

This paper reports on a 3D photomosaicing pipeline using data collected from an autonomous underwater vehicle performing simultaneous localization and mapping (SLAM). The pipeline projects and blends 2D imaging sonar data onto a large-scale 3D mesh that is either given a priori or derived from SLAM.…

Cited by 33SourceScholar