← Search

Justin Kerr

22 accepted papers

2026

Viser: Imperative, Web-based 3D Visualization for Python

RSS 2026poster

We present Viser, a toolkit for 3D visualization in robotics and computer vision. Viser aims to bring easy and extensible 3D visualization to Python: we provide comprehensive 3D scene and 2D GUI primitives, which can be used independently with minimal setup or composed to build specialized interface…

Cited by 0SourceScholar
2025

Blox-Net: Generative Design-for-Robot-Assembly Using VLM Supervision, Physics Simulation, and a Robot with Reset

ICRA 2025

Generative AI systems have shown impressive capabilities in creating text, code, and images. Inspired by the importance of research in industrial Design for Assembly, we introduce a novel problem: Generative Design-for-RobotAssembly (GDfRA). The task is to generate an assembly based on a natural lan

Cited by 16SourceScholar
2025

Botany-Bot: Digital Twin Monitoring of Occluded and Underleaf Plant Structures with Gaussian Splats

IROS 2025

Commercial plant phenotyping systems using fixed cameras cannot perceive many plant details due to leaf occlusion. In this paper, we present Botany-Bot, a system for building detailed “annotated digital twins” of living plants using two stereo cameras, a digital turntable inside a lightbox, an indus

Cited by 0SourcecodeScholar
2025

Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop

CoRL 2025poster

Humans do not passively observe the visual world---we actively look in order to act. Motivated by this principle, we introduce EyeRobot, a robotic system with gaze behavior that emerges from the need to complete real-world tasks. We develop a mechanical eyeball that can freely rotate to observe its…

Cited by 0SourceScholar
2025

Omni-Scan: Creating Visually-Accurate Digital Twin Object Models Using a Bimanual Robot with Handover and Gaussian Splat Merging

IROS 2025

3D Gaussian Splats (3DGSs) are 3D object models derived from multi-view images. Such “digital twins” are useful for simulations, virtual reality, E-commerce, robot policy fine-tuning, and part inspection. 3D object scanning usually requires multi-camera arrays, precise laser scanners, or robot wrist

Cited by 2SourcecodeScholar
2025

Persistent Object Gaussian Splat (POGS) for Tracking Human and Robot Manipulation of Irregularly Shaped Objects

ICRA 2025

Tracking and manipulating irregularly-shaped, previously unseen objects in dynamic environments is important for robotic applications in manufacturing, assembly, and logistics. Recently introduced Gaussian Splats [1] efficiently model object geometry, but lack persistent state estimation for taskori

Cited by 11SourceScholar
2025

Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding

ICCV 2025poster

Whether snipping with scissors or opening a box, humans can quickly understand the 3D configurations of familiar objects. For novel objects, we can resort to long-form inspection to build intuition. The more we observe the object, the better we get at predicting its 3D state immediately. Existing sy…

Cited by 0SourcePDFScholar
2024

GARField: Group Anything with Radiance Fields

CVPR 2024poster

Grouping is inherently ambiguous due to the multiple levels of granularity in which one can decompose a scene --- should the wheels of an excavator be considered separate or part of the whole? We propose Group Anything with Radiance Fields (GARField) an approach for decomposing 3D scenes into a hier…

2024

Language-Embedded Gaussian Splats (LEGS): Incrementally Building Room-Scale Representations with a Mobile Robot

IROS 2024

Building semantic 3D maps is valuable for searching for objects of interest in offices, warehouses, stores, and homes. We present a mapping system that incrementally builds a Language-Embedded Gaussian Splat (LEGS): a detailed 3D scene representation that encodes both appearance and semantics in a u

Cited by 28SourcecodeScholar
2024

Learning Robotic Locomotion Affordances and Photorealistic Simulators from Human-Captured Data

CoRL 2024poster

Learning reliable affordance models which satisfy human preferences is often hindered by a lack of high-quality training data. Similarly, learning visuomotor policies in simulation can be challenging due to the high cost of photo-realistic rendering. We present PAWS: a comprehensive robot learning f…

Cited by 1SourceScholar
2024

Lifelong LERF: Local 3D Semantic Inventory Monitoring Using FogROS2

ICRA 2024poster

Inventory monitoring in homes, factories, and retail stores relies on maintaining data despite objects being swapped, added, removed, or moved. We introduce Lifelong LERF, a method that allows a mobile robot with minimal compute to jointly optimize a dense language and geometric representation of it…

Cited by 6SourceScholar
2024

Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction

CoRL 2024poster

Humans can learn to manipulate new objects by simply watching others; providing robots with the ability to learn from such demonstrations would enable a natural interface specifying new behaviors. This work develops Robot See Robot Do (RSRD), a method for imitating articulated object manipulation fr…

Cited by 15SourcecodeScholar
2023

HANDLOOM: Learned Tracing of One-Dimensional Objects for Inspection and Manipulation

CoRL 2023oral

Tracing – estimating the spatial state of – long deformable linear objects such as cables, threads, hoses, or ropes, is useful for a broad range of tasks in homes, retail, factories, construction, transportation, and healthcare. For long deformable linear objects (DLOs or simply cables) with many (o…

Cited by 6SourcecodeScholar
2023

Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping

CoRL 2023oral

Grasping objects by a specific subpart is often crucial for safety and for executing downstream tasks. We propose LERF-TOGO, Language Embedded Radiance Fields for Task-Oriented Grasping of Objects, which uses vision-language models zero-shot to output a grasp distribution over an object given a natu…

Cited by 88SourcecodeScholar
2023

SGTM 2.0: Autonomously Untangling Long Cables using Interactive Perception

ICRA 2023poster

Cables are commonplace in homes, hospitals, and industrial warehouses and are prone to tangling. This paper extends prior work on autonomously untangling long cables by introducing novel uncertainty quantification metrics and actions that interact with the cable to reduce perception uncertainty. We…

Cited by 20SourceScholar
2023

Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

RSS 2023poster

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work demonstrates the efficacy of tactile sensing for precise manip…

Cited by 33SourcePDFScholar
2022

All You Need is LUV: Unsupervised Collection of Labeled Images Using UV-Fluorescent Markings

IROS 2022poster

Learning-based perception systems in robotics often requires large-scale image segmentation annotation. Current approaches rely on human labelers, which can be expensive, or simulation data, which can visually differ from real data. This paper proposes Labels from UltraViolet (LUV), a novel framewor…

Cited by 12SourceScholar
2022

Evo-NeRF: Evolving NeRF for Sequential Robot Grasping of Transparent Objects

CoRL 2022oral

Sequential robot grasping of transparent objects, where a robot removes objects one by one from a workspace, is important in many industrial and household scenarios. We propose Evolving NeRF (Evo-NeRF), leveraging recent speedups in NeRF training and further extending it to rapidly train the NeRF re…

Cited by 100SourceScholar
2022

Learning to Localize, Grasp, and Hand Over Unmodified Surgical Needles

ICRA 2022poster

Robotic Surgical Assistants (RSAs) are commonly used to perform minimally invasive surgeries by expert surgeons. However, long procedures filled with tedious and repetitive tasks such as suturing can lead to surgeon fatigue, motivating the automation of suturing. As visual tracking of a thin reflect…

Cited by 34SourceScholar
2021

Dex-NeRF: Using a Neural Radiance Field to Grasp Transparent Objects

CoRL 2021poster

The ability to grasp and manipulate transparent objects is a major challenge for robots. Existing depth cameras have difficulty detecting, localizing, and inferring the geometry of such objects. We propose using neural radiance fields (NeRF) to detect, localize, and infer the geometry of transparent…

Cited by 196SourceScholar
2019

PRIMAL: Pathfinding via Reinforcement and Imitation Multi-Agent Learning

RA-L 2019

Multi-agent path finding (MAPF) is an essential component of many large-scale, real-world robot deployments, from aerial swarms to warehouse automation. However, despite the community's continued efforts, most state-of-the-art MAPF planners still rely on centralized planning and scale poorly past a

Cited by 398SourceScholar