← Search

Jean Oh

53 accepted papers

2026

DYMO-Hair: Generalizable Volumetric Dynamics Modeling for Robot Hair Manipulation

ICRA 2026poster

Hair care is an essential daily activity, yet it remains inaccessible to individuals with limited mobility and challenging for autonomous robot systems due to the fine-grained physical structure and complex dynamics of hair. In this work, we present DYMO-Hair, a model-based robot hair care system. W…

2026

Functional Force-Aware Retargeting from Virtual Human Demos to Soft Robot Policies

RSS 2026poster

We introduce SoftAct, a framework for teaching soft robot hands to perform human-like manipulation skills by explicitly reasoning about contact forces. Leveraging immersive virtual reality, our system captures rich human demonstrations, including hand kinematics, object motion, dense contact patches…

Cited by 0SourceScholar
2026

GRAPPA: Generalizing and Adapting Robot Policies via Online Agentic Guidance

RA-L 2026

Robot learning approaches such as behavior cloning and reinforcement learning have shown great promise in synthesizing robot skills from human demonstrations in specific environments. However, these approaches often struggle to generalize to unseen real-world settings because they rely on task-speci

Cited by 0SourceScholar
2026

Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications

ICRA 2026poster

Mobile robot navigation in dynamic human environments requires policies that balance adaptability to diverse behaviors with compliance to safety constraints. We hypothesize that integrating data-driven rewards with rule-based objectives enables navigation policies to achieve a more effective balance…

2026

RIO: Flexible Real-time Robot I/O for Cross-Embodiment Robot Learning

RSS 2026poster

Despite recent efforts to collect multi-task or multiembodiment datasets, to design efficient recipes for training Vision-Language-Action models (VLAs), and to showcase these models on selected robot platforms, generalist robot capabilities and cross-embodiment transfer remain largely elusive ideals…

Cited by 0SourceScholar
2026

SEAL: Towards Safe Autonomous Driving Via Skill-Enabled Adversary Learning for Closed-Loop Scenario Generation

ICRA 2026poster

Verification and validation of autonomous driving (AD) systems and components is of increasing importance, as such technology increases in real-world prevalence. Safety-critical scenario generation is a key approach to robustify AD policies through closed-loop training. However, existing approaches …

2026

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

ICRA 2026poster

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs to navigation presents two key challenges: effectively parsing and structuring complex environment information and determ…

2026

Spline-FRIDA: Towards Diverse, Humanlike Robot Painting Styles with a Sample-Efficient, Differentiable Brush Stroke Model

ICRA 2026poster

A painting is more than just a picture on a wall; a painting is a process comprised of many intentional brush strokes, the shapes of which are an important component of a painting's overall style and message. Prior work in modeling brush stroke trajectories either does not work with real-world robot…

2025

Bridging Spectral-Wise and Multi-Spectral Depth Estimation Via Geometry-Guided Contrastive Learning

ICRA 2025

Deploying depth estimation networks in the real world requires high-level robustness against various adverse conditions to ensure safe and reliable autonomy. For this purpose, many autonomous vehicles employ multi-modal sensor systems, including an RGB camera, NIR camera, thermal camera, LiDAR, or R

Cited by 3SourcecodeScholar
2025

FIReStereo: Forest InfraRed Stereo Dataset for UAS Depth Perception in Visually Degraded Environments

RA-L 2025

Robust depth perception in visually-degraded environments is crucial for autonomous aerial systems. Thermal imaging cameras, which capture infrared radiation, are robust to visual degradation. However, due to lack of a large-scale dataset, the use of thermal cameras for uncrewed aerial system (UAS)

Cited by 8SourceScholar
2025

Flow4D: Leveraging 4D Voxel Network for LiDAR Scene Flow Estimation

RA-L 2025

Understanding the motion states of the surrounding environment is critical for safe autonomous driving. These motion states can be accurately derived from scene flow, which captures the three-dimensional motion field of points. Existing LiDAR scene flow methods extract spatial features from each poi

Cited by 20SourcecodeScholar
2025

KineSoft: Learning Proprioceptive Manipulation Policies with Soft Robot Hands

CoRL 2025oral

Underactuated soft robot hands offer inherent safety and adaptability advantages over rigid systems, but developing dexterous manipulation skills remains challenging. While imitation learning shows promise for complex manipulation tasks, traditional approaches struggle with soft systems due to demon…

Cited by 0SourceScholar
2025

MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments

ICCV 2025poster

We introduce a diffusion-based approach for generating privacy-preserving digital twins of multi-room indoor environments from depth images only. Central to our approach is a novel Multi-view Overlapped Scene Alignment with Implicit Consistency (MOSAIC) model that explicitly considers cross-view dep…

Cited by 0SourcePDFScholar
2025

SEAL: Towards Safe Autonomous Driving via Skill-Enabled Adversary Learning for Closed-Loop Scenario Generation

RA-L 2025

Verification and validation of autonomous driving (AD) systems and components is of increasing importance, as such technology increases in real-world prevalence. Safety-critical scenario generation is a key approach to robustify AD policies through closed-loop training. However, existing approaches

Cited by 10SourceScholar
2025

SonicBoom: Contact Localization Using Array of Microphones

RA-L 2025

In cluttered environments where visual sensors encounter heavy occlusion, such as in agricultural settings, tactile signals can provide crucial spatial information for the robot to locate rigid objects and maneuver around them. We introduce SonicBoom, a holistic hardware and learning pipeline that e

Cited by 6SourcecodeScholar
2025

Spline-FRIDA: Towards Diverse, Humanlike Robot Painting Styles With a Sample-Efficient, Differentiable Brush Stroke Model

RA-L 2025

A painting is more than just a picture on a wall; a painting is a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">process</i> comprised of many intentional brush strokes, the shapes of which are an important component of a painting's overall style an

Cited by 5SourceScholar
2025

Stabilizing Reinforcement Learning in Differentiable Multiphysics Simulation

ICLR 2025spotlight

Recent advances in GPU-based parallel simulation have enabled practitioners to collect large amounts of data and train complex control policies using deep reinforcement learning (RL), on commodity GPUs. However, such successes for RL in robotics have been limited to tasks sufficiently simulated by f…

2025

VPOcc: Exploiting Vanishing Point for 3D Semantic Occupancy Prediction

IROS 2025

Understanding 3D scenes semantically and spatially is crucial for the safe navigation of robots and autonomous vehicles, aiding obstacle avoidance and accurate trajectory planning. Camera-based 3D semantic occupancy prediction, which infers complete voxel grids from 2D images, is gaining importance

Cited by 1SourcecodeScholar
2024

CoFRIDA: Self-Supervised Fine-Tuning for Human-Robot Co-Painting

ICRA 2024poster

Prior robot painting and drawing work, such as FRIDA, has focused on decreasing the sim-to-real gap and expanding input modalities for users, but the interaction with these systems generally exists only in the input stages. To support interactive, human-robot collaborative painting, we introduce the…

Cited by 15SourcecodeScholar
2024

Complementary Random Masking for RGB-Thermal Semantic Segmentation

ICRA 2024poster

RGB-thermal semantic segmentation is one potential solution to achieve reliable semantic scene understanding in adverse weather and lighting conditions. However, the previous studies mostly focus on designing a multi-modal fusion module without consideration of the nature of multi-modality inputs. T…

Cited by 28SourcecodeScholar
2024

Density-aware Domain Generalization for LiDAR Semantic Segmentation

IROS 2024poster

3D LiDAR-based perception has made remarkable advancements, leading to the widespread adoption of LiDAR in autonomous driving systems. Despite these technological strides, variations in LiDAR sensors and environmental conditions can significantly deteriorate the performance of perception models, pri…

Cited by 2SourceScholar
2024

POE: Acoustic Soft Robotic Proprioception for Omnidirectional End-effectors

ICRA 2024poster

Shape estimation is crucial for precise control of soft robots. However, soft robot shape estimation and proprioception are challenging due to their complex deformation behaviors and infinite degrees of freedom. Their continuously deforming bodies complicate integrating rigid sensors and reliably es…

Cited by 9SourceScholar
2024

SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation

CVPR 2024poster

Accurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with harmful stereotypes and misrepresentations of cultures. We improve inclusive representation i…

Cited by 17SourcePDFScholar
2024

SoRTS: Learned Tree Search for Long Horizon Social Robot Navigation

RA-L 2024

The fast-growing demand for fully autonomous robots in shared spaces calls for developing trustworthy agents that can safely and seamlessly navigate crowded environments. Recent models for motion prediction show promise in characterizing social interactions in such environments. However, using them

Cited by 5SourcecodeScholar
2024

Towards Human-Centered Construction Robotics: A Reinforcement Learning-Driven Companion Robot for Contextually Assisting Carpentry Workers

IROS 2024poster

In the dynamic construction industry, traditional robotic integration has primarily focused on automating specific tasks, often overlooking the complexity and variability of human aspects in construction workflows. This paper introduces a human-centered approach with a "work companion rover" designe…

Cited by 3SourceScholar
2024

Translating Agent-Environment Interactions from Humans to Robots

IROS 2024poster

Humans are remarkably adept at imitating other people performing tasks, afforded by their ability to abstract away irrelevant details and focus on the task strategy of the demonstrator. In this paper, we take steps towards enabling robots with this ability, and present a framework, TransAct to do so…

Cited by 0SourceScholar
2023

Core Challenges in Embodied Vision-Language Planning (Extended Abstract)

IJCAI 2023poster

Recent advances in the areas of Multimodal Machine Learning and Artificial Intelligence (AI) have led to the development of challenging tasks at the intersection of Computer Vision, Natural Language Processing, and Robotics. Whereas many approaches and previous survey pursuits have characterised one…

Cited by 0SourcePDFScholar
2023

FRIDA: A Collaborative Robot Painter with a Differentiable, Real2Sim2Real Planning Environment

ICRA 2023poster

Painting is an artistic process of rendering visual content that achieves the high-level communication goals of an artist that may change dynamically throughout the creative process. In this paper, we present a Framework and Robotics Initiative for Developing Arts (FRIDA) that enables humans to prod…

Cited by 29SourcecodeScholar
2023

Follow The Rules: Online Signal Temporal Logic Tree Search for Guided Imitation Learning in Stochastic Domains

ICRA 2023poster

Seamlessly integrating rules in Learning-from-Demonstrations (LfD) policies is a critical requirement to enable the real-world deployment of AI agents. Recently, Signal Temporal Logic (STL) has been shown to be an effective language for encoding rules as spatio-temporal constraints. This work uses M…

Cited by 14SourcecodeScholar
2023

T2FPV: Dataset and Method for Correcting First-Person View Errors in Pedestrian Trajectory Prediction

IROS 2023poster

Predicting pedestrian motion is essential for developing socially-aware robots that interact in a crowded environment. While the natural visual perspective for a social interaction setting is an egocentric view, the majority of existing work in trajectory prediction therein has been investigated pur…

Cited by 5SourcecodeScholar
2022

Autonomous Exploration Development Environment and the Planning Algorithms

ICRA 2022poster

Autonomous Exploration Development Environment is an open-source repository released to facilitate development of high-level planning algorithms and integration of com-plete autonomous navigation systems. The repository contains representative simulation environment models, fundamental navigation mo…

Cited by 97SourceScholar
2022

FAR Planner: Fast, Attemptable Route Planner using Dynamic Visibility Update

IROS 2022poster

Path planning in unknown environments remains a challenging problem, as the environment is gradually observed during the navigation, the underlying planner has to update the environment representation and replan, promptly and constantly, to account for the new observations. In this paper, we present…

Cited by 58SourcecodeScholar
2022

Predicting Like A Pilot: Dataset and Method to Predict Socially-Aware Aircraft Trajectories in Non-Towered Terminal Airspace

ICRA 2022poster

Pilots operating aircraft in non-towered terminal airspace rely on their situational awareness and prior knowledge to predict the future trajectories of other agents. These predictions are conditioned on the past trajectories of other agents, agent-agent social interactions and environmental context…

Cited by 25SourcecodeScholar
2022

StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Translation

IJCAI 2022poster

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control of the style of image to be generated. We present an approach for generating styl…

2022

Translating Robot Skills: Learning Unsupervised Skill Correspondences Across Robots

ICML 2022spotlight

In this paper, we explore how we can endow robots with the ability to learn correspondences between their own skills, and those of morphologically different robots in different domains, in an entirely unsupervised manner. We make the insight that different morphological robots use similar task strat…

Cited by 9SourcePDFScholar
2021

Content Masked Loss: Human-Like Brush Stroke Planning in a Reinforcement Learning Painting Agent

AAAI 2021technical

The objective of most Reinforcement Learning painting agents is to minimize the loss between a target image and the paint canvas. Human painter artistry emphasizes important features of the target image rather than simply reproducing it. Using adversarial or L2 losses in the RL painting models, al…

2020

Learning Shape-based Representation for Visual Localization in Extremely Changing Conditions

ICRA 2020poster

Visual localization is an important task for applications such as navigation and augmented reality, but is a challenging problem when there are changes in scene appearances through day, seasons, or environments. In this paper, we present a convolutional neural network (CNN)-based approach for visual…

Cited by 6SourceScholar
2015

Inferring door locations from a teammate's trajectory in stealth human-robot team operations

IROS 2015poster

Robot perception is generally viewed as the interpretation of data from various types of sensors such as cameras. In this paper, we study indirect perception where a robot can perceive new information by making inferences from non-visual observations of human teammates. As a proof-of-concept study,…

Cited by 5SourceScholar