← Search

Songhwai Oh

83 accepted papers

2026

Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement Learning

ICML 2026poster

In offline goal-conditioned reinforcement learning (GCRL), where one relies on a limited reward-free dataset to learn a generalist goal-reaching agent, compositional generalization becomes essential for reaching unseen goals under novel contextual variations. Most prior approaches pursue this via tr…

Cited by 0SourceScholar
2026

Memory-Efficient Voxelized Renderable Neural 3D Spatial Representation for Vision-Based Robotics

RA-L 2026

In this paper, we introduce a novel approach for modeling a memory-efficient spatial representation with 3D Gaussian splatting. Efficient vision-based spatial representation poses a significant challenge due to the memory demands of visual information. Recent advances in 3D rendering technologies, s

Cited by 0SourceScholar
2026

Memory-Efficient Voxelized Renderable Neural 3D Spatial Representation for Vision-Based Robotics

ICRA 2026poster

In this paper, we introduce a novel approach for modeling a memory-efficient spatial representation with 3D Gaussian splatting. Efficient vision-based spatial representation poses a significant challenge due to the memory demands of visual information. Recent advances in 3D rendering technologies, s…

Cited by 0SourceScholar
2026

Playbook: Scalable Discrete Skill Discovery from Unstructured Datasets for Long-Horizon Decision-Making Problems

ICRA 2026poster

Skill discovery methods enable agents to tackle intricate tasks by acquiring diverse and useful skills from task-agnostic datasets in an unsupervised manner. To apply these methods to more general and everyday tasks, the skill set must be scalable. However, current approaches struggle with this scal…

Cited by 0SourceScholar
2026

Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement

ICRA 2026poster

In this paper, we present the tidiness score-guided Monte Carlo tree search (TSMCTS), a novel framework designed to address the tabletop tidying up problem using only an RGB-D camera. We address two major problems for tabletop tidying up problem: (1) the lack of public datasets and benchmarks, and (…

2025

Automatic Real-to-Sim-to-Real System through Iterative Interactions for Robust Robot Manipulation Policy Learning with Unseen Objects

IROS 2025

Real-to-sim-to-real systems have been studied to overcome the challenges of robot policy learning in the real world by creating a virtual environment that mimics the actual workspace. However, previous studies have limitations, requiring human assistance, such as observing the workspace with a hand-

Cited by 0SourceScholar
2025

Conflict-Averse Gradient Aggregation for Constrained Multi-Objective Reinforcement Learning

ICLR 2025poster

In real-world applications, a reinforcement learning (RL) agent should consider multiple objectives and adhere to safety guidelines. To address these considerations, we propose a constrained multi-objective RL algorithm named constrained multi-objective gradient aggregator (CoMOGA). In the field of…

Cited by 0SourcePDFScholar
2025

Playbook: Scalable Discrete Skill Discovery From Unstructured Datasets for Long-Horizon Decision-Making Problems

RA-L 2025

Skill discovery methods enable agents to tackle intricate tasks by acquiring diverse and useful skills from task-agnostic datasets in an unsupervised manner. To apply these methods to more general and everyday tasks, the skill set must be scalable. However, current approaches struggle with this scal

Cited by 0SourcecodeScholar
2025

Stage-Wise Reward Shaping for Acrobatic Robots: A Constrained Multi-Objective Reinforcement Learning Approach

ICRA 2025

As the complexity of tasks addressed through reinforcement learning (RL) increases, the definition of reward functions also has become highly complicated. We introduce an RL method aimed at simplifying the reward-shaping process through intuitive strategies. Initially, instead of a single reward fun

Cited by 16SourcecodeScholar
2025

Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement

RA-L 2025

In this paper, we present the tidiness score-guided Monte Carlo tree search (TSMCTS), a novel framework designed to address the tabletop tidying up problem using only an RGB-D camera. We address two major problems for tabletop tidying up problem: (1) the lack of public datasets and benchmarks, and (

Cited by 2SourcecodeScholar
2024

Adversarial Environment Design via Regret-Guided Diffusion Models

NeurIPS 2024spotlight

Training agents that are robust to environmental changes remains a significant challenge in deep reinforcement learning (RL). Unsupervised environment design (UED) has recently emerged to address this issue by generating a set of training environments tailored to the agent's capabilities. While prio…

Cited by 0SourcePDFScholar
2024

Gradual Receptive Expansion Using Vision Transformer for Online 3D Bin Packing

IROS 2024poster

The bin packing problem (BPP) is a challenging combinatorial optimization problem with a number of practical applications. This paper focuses on online 3D-BPP, where the packer makes immediate decisions for a loading position as items continually arrive. We propose a novel reinforcement learning alg…

Cited by 0SourceScholar
2024

MAC-ID: Multi-Agent Reinforcement Learning with Local Coordination for Individual Diversity

ICRA 2024poster

With the increase of robots navigating through crowded environments in our daily lives, the demand for designing a socially-aware navigation method considering humanrobot interaction has risen. When developing and assessing socially-aware navigation methods, pedestrian motion modeling plays a signif…

Cited by 0SourceScholar
2024

RNR-Nav: A Real-World Visual Navigation System Using Renderable Neural Radiance Maps

IROS 2024

We propose a novel visual localization and navigation framework for real-world environments directly integrating observed visual information into the bird-eye-view map. While the renderable neural radiance map (RNR-Map) [1] shows considerable promise in simulated settings, its deployment in real-wor

Cited by 1SourceScholar
2024

Renderable Street View Map-Based Localization: Leveraging 3D Gaussian Splatting for Street-Level Positioning

IROS 2024poster

In this paper, we introduce a new method that first utilizes 3D Gaussian splatting in street-level localization problem. Robust localization with street-level real-world images such as street view is a major issue for autonomous vehicle, augmented reality (AR) navigation, and outdoor mobile robots.…

Cited by 1SourceScholar
2024

Safe CoR: A Dual-Expert Approach to Integrating Imitation Learning and Safe Reinforcement Learning Using Constraint Rewards

IROS 2024poster

In the realm of autonomous agents, ensuring safety and reliability in complex and dynamic environments remains a paramount challenge. Safe reinforcement learning addresses these concerns by introducing safety constraints, but still faces challenges in navigating intricate environments such as comple…

Cited by 1SourceScholar
2024

Spectral-Risk Safe Reinforcement Learning with Convergence Guarantees

NeurIPS 2024poster

The field of risk-constrained reinforcement learning (RCRL) has been developed to effectively reduce the likelihood of worst-case scenarios by explicitly handling risk-measure-based constraints. However, the nonlinearity of risk measures makes it challenging to achieve convergence and optimality. To…

2024

Unsupervised 3D Part Decomposition via Leveraged Gaussian Splatting

IROS 2024poster

We propose a novel unsupervised method for motion-based 3D part decomposition of articulated objects using a single monocular video of a dynamic scene. In contrast to existing unsupervised methods relying on optical flow or tracking techniques, our approach addresses this problem without additional…

Cited by 0SourcecodeScholar
2024

WayIL: Image-based Indoor Localization with Wayfinding Maps

ICRA 2024poster

This paper tackles a localization problem in large-scale indoor environments with wayfinding maps. A wayfinding map abstractly portrays the environment, and humans can localize themselves based on the map. However, when it comes to using it for robot localization, large geometrical discrepancies bet…

Cited by 3SourcecodeScholar
2023

Dual Variable Actor-Critic for Adaptive Safe Reinforcement Learning

IROS 2023poster

Satisfying safety constraints in reinforcement learning (RL) is an important issue, especially in real-world applications. Many studies have approached safe RL with the Lagrangian method, which introduces dual variables. However, applying a trained policy with the optimal dual variable to a new envi…

Cited by 1SourceScholar
2023

Meta-Explore: Exploratory Hierarchical Vision-and-Language Navigation Using Scene Object Spectrum Grounding

CVPR 2023poster

The main challenge in vision-and-language navigation (VLN) is how to understand natural-language instructions in an unseen environment. The main limitation of conventional VLN algorithms is that if an action is mistaken, the agent fails to follow the instructions or explores unnecessary regions, lea…

Cited by 20SourcePDFScholar
2023

Object Rearrangement Planning for Target Retrieval in a Confined Space with Lateral View

IROS 2023poster

In this paper, we perform an object rearrangement task for target retrieval in an environment with a confined space and limited observation directions. The agent must create a collision-free path to bring out the target object by relocating the surrounding objects using the prehensile action, i.e.,…

Cited by 1SourceScholar
2023

RIANet++: Road Graph and Image Attention Networks for Robust Urban Autonomous Driving Under Road Changes

RA-L 2023

The structure of roads plays an important role in designing autonomous driving algorithms. We propose a novel road graph based driving framework, named RIANet++. The proposed framework considers the road structural scene context by incorporating both graphical features of the road and visual informa

Cited by 4SourceScholar
2023

SCAN: Socially-Aware Navigation Using Monte Carlo Tree Search

ICRA 2023poster

Designing a socially-aware navigation method for crowded environments has become a critical issue in robotics. In order to perform navigation in a crowded environment without causing discomfort to nearby pedestrians, it is necessary to design a global planner that is able to consider both human-robo…

Cited by 7SourceScholar
2023

SDF-Based Graph Convolutional Q-Networks for Rearrangement of Multiple Objects

ICRA 2023poster

In this paper, we propose a signed distance field (SDF)-based deep Q-learning framework for multi-object re-arrangement. Our method learns to rearrange objects with non-prehensile manipulation, e.g., pushing, in unstructured environments. To reliably estimate Q-values in various scenes, we train the…

Cited by 3SourceScholar
2023

Sequential Preference Ranking for Efficient Reinforcement Learning from Human Feedback

NeurIPS 2023poster

Reinforcement learning from human feedback (RLHF) alleviates the problem of designing a task-specific reward function in reinforcement learning by learning it from human preference. However, existing RLHF models are considered inefficient as they produce only a single preference data from each human…

Cited by 11SourcePDFScholar
2023

Trust Region-Based Safe Distributional Reinforcement Learning for Multiple Constraints

NeurIPS 2023poster

In safety-critical robotic tasks, potential failures must be reduced, and multiple constraints must be met, such as avoiding collisions, limiting energy consumption, and maintaining balance. Thus, applying safe reinforcement learning (RL) in such robotic tasks requires to handle multiple constraints…

2022

Dynamics-Aware Metric Embedding: Metric Learning in a Latent Space for Visual Planning

RA-L 2022

In this letter, we consider vision-based control tasks of which the desired goals are given as target images. The problems are often addressed by an autonomous agent which optimizes a trajectory to minimize a manually designed cost function. However, it is challenging to design a suitable cost funct

Cited by 3SourceScholar
2022

Grasp Planning for Occluded Objects in a Confined Space with Lateral View Using Monte Carlo Tree Search

IROS 2022poster

In the lateral access environment, the robot be-havior should be planned considering surrounding objects and obstacles because object observation directions and approach angles are limited. To safely retrieve a partially occluded target object in these environments, we have to relocate objects using…

Cited by 6SourceScholar
2022

RIANet: Road Graph and Image Attention Network for Urban Autonomous Driving

IROS 2022poster

In this paper, we present a novel autonomous driving framework, called a road graph and image attention network (RIANet), which computes the attention scores of objects in the image using the road graph feature. The process of the proposed method is as follows: First, the feature encoder module enco…

Cited by 1SourceScholar
2022

Texture Generation Using Dual-Domain Feature Flow with Multi-View Hallucinations

AAAI 2022technical

We propose a dual-domain generative model to estimate a texture map from a single image for colorizing a 3D human model. When estimating a texture map, a single image is insufficient as it reveals only one facet of a 3D object. To provide sufficient information for estimating a complete texture map,…

Cited by 2SourcePDFScholar
2022

Topological Semantic Graph Memory for Image-Goal Navigation

CoRL 2022oral

A novel framework is proposed to incrementally collect landmark-based graph memory and use the collected memory for image goal navigation. Given a target image to search, an embodied robot utilizes semantic memory to find the target in an unknown environment. In this paper, we present a topological…

Cited by 58SourceScholar
2022

Towards Defensive Autonomous Driving: Collecting and Probing Driving Demonstrations of Mixed Qualities

IROS 2022poster

Designing or learning an autonomous driving policy is undoubtedly a challenging task as the policy has to maintain its safety in all corner cases. In order to secure safety in autonomous driving, the ability to detect hazardous situations, which can be seen as an out-of-distribution (OOD) detection…

Cited by 2SourcecodeScholar
2022

Unsupervised 3D Link Segmentation of Articulated Objects With a Mixture of Coherent Point Drift

RA-L 2022

In this letter, we address the 3D link segmentation problem of articulated objects using multiple point sets with different configurations. We are motivated by the fact that a point set of an object can be aligned to point sets with different configurations by applying rigid transformations to links

Cited by 1SourceScholar
2022

Visually Grounding Language Instruction for History-Dependent Manipulation

ICRA 2022poster

This paper emphasizes the importance of a robot's ability to refer to its task history, especially when it exe-cutes a series of pick-and-place manipulations by following language instructions given one by one. The advantage of referring to the manipulation history can be categorized into two folds:…

Cited by 7SourceScholar
2021

Visual Graph Memory With Unsupervised Representation for Visual Navigation

ICCV 2021poster

We present a novel graph-structured memory for visual navigation, called visual graph memory (VGM), which consists of unsupervised image representations obtained from navigation history. The proposed VGM is constructed incrementally based on the similarities among the unsupervised representations of…

Cited by 81PDFcodeScholar
2020

Generalized Tsallis Entropy Reinforcement Learning and Its Application to Soft Mobile Robots

RSS 2020poster

In this paper, we present a new class of Markov decision processes (MDPs), called Tsallis MDPs, with Tsallis entropy maximization, which generalizes existing maximum entropy reinforcement learning (RL). A Tsallis MDP provides a unified framework for the original RL problem and RL with various types…

2020

Hierarchical 6-DoF Grasping with Approaching Direction Selection

ICRA 2020poster

In this paper, we tackle the problem of 6-DoF grasp detection which is crucial for robot grasping in cluttered real-world scenes. Unlike existing approaches which synthesize 6-DoF grasp data sets and train grasp quality networks with input grasp representations based on point clouds, we rather take…

Cited by 8SourceScholar
2020

Learning to Walk a Tripod Mobile Robot Using Nonlinear Soft Vibration Actuators With Entropy Adaptive Reinforcement Learning

RA-L 2020

Soft mobile robots have shown great potential in unstructured and confined environments by taking advantage of their excellent adaptability and high dexterity. However, there are several issues to be addressed, such as actuating speeds and controllability, in soft robots. In this letter, a new vibra

Cited by 17SourceScholar
2020

MixGAIL: Autonomous Driving Using Demonstrations with Mixed Qualities

IROS 2020poster

In this paper, we consider autonomous driving of a vehicle using imitation learning. Generative adversarial imitation learning (GAIL) is a widely used algorithm for imitation learning. This algorithm leverages positive demonstrations to imitate the behavior of an expert. In this paper, we propose a…

Cited by 25SourceScholar
2020

No-Regret Shannon Entropy Regularized Neural Contextual Bandit Online Learning for Robotic Grasping

IROS 2020poster

In this paper, we propose a novel contextual bandit algorithm that employs a neural network as a reward estimator and utilizes Shannon entropy regularization to encourage exploration, which is called Shannon entropy regularized neural contextual bandits (SERN). In many learning-based algorithms for…

Cited by 2SourceScholar
2020

Pedestrian Intention Prediction for Autonomous Driving Using a Multiple Stakeholder Perspective Model

IROS 2020poster

This paper proposes a multiple stakeholder perspective model (MSPM) which predicts the future pedestrian trajectory observed from vehicle's point of view. For the vehicle-pedestrian interaction, the estimation of the pedestrian's intention is a key factor. However, even if this interaction is common…

Cited by 19SourceScholar
2019

Deep Predictive Autonomous Driving Using Multi-Agent Joint Trajectory Prediction and Traffic Rules

IROS 2019poster

Autonomous driving is a challenging problem because the autonomous vehicle must understand complex and dynamic environment. This understanding consists of predicting future behavior of nearby vehicles and recognizing predefined rules. It is observed that not all rules have equivalent values, and the…

Cited by 48SourceScholar
2019

Deep Virtual Networks for Memory Efficient Inference of Multiple Tasks

CVPR 2019poster

Deep networks consume a large amount of memory by their nature. A natural question arises can we reduce that memory requirement whilst maintaining performance. In particular, in this work we address the problem of memory efficient learning for multiple tasks. To this end, we propose a novel network…

Cited by 12PDFScholar
2018

Interactive Text2Pickup Networks for Natural Language-Based Human-Robot Collaboration

RA-L 2018

In this letter, we propose the Interactive Text2Pickup (IT2P) network for human-robot collaboration that enables an effective interaction with a human user despite the ambiguity in user's commands. We focus on the task where a robot is expected to pick up an object instructed by a human, and to inte

Cited by 26SourceScholar
2018

Sparse Markov Decision Processes With Causal Sparse Tsallis Entropy Regularization for Reinforcement Learning

RA-L 2018

In this letter, a sparse Markov decision process (MDP) with novel causal sparse Tsallis entropy regularization is proposed. The proposed policy regularization induces a sparse and multimodal optimal policy distribution of a sparse MDP. The full mathematical analysis of the proposed sparse MDP is pro

Cited by 70SourceScholar
2018

Text2Action: Generative Adversarial Synthesis from Language to Action

ICRA 2018poster

In this paper, we propose a generative model which learns the relationship between language and human action in order to generate a human action sequence given a sentence describing human behavior. The proposed generative model is a generative adversarial network (GAN), which is based on the sequenc…

Cited by 186SourceScholar
2018

Uncertainty-Aware Learning from Demonstration Using Mixture Density Networks with Sampling-Free Variance Modeling

ICRA 2018poster

In this paper, we propose an uncertainty-aware learning from demonstration method by presenting a novel uncertainty estimation method utilizing a mixture density network appropriate for modeling complex and noisy human behaviors. The proposed uncertainty acquisition can be done with a single forward…

Cited by 138SourceScholar
2018

Unsupervised holistic image generation from key local patches

ECCV 2018poster

We introduce a new problem of generating an image based on a small number of key local patches without any geometric prior. In this work, key local patches are defined as informative regions of the target object or scene. This is a challenging problem since it requires generating realistic images an…

Cited by 16SourcePDFScholar
2017

Scalable robust learning from demonstration with leveraged deep neural networks

IROS 2017poster

In this paper, we propose a novel algorithm for learning from demonstration, which can learn a policy function robustly from a large number of demonstrations with mixed qualities. While most of the existing approaches assume that demonstrations are collected from skillful experts, the proposed metho…

Cited by 4SourceScholar
2016

Robust learning from demonstration using leveraged Gaussian processes and sparse-constrained optimization

ICRA 2016

In this paper, we propose a novel method for robust learning from demonstration using leveraged Gaussian process regression. While existing learning from demonstration (LfD) algorithms assume that demonstrations are given from skillful experts, the proposed method alleviates such assumption by allow

Cited by 30SourceScholar
2016

Robust modeling and prediction in dynamic environments using recurrent flow networks

IROS 2016poster

To enable safe motion planning in a dynamic environment, it is vital to anticipate and predict object movements. In practice, however, an accurate object identification among multiple moving objects is extremely challenging, making it infeasible to accurately track and predict individual objects. Fu…

Cited by 5SourceScholar
2015

Leveraged non-stationary Gaussian process regression for autonomous robot navigation

ICRA 2015poster

In this paper, we propose a novel regression method that can incorporate both positive and negative training data into a single regression framework. In detail, a leveraged kernel function for non-stationary Gaussian process regression is proposed. With this new kernel function, we can vary the corr…

Cited by 14SourceScholar
2015

Structured low-rank matrix approximation in Gaussian process regression for autonomous robot navigation

ICRA 2015poster

This paper considers the problem of approximating a kernel matrix in an autoregressive Gaussian process regression (AR-GP) in the presence of measurement noises or natural errors for modeling complex motions of pedestrians in a crowded environment. While a number of methods have been proposed to rob…

Cited by 4SourceScholar