← Search

Harold Soh

48 accepted papers

2026

Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion Policies

ICLR 2026poster

Diffusion-based policies have recently shown strong results in robot manipulation, but their extension to multi-task scenarios is hindered by the high cost of scaling model size and demonstrations. We introduce Skill Mixture-of-Experts Policy (SMP), a diffusion-based mixture-of-experts policy that l…

Cited by 0SourceScholar
2026

Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards

ICML 2026poster

Inference-time guided sampling steers state-of-the-art diffusion and flow models without fine-tuning by interpreting the generation process as a controllable trajectory. This provides a simple and flexible way to inject external constraints (e.g., cost functions or pre-trained verifiers) for control…

Cited by 0SourceScholar
2026

DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting

ICRA 2026poster

Diffusion policies have demonstrated strong performance in generative modeling, making them promising for robotic manipulation guided by natural language instructions. However, generalizing language-conditioned diffusion policies to open-vocabulary instructions in everyday scenarios remains challeng…

2026

GENIE: A Generalizable Navigation System for In-The-Wild Environments

ICRA 2026poster

Reliable navigation in unstructured, real-world environments remains a significant challenge for embodied agents, especially when operating across diverse terrains, weather conditions, and sensor configurations. In this paper, we introduce GeNIE (Generalizable Navigation System for In-the-Wild Envir…

2026

SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse

RSS 2026poster

Recent progress in vision–language–action (VLA) models has demonstrated strong potential for dual-arm manipulation, enabling complex behaviors and generalization to unseen environments. However, mainstream bimanual VLA formulations largely overlook the critical challenge of combinatorial diversity. …

Cited by 0SourceScholar
2026

TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing

ICRA 2026poster

Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standardized evaluation metrics, datasets, and protocols. Existing systems are evaluated in different environments under different criteria, preventing fair and reprodu…

2025

Arena 4.0: a Comprehensive Ros2 Development and Benchmarking Platform for Human-Centric Navigation Using Generative-Model-Based Environment Generation

ICRA 2025

Building upon the foundations laid by our previous work, this paper introduces Arena 4.0, a significant advancement of Arena 3.0 [1], Arena-Bench [2], Arena 1.0 [3], and Arena 2.0 [4]. Arena 4.0 provides three main novel contributions: 1) a generative-model-based world and scenario generation approa

Cited by 12SourceScholar
2025

DISCO: Language-Guided Manipulation With Diffusion Policies and Constrained Inpainting

RA-L 2025

Diffusion policies have demonstrated strong performance in generative modeling, making them promising for robotic manipulation guided by natural language instructions. However, generalizing language-conditioned diffusion policies to open-vocabulary instructions in everyday scenarios remains challeng

Cited by 8SourceScholar
2025

Demonstrating Arena 5.0: A Photorealistic ROS2 Simulation Framework for Developing and Benchmarking Social Navigation

RSS 2025poster

Building upon the foundations laid by our previous work, this paper introducesArena 5.0, the fifth iteration of our framework for robotics social navigation development and benchmarking. Arena 5.0 provides three main contributions: 1) The complete integration of NVIDIA Isaac Gym, enabling photoreali…

Cited by 0PDFScholar
2025

Diffusion Meets Options: Hierarchical Generative Skill Composition for Temporally-Extended Tasks

ICRA 2025

Safe and successful deployment of robots requires not only the ability to generate complex plans but also the capacity to frequently replan and correct execution errors. This paper addresses the challenge of long-horizon trajectory planning under temporally extended objectives in a receding horizon

Cited by 9SourceScholar
2025

Imitation Learning with Limited Actions via Diffusion Planners and Deep Koopman Controllers

ICRA 2025

Recent advances in diffusion-based robot policies have demonstrated significant potential in imitating multi-modal behaviors. However, these approaches typically require large quantities of demonstration data paired with corresponding robot action labels, creating a substantial data collection burde

Cited by 4SourcecodeScholar
2025

NUSense: Shear Based Robust Optical Tactile Sensor

IROS 2025

While most optical tactile sensors rely on measuring surface displacement, insights from continuum mechanics suggest that measuring shear strain provides key information for tactile sensing. In this work, we introduce an optical tactile sensing principle based on shear strain detection. A silicone r

Cited by 0SourceScholar
2025

OpenRoboCare: A Multimodal Multi-Task Expert Demonstration Dataset for Robot Caregiving

IROS 2025

We present OpenRoboCare, a multimodal dataset for robot caregiving, capturing expert occupational therapist demonstrations of Activities of Daily Living (ADLs). Caregiving tasks involve complex physical human-robot interactions, requiring precise perception under occlusions, safe physical contact, a

Cited by 3SourceScholar
2024

Demonstrating Arena 3.0: Advancing Social Navigation in Collaborative and Highly Dynamic Environments

RSS 2024poster

Building upon our previous contributions, this paper introduces Arena 3.0, an extension of Arena-Bench, Arena 1.0, and Arena 2.0 focusing on the development, simulation, and benchmarking of social navigation approaches in collaborative environments. We significantly enhance the realism of human beha…

2024

Don't Start From Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion

RSS 2024poster

Imitation learning empowers artificial agents to mimic behavior by learning from demonstrations. Recently, diffusion models, which have the ability to model high-dimensional and multimodal distributions, have shown impressive performance on imitation learning tasks. These models learn to shape a pol…

2024

Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction

EMNLP 2024main

In this work, we are interested in automated methods for knowledge graph creation (KGC) from input text. Progress on large language models (LLMs) has prompted a series of recent works applying them to KGC, e.g., via zero/few-shot prompting. Despite successes on small domain-specific datasets, these…

2024

GRaCE: Balancing Multiple Criteria to Achieve Stable, Collision-Free, and Functional Grasps

RSS 2024poster

Abstract—This paper addresses the multi-faceted problem of robot grasping, where multiple criteria may conflict and differ in importance. We introduce a probabilistic framework, Grasp Ranking and Criteria Evaluation (GRaCE), which employs hierarchical rule-based logic and a rank-preserving utility f…

2024

LTLDoG: Satisfying Temporally-Extended Symbolic Constraints for Safe Diffusion-Based Planning

RA-L 2024

Operating effectively in complex environments while complying with specified constraints is crucial for the safe and successful deployment of robots that interact with and operate around people. In this letter, we focus on generating long-horizon trajectories that adhere to static and temporally-ext

Cited by 14SourcecodeScholar
2024

Octopi: Object Property Reasoning with Large Tactile-Language Models

RSS 2024poster

Physical reasoning is important for effective robot manipulation. Recent work has investigated both vision and language modalities for physical reasoning; vision can reveal information about objects in the environment and language serves as an abstraction and communication medium for additional cont…

2024

Out-of-Distribution Detection with a Single Unconditional Diffusion Model

NeurIPS 2024poster

Out-of-distribution (OOD) detection is a critical task in machine learning that seeks to identify abnormal samples. Traditionally, unsupervised methods utilize a deep generative model for OOD detection. However, such approaches require a new model to be trained for each inlier dataset. This paper ex…

2023

Latent Emission-Augmented Perspective-Taking (LEAPT) for Human-Robot Interaction

IROS 2023poster

Perspective-taking is the ability to perceive or understand a situation or concept from another individual's point of view, and is crucial in daily human interactions. Enabling robots to perform perspective-taking remains an unsolved problem; existing approaches that use deterministic or handcrafted…

Cited by 0SourceScholar
2023

Neural Continuous-Discrete State Space Models for Irregularly-Sampled Time Series

ICML 2023oral

Learning accurate predictive models of real-world dynamic phenomena (e.g., climate, biological) remains a challenging task. One key issue is that the data generated by both natural and artificial processes often comprise time series that are irregularly sampled and/or contain missing observations. I…

2023

Refining 6-DoF Grasps with Context-Specific Classifiers

IROS 2023poster

In this work, we present GraspFlow, a refinement approach for generating context-specific grasps. We formulate the problem of grasp synthesis as a sampling problem: we seek to sample from a context-conditioned probability distribution of successful grasps. However, this target distribution is unknow…

Cited by 2SourcecodeScholar
2023

Selective Amnesia: A Continual Learning Approach to Forgetting in Deep Generative Models

NeurIPS 2023spotlight

The recent proliferation of large-scale text-to-image models has led to growing concerns that such models may be misused to generate harmful, misleading, and inappropriate content. Motivated by this issue, we derive a technique inspired by continual learning to selectively forget concepts in pretrai…

2023

The Best of Both Worlds in Network Population Games: Reaching Consensus and Convergence to Equilibrium

NeurIPS 2023poster

Reaching consensus and convergence to equilibrium are two major challenges of multi-agent systems. Although each has attracted significant attention, relatively few studies address both challenges at the same time. This paper examines the connection between the notions of consensus and equilibrium i…

Cited by 6SourcePDFScholar
2022

MIRROR: Differentiable Deep Social Projection for Assistive Human-Robot Communication

RSS 2022poster

Communication is a hallmark of intelligence. In this work, we present MIRROR, an approach to (i) quickly learn human models from human demonstrations, and (ii) use the models for subsequent communication planning in assistive shared-control settings. MIRROR is inspired by social projection theory, w…

2021

Deep Explicit Duration Switching Models for Time Series

NeurIPS 2021poster

Many complex time series can be effectively subdivided into distinct regimes that exhibit persistent dynamics. Discovering the switching behavior and the statistical patterns in these regimes is important for understanding the underlying dynamical system. We propose the Recurrent Explicit Duration S…

2021

Extended Tactile Perception: Vibration Sensing through Tools and Grasped Objects

IROS 2021poster

Humans display the remarkable ability to sense the world through tools and other held objects. For example, we are able to pinpoint impact locations on a held rod and tell apart different textures using a rigid probe. In this work, we consider how we can enable robots to have a similar capacity, i.e…

Cited by 26SourcecodeScholar
2021

Multi-Modal Mutual Information (MuMMI) Training for Robust Self-Supervised Deep Reinforcement Learning

ICRA 2021poster

This work focuses on learning useful and robust deep world models using multiple, possibly unreliable, sensors. We find that current methods do not sufficiently encourage a shared representation between modalities; this can cause poor performance on downstream tasks and over-reliance on specific sen…

Cited by 26SourceScholar
2021

Refining Deep Generative Models via Discriminator Gradient Flow

ICLR 2021poster

Deep generative modeling has seen impressive advances in recent years, to the point where it is now commonplace to see simulated samples (e.g., images) that closely resemble real-world data. However, generation quality is generally inconsistent for any given model and can vary dramatically between s…

2020

A Characteristic Function Approach to Deep Implicit Generative Modeling

CVPR 2020oral

Implicit Generative Models (IGMs) such as GANs have emerged as effective data-driven models for generating samples, particularly images. In this paper, we formulate the problem of learning an IGM as minimizing the expected distance between characteristic functions. Specifically, we minimize the dist…

Cited by 46PDFcodeScholar
2020

Efficient Exploration of Reward Functions in Inverse Reinforcement Learning via Bayesian Optimization

NeurIPS 2020poster

The problem of inverse reinforcement learning (IRL) is relevant to a variety of tasks including value alignment and robot learning from demonstration. Despite significant algorithmic contributions in recent years, IRL remains an ill-posed problem at its core; multiple reward functions coincide with…

Cited by 36SourcePDFScholar
2020

Event-Driven Visual-Tactile Sensing and Learning for Robots

RSS 2020poster

This work contributes an event-driven visual-tactile perception system, comprising a novel biologically-inspired tactile sensor and multi-modal spike-based learning. Our neuromorphic fingertip tactile sensor, NeuTouch, scales well with the number of taxels thanks to its event-based nature. Likewise,…

Cited by 125SourcePDFScholar
2020

Fast Texture Classification Using Tactile Neural Coding and Spiking Neural Network

IROS 2020poster

Touch is arguably the most important sensing modality in physical interactions. However, tactile sensing has been largely under-explored in robotics applications owing to the complexity in making perceptual inferences until the recent advancements in machine learning or deep learning in particular.…

Cited by 35SourceScholar
2020

Getting to Know One Another: Calibrating Intent, Capabilities and Trust for Human-Robot Collaboration

IROS 2020poster

Common experience suggests that agents who know each other well are better able to work together. In this work, we address the problem of calibrating intention and capabilities in human-robot collaboration. In particular, we focus on scenarios where the robot is attempting to assist a human who is u…

Cited by 19SourcecodeScholar
2020

TactileSGNet: A Spiking Graph Neural Network for Event-based Tactile Object Recognition

IROS 2020poster

Tactile perception is crucial for a variety of robot tasks including grasping and in-hand manipulation. New advances in flexible, event-driven, electronic skins may soon endow robots with touch perception capabilities similar to humans. These electronic skins respond asynchronously to changes (e.g.,…

Cited by 44SourcecodeScholar
2019

Embedding Symbolic Knowledge into Deep Networks

NeurIPS 2019poster

In this work, we aim to leverage prior symbolic knowledge to improve the performance of deep models. We propose a graph embedding network that projects propositional formulae (and assignments) onto a manifold via an augmented Graph Convolutional Network (GCN). To generate semantically-faithful embed…

2019

Towards Effective Tactile Identification of Textures using a Hybrid Touch Approach

ICRA 2019poster

The sense of touch is arguably the first human sense to develop. Empowering robots with the sense of touch may augment their understanding of interacted objects and the environment beyond standard sensory modalities (e.g., vision). This paper investigates the effect of hybridizing touch and sliding…

Cited by 52SourceScholar