← Search

David Hsu

60 accepted papers

2026

10 Open Challenges Steering the Future of Vision-Language-Action Models

AAAI 2026technical

Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly preva- lent in the embodied AI arena, following the widespread suc- cess of their precursors—LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing develop- men

Cited by 3SourcePDFScholar
2026

CANINE: Coaching Visually Impaired Users for Interactive Navigation with a Robot Guide Dog

RSS 2026poster

Robot guide dogs offer navigation assistance that will greatly expand the independent mobility of the visually impaired, but their effective use requires subtle human-robot coordination that is difficult for users to learn from generic verbal instructions. To tackle the challenge, we present CANINE,…

Cited by 0SourceScholar
2026

Differentiable Contact Dynamics for Stable Object Placement Under Geometric Uncertainties

RA-L 2026

From serving a cup of coffee to positioning mechanical parts during assembly, stable object placement is a crucial skill for future robots. It becomes particularly challenging under geometric uncertainties, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xl

Cited by 2SourceScholar
2026

Differentiable Contact Dynamics for Stable Object Placement under Geometric Uncertainties

ICRA 2026poster

From stacking a tower of blocks to serving a cup of coffee, stable object placement is a crucial skill for future robots. It becomes particularly challenging under geometric uncertainties, e.g., when the object pose or shape is not known accurately. This work leverages a differentiable simulation mo…

2026

From Obstacles to Etiquette: Robot Social Navigation With VLM-Informed Path Selection

RA-L 2026

Navigating socially in human environments requires more than satisfying geometric constraints, as collision-free paths may still interfere with ongoing activities or conflict with social norms. Addressing this challenge calls for analyzing interactions between agents and incorporating common-sense r

Cited by 1SourcecodeScholar
2026

Hypothesis-driven Model Expansion under Uncertainty for Open-World Robot Planning

RSS 2026poster

We consider an open-world planning setting in which service robots must operate in unknown environments with incomplete knowledge of objects and actions. Traditional closed-world approaches with pre-programmed knowledge bases fail when robots encounter unexpected situations and tasks, posing a funda…

Cited by 0SourceScholar
2026

SignLoc: Robust Localization Using Navigation Signs and Public Maps

ICRA 2026poster

Navigation signs and maps, such as floor plans and street maps, are widely available and serve as ubiquitous aids for way-finding in human environments. Yet, they are rarely used by robot systems. This paper presents SignLoc, a global localization method that leverages navigation signs to localize t…

2025

MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence

CoRL 2025poster

Imitating tool manipulation from human videos offers an intuitive approach to teaching robots, while also providing a promising and scalable alternative to labor-intensive teleoperation data collection for visuomotor policy learning. While humans can mimic tool manipulation behavior by observing oth…

Cited by 0SourceScholar
2025

Neuralized Markov Random Field for Interaction-Aware Stochastic Human Trajectory Prediction

ICLR 2025poster

Interactive human motions and the continuously changing nature of intentions pose significant challenges for human trajectory prediction. In this paper, we present a neuralized Markov random field (MRF)-based motion evolution method for probabilistic interaction-aware human trajectory prediction. We…

2025

Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

ICRA 2025

Imagine a future when we can Zoom-call a robot to manage household chores remotely. This work takes one step in this direction. Robi Butler is a new household robot assistant that enables seamless multimodal remote interaction. It allows the human user to monitor its environment from a first-person

Cited by 6SourceScholar
2025

Walk in Others’ Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective Transformation

ACL 2025long

Visual perspective-taking, an ability to envision others’ perspectives from a single self-perspective, is vital in human-robot interactions. Thus, we introduce a human-centric visual grounding task and a dataset to evaluate this ability. Recent advances in vision-language models (VLMs) have shown po…

2025

“Stack It Up!”: 3D Stable Structure Generation from 2D Hand-drawn Sketch

CoRL 2025oral

Imagine a child sketching the Eiffel Tower and asking a robot to bring it to life. Today’s robot manipulation systems can’t act on such sketches directly—they require precise 3D block poses as goals, which in turn demand structural analysis and expert tools like CAD. We present *StackItUp*, a system…

Cited by 0SourceScholar
2024

On the Empirical Complexity of Reasoning and Planning in LLMs

EMNLP 2024finding

Chain-of-thought (CoT), tree-of-thought (ToT), and related techniques work surprisingly well in practice for some complex reasoning tasks with Large Language Models (LLMs), but why? This work seeks the underlying reasons by conducting experimental case studies and linking the performance benefits to…

2024

Set It Up!: Functional Object Arrangement with Compositional Generative Models

RSS 2024poster

This paper studies the challenge of developing robots capable of understanding under-specified instructions for creating functional object arrangements, such as "set up a dining table for two"; previous arrangement approaches have focused on much more explicit instructions, such as "put object A on…

2023

DaxBench: Benchmarking Deformable Object Manipulation with Differentiable Physics

ICLR 2023top-5%

Deformable object manipulation (DOM) is a long-standing challenge in robotics and has attracted significant interest recently. This paper presents DaXBench, a differentiable simulation framework for DOM. While existing work often focuses on a specific type of deformable objects, DaXBench supports fl…

2023

Differentiable Parsing and Visual Grounding of Natural Language Instructions for Object Placement

ICRA 2023poster

We present a new method, PARsing And visual GrOuNding (PARAGON), for grounding natural language in object placement tasks. Natural language generally describes objects and spatial relations with compositionality and ambiguity, two major obstacles to effective language grounding. For compositionality…

Cited by 11SourceScholar
2023

What Truly Matters in Trajectory Prediction for Autonomous Driving?

NeurIPS 2023poster

Trajectory prediction plays a vital role in the performance of autonomous driving systems, and prediction accuracy, such as average displacement error (ADE) or final displacement error (FDE), is widely used as a performance metric. However, a significant disparity exists between the accuracy of pred…

2022

LEADER: Learning Attention over Driving Behaviors for Planning under Uncertainty

CoRL 2022oral

Uncertainty in human behaviors poses a significant challenge to autonomous driving in crowded urban environments. The partially observable Markov decision process (POMDP) offers a principled general framework for decision making under uncertainty and achieves real-time performance for complex tasks…

Cited by 12SourcecodeScholar
2021

Interactive Planning for Autonomous Urban Driving in Adversarial Scenarios

ICRA 2021poster

Autonomous urban driving among human-driven cars requires a holistic understanding of road rules, driver intents and driving styles. This is challenging as a short-term, single instance, driver intent of lane change may not correspond to their driving styles for a longer duration. This paper present…

Cited by 13SourceScholar
2020

Contrastive Variational Reinforcement Learning for Complex Observations

CoRL 2020

Deep reinforcement learning (DRL) has achieved significant success in various robot tasks: manipulation, navigation, etc. However, complex visual observations in natural environments remains a major challenge. This paper presents Contrastive Variational Reinforcement Learning (CVRL), a model-based m

2020

Discriminative Particle Filter Reinforcement Learning for Complex Partial observations

ICLR 2020poster

Deep reinforcement learning is successful in decision making for sophisticated games, such as Atari, Go, etc. However, real-world decision making often requires reasoning with partial information extracted from complex visual observations. This paper presents Discriminative Particle Filter Reinfor…

Cited by 45SourcecodeScholar
2019

Context and Intention Aware Planning for Urban Driving

IROS 2019poster

We present a novel autonomous driving system which uses the road contextual information and intentions of other road users for urban driving. Unlike highways, urban environments require the drivers to follow traffic signs and signals while using their best judgment for anomalous situations. In such…

Cited by 25SourceScholar
2019

DESPOT-Alpha: Online POMDP Planning with Large State and Observation Spaces

RSS 2019poster

State-of-the-art sampling-based online POMDP solvers compute near-optimal policies for POMDPs with very large state spaces. However, when faced with large observation spaces, they may become overly optimistic and compute sub-optimal policies, because of particle divergence. This paper presents a new…

Cited by 73SourcePDFScholar
2019

Differentiable Algorithm Networks for Composable Robot Learning

RSS 2019poster

This paper introduces the Differentiable Algorithm Network (DAN), a composable architecture for robot learning systems. A DAN is composed of neural network modules, each encoding a differentiable robot algorithm and an associated model; and it is trained end-to-end from data. DAN combines the streng…

Cited by 81SourcePDFScholar
2019

Factored Contextual Policy Search with Bayesian optimization

ICRA 2019poster

Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning and generalization by explicitly conditioning the policy on a parametric context…

Cited by 8SourceScholar
2019

LeTS-Drive: Driving in a Crowd by Learning from Tree Search

RSS 2019poster

Autonomous driving in a crowded environment, e.g., a busy traffic intersection, is an unsolved challenge for robotics. The robot vehicle must contend with a dynamic and partially observable environment, noisy sensors, and many agents. A principled approach is to formalize it as a Partially Observabl…

Cited by 40SourcePDFScholar
2018

HyP-DESPOT: A Hybrid Parallel Algorithm for Online Planning under Uncertainty

RSS 2018poster

Planning under uncertainty is critical for robust robot performance in uncertain, dynamic environments, but it incurs high computational cost. State-of-the-art online search algorithms, such as DESPOT, have vastly improved the computational efficiency of planning under uncertainty and made it a valu…

2018

PORCA: Modeling and Planning for Autonomous Driving Among Many Pedestrians

RA-L 2018

This letter presents a planning system for autonomous driving among many pedestrians. A key ingredient of our approach is Pedestrian Optimal Reciprocal Collision Avoidance, a pedestrian motion prediction model that accounts for both a pedestrian's global navigation intention and local interactions w

Cited by 191SourceScholar
2017

Intention-Net: Integrating Planning and Deep Learning for Goal-Directed Autonomous Navigation

CoRL 2017

How can a delivery robot navigate reliably to a destination in a new office building, with minimal prior information? To tackle this challenge, this paper introduces a two-level hierarchical approach, which integrates model-free deep learning and model-based path planning. At the low level, a neural

Cited by 0SourcePDFScholar
2017

QMDP-Net: Deep Learning for Planning under Partial Observability

NeurIPS 2017poster

This paper introduces the QMDP-net, a neural network architecture for planning under partial observability. The QMDP-net combines the strengths of model-free learning and model-based planning. It is a recurrent policy network, but it represents a policy for a parameterized set of tasks by connecting…

2017

XPose: Reinventing User Interaction with Flying Cameras

RSS 2017poster

XPose is a new touch-based interactive system for photo taking, designed to take advantage of the autonomous flying capability of a drone-mounted camera. It enables the user to interact with photos directly and focus on taking photos instead of piloting the drone. XPose introduces a two-stage eXplor…

2015

Intention-aware online POMDP planning for autonomous driving in a crowd

ICRA 2015poster

This paper presents an intention-aware online planning approach for autonomous driving amid many pedestrians. To drive near pedestrians safely, efficiently, and smoothly, autonomous vehicles must estimate unknown pedestrian intentions and hedge against the uncertainty in intention estimates in order…

Cited by 432SourceScholar
2015

Towards autonomous navigation of unsignalized intersections under uncertainty of human driver intent

IROS 2015poster

In a mixed environment of autonomous driverless vehicles and human driven vehicles operating on the same road, identifying intentions of human drivers and interacting with them in a compliant and responsible manner becomes a challenging problem for the driverless vehicles. In this paper, the problem…

Cited by 86SourceScholar