← Search

Pinhao Song

9 accepted papers

2026

Mini Diffuser: Fast Multi-Task Diffusion Policy Training Using Two-Level Mini-Batches

RA-L 2026

We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from a previously underexplored distinction between action diffusion and the image diffusion techniques that inspired it: In

Cited by 0SourcecodeScholar
2026

Mini Diffuser: Fast Multi-Task Diffusion Policy Training Using Two-Level Mini-Batches

ICRA 2026poster

We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from a previously underexplored distinction between action diffusion and the image diffusion techniques that inspired it: In …

2025

See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model

NeurIPS 2025poster

We introduce See&Trek, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMs) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely…

Cited by 0SourceScholar
2024

Implicit Grasp Diffusion: Bridging the Gap between Dense Prediction and Sampling-based Grasping

CoRL 2024poster

There are two dominant approaches in modern robot grasp planning: dense prediction and sampling-based methods. Dense prediction calculates viable grasps across the robot’s view but is limited to predicting one grasp per voxel. Sampling-based methods, on the other hand, encode multi-modal grasp distr…

Cited by 2SourceScholar
2024

OTOcc: Optimal Transport for Occupancy Prediction

IJCAI 2024poster

The autonomous driving community is highly interested in 3D occupancy prediction due to its outstanding geometric perception and object recognition capabilities. However, previous methods are limited to existing semantic conversion mechanisms for solving sparse ground truths problem, causing excessi…

2024

Robot Trajectron: Trajectory Prediction-based Shared Control for Robot Manipulation

ICRA 2024poster

We address the problem of (a) predicting the trajectory of an arm reaching motion, based on a few seconds of the motion’s onset, and (b) leveraging this predictor to facilitate shared-control manipulation tasks, by reducing the operator’s cognitive load through assistance in their anticipated direct…

Cited by 24SourcecodeScholar
2023

Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes

ICASSP 2023accepted

Generic object detection methods have achieved preferable results, but it is still challenging to detect objects from complicated traffic scenes like extreme illumination and adverse weather. The existing methods are not robust enough to be extended to new complex traffic scenes. To address this iss…

Cited by 0SourceScholar
2022

Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on Transformer

AAAI 2022technical

Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based method…