← Search

Liang Han

12 accepted papers

2026

PIDiff: Integrating a High-Performance Transformer Into Diffusion Models for Robust and Efficient Imitation Learning

RA-L 2026

Imitation learning is a critical approach for robots to acquire skills by mimicking human behavior. However, traditional imitation learning frameworks often exhibit poor action prediction accuracy and low robustness when handling complex tasks. To tackle these limitations, we propose the PIDiff poli

Cited by 1SourceScholar
2026

VGGS: VGGT-guided Gaussian Splatting for Efficient and Faithful Sparse-View Surface Reconstruction

AAAI 2026technical

Reconstructing a faithful geometric surface from sparse images remains a fundamental challenge in 3D computer vision. While recent methods have achieved remarkable progress, they still struggle to recover reliable geometry due to the lack of multi-view geometric cues, particularly in non-overlapping

Cited by 0SourcePDFScholar
2025

CushionCatch: A Compliant Catching Mechanism for Mobile Manipulators via Combined Optimization and Learning

IROS 2025

Catching flying objects with a cushioning process is a skill commonly performed by humans, yet it remains a significant challenge for robots. In this paper, we present a framework that combines optimization and learning to achieve compliant catching on mobile manipulators (CCMM). First, we propose a

Cited by 1SourceScholar
2025

MonoInstance: Enhancing Monocular Priors via Multi-view Instance Alignment for Neural Rendering and Reconstruction

CVPR 2025poster

Monocular depth priors have been widely adopted by neural rendering in multi-view based tasks such as 3D reconstruction and novel view synthesis. However, due to the inconsistent prediction on each view, how to more effectively leverage monocular cues in a multi-view context remains a challenge. Cur…

Cited by 4SourcePDFScholar
2025

SparseRecon: Neural Implicit Surface Reconstruction from Sparse Views with Feature and Depth Consistencies

ICCV 2025poster

Surface reconstruction from sparse views aims to reconstruct a 3D shape or scene from few RGB images. The latest methods are either generalization-based or overfitting-based. However, the generalization-based methods do not generalize well on views that were unseen during training, while the reconst…

2025

UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces

NeurIPS 2025poster

Building a universal trajectory foundation model is a promising solution to address the limitations of existing trajectory modeling approaches, such as task specificity, regional dependency, and data sensitivity. Despite its potential, data preparation, pre-training strategy development, and archite…

Cited by 0SourcecodeScholar
2024

Binocular-Guided 3D Gaussian Splatting with View Consistency for Sparse View Synthesis

NeurIPS 2024poster

Novel view synthesis from sparse inputs is a vital yet challenging task in 3D computer vision. Previous methods explore 3D Gaussian Splatting with neural priors (e.g. depth priors) as an additional supervision, demonstrating promising quality and efficiency compared to the NeRF based methods. Howeve…

Cited by 8SourcePDFScholar
2023

Event-Triggered Optimal Formation Tracking Control Using Reinforcement Learning for Large-Scale UAV Systems

ICRA 2023poster

Large-scale UAV switching formation tracking control has been widely applied in many fields such as search and rescue, cooperative transportation, and UAV light shows. In order to optimize the control performance and reduce the computational burden of the system, this study proposes an event-trigger…

Cited by 3SourceScholar
2018

A Lightweight Redundant Manipulator with High Stable Wireless Communication and Compliance Control

IROS 2018poster

For traditional manipulators, there is a large number of electrical cables between the motion controller and the joint servo controllers. It is very inconvenient for maintenance, update, and safe operation. In this paper, we develop a lightweight redundant manipulator with high stable wireless commu…

Cited by 6SourceScholar
2017

Cross-modality matching based on Fisher Vector with neural word embeddings and deep image features

ICASSP 2017accepted

Cross-modal retrieval, which aims to solve the problem that the query and the retrieved results are from different modality, becomes more and more essential with the development of the Internet. In this paper, we mainly focus on the exploration of high-level semantic representation of image and text…

Cited by 0SourceScholar