← Search

Huaping Liu

93 accepted papers

2026

A Dual-Mode Electrical Capacitance Tomography Sensor for Robotic Proximity Servoing and Grasping

RSS 2026poster

Tactile and proximity sensing is fundamental for achieving autonomous robotic manipulation and safe human-robot interaction. However, traditional dual-mode sensors often face challenges such as environmental interference and the perception gap between far-field vision and near-field contact. This st…

Cited by 0SourceScholar
2026

A Robust and Efficient Visual-Inertial SLAM Using Hybrid Point-Line Features

RA-L 2026

Visual simultaneous localization and mapping (VSLAM) is a foundational technology in robotics, providing an optimal balance of cost and accuracy. However, existing systems often lack robustness in environments with fast motion, dynamic lighting, or low texture. This letter introduces ML-SLAM, a hybr

Cited by 0SourceScholar
2026

CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human

ICRA 2026poster

In this work, we present CollabVLA, a self-reflective vision-language-action framework that transforms a standard visuomotor policy into a collaborative assistant. CollabVLA tackles key limitations of prior VLAs, including domain overfitting, non-interpretable reasoning, and the high latency of auxi…

2026

FlowDreamer: A RGB-D World Model With Flow-Based Motion Representations for Robot Manipulation

RA-L 2026

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canon

Cited by 12SourcecodeScholar
2026

FlowDreamer: A RGB-D World Model with Flow-Based Motion Representations for Robot Manipulation

ICRA 2026poster

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canon…

2026

OWOD-FSL: Open-World Object Detection Via Few-Shot Learning and Dynamic Prototypes

ICRA 2026poster

Open-World Object Detection (OWOD) presents a critical challenge for modern computer vision systems: detecting known classes, identifying unknown objects, and incrementally learning to recognize them over time. However, current approaches have two fundamental limitations: (1) the fixed-dimensional c…

Cited by 0Scholar
2026

RoboOmni: Actions Are Just Another Modality for Your Vision-Language Models

ICML 2026poster

Integrating Vision-Language Models (VLMs) into robotics has facilitated the development of generalizable Vision-Language Action (VLA) policies. However, unified discrete frameworks lag behind decoupled continuous designs due to limitations in action chunking and temporal modeling. To address this, w…

Cited by 0SourceScholar
2025

A Novel Terrain Classification System with Planar ECT Sensor

IROS 2025

Terrain classification is crucial for robotic navigation especially in unknown environment. Existing terrain classification methods usually have high requirements for environment conditions and robot motions, making them challenging to apply to real-world scenarios. In this paper, we develop a novel

Cited by 0SourceScholar
2025

A Patch-Based Transformer Method for Electrical Capacitance Tomography Image Reconstruction

IROS 2025

Electrical capacitance tomography (ECT) is a contactless and non-invasive imaging technique, which visualizes the internal permittivity distribution around a region utilizing boundary capacitance measurements. It has been widely used in the fields of object classification, tactile sensing and multip

Cited by 0SourceScholar
2025

AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environments

IROS 2025

Current service robots suffer from limited natural language communication abilities, heavy reliance on predefined commands, ongoing human intervention, and, most notably, a lack of proactive collaboration awareness in human-populated environments. This results in narrow applicability and low utility

Cited by 6SourcecodeScholar
2025

Bayesian Morphology Optimization for Musculoskeletal Systems

IROS 2025

In this study, we focus on enhancing the policy of a musculoskeletal arm to develop grasping abilities for objects of varying weights. The agent is modeled using MyoSuite, a platform with realistic biomechanics where muscles drive skeletal movement. We observed that optimizing only the control polic

Cited by 2SourceScholar
2025

Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach

AAAI 2025technical

Graph representation learning methods are highly effective in handling complex non-Euclidean data by capturing intricate relationships and features within graph structures. However, traditional methods face challenges when dealing with heterogeneous graphs that contain various types of nodes and edg…

2025

LLM Enhancers for GNNs: An Analysis from the Perspective of Causal Mechanism Identification

ICML 2025poster

The use of large language models (LLMs) as feature enhancers to optimize node representations, which are then used as inputs for graph neural networks (GNNs), has shown significant potential in graph representation learning. However, the fundamental properties of this approach remain underexplored.…

Cited by 0SourcePDFScholar
2025

Learn to Think: Bootstrapping LLM Logic Through Graph Representation Learning

IJCAI 2025

Large Language Models (LLMs) have achieved remarkable success across various domains. However, they still face significant challenges, including high computational costs for training and limitations in solving complex reasoning problems. Although existing methods have extended the reasoning capabili

2025

Lifelong Morphology Learning for Deformable Embodied Agents

IROS 2025

A deformable agent can continuously adjust its morphology during training, allowing it to discover more suitable structures and outperform fixed-morphology counterparts in terrain-specific tasks. This adaptability is achieved through a joint optimization process consisting of two stages: the Skeleto

Cited by 0SourcecodeScholar
2025

MADI: Malicious Agent Detection and Isolation in Mixed Autonomy Traffic Systems

IROS 2025

Mixed autonomy traffic systems face significant security challenges when malicious agents disrupt coordination between autonomous and human-driven vehicles. We present Malicious Agent Detection and Isolation (MADI), a framework addressing two critical forms of disruptive behavior: path order violati

Cited by 0SourceScholar
2025

MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception

CVPR 2025poster

Advanced driver assistance systems require a comprehensive understanding of the driver's mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper proposes MMTL-UniAD, a unified multi-modal multi-task learning…

2025

MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation

CVPR 2025highlight

In this paper, we introduce MeshGen, an advanced image-to-3D pipeline that generates high-quality 3D meshes with detailed geometry and physically based rendering (PBR) textures. Addressing the challenges faced by existing 3D native diffusion models, such as suboptimal auto-encoder performance, limit…

2025

Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

RA-L 2025

In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this letter, we investigate the problem of robotic manipulation under lim

Cited by 12SourceScholar
2025

ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained Environments

EMNLP 2025

We introduce ProcWorld, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM). ProcWorld features a wide range of challenging embodied navigation and object manipulation tasks, covering 16

Cited by 0SourcePDFScholar
2025

Sparse Spectral Training and Inference on Euclidean and Hyperbolic Neural Networks

ICML 2025poster

The growing demands on GPU memory posed by the increasing number of neural network parameters call for training approaches that are more memory-efficient. Previous memory reduction training techniques, such as Low-Rank Adaptation (LoRA) and ReLoRA, face challenges, with LoRA being constrained by its…

Cited by 1SourcePDFScholar
2025

TEM3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive Driving

IROS 2025

Multi-task learning (MTL) can advance assistive driving by exploring inter-task correlations through shared representations. However, existing methods face two critical limitations: single-modality constraints limiting comprehensive scene understanding and inefficient architectures impeding real-tim

Cited by 3SourcecodeScholar
2024

A Large-area Tactile Sensor for Distributed Force Sensing Using Highly Sensitive Piezoresistive Sponge

ICRA 2024poster

Tactile sensing plays a critical role in enabling robots to interact safely with target objects in dynamic and unstructured environments. While various tactile sensors based on different sensing principles or different sensitive materials have been proposed, the development of flexible large-area ta…

Cited by 1SourceScholar
2024

Automatic Captioning based on Visible and Infrared Images

ICRA 2024poster

In this paper, we tackle the task of image captioning with the complementarity of visible light images and infrared images. To address this problem, we propose an RGBIR image fusion captioning model, which can take full advantage of visible light images and infrared images under different conditions…

Cited by 1SourceScholar
2024

Bionic Soft Fingers with Hybrid Variable Stiffness Mechanisms for Multimode Grasping

ICRA 2024poster

This paper presents a novel Bionic Soft Finger (BSF) that aims to overcome the limitations of conventional rigid manipulators in terms of adaptability and safety, as well as the challenges faced by soft hands regarding carrying capacity and stability. The BSF design uses a hybrid variable stiffness…

Cited by 1SourceScholar
2024

CompetEvo: Towards Morphological Evolution from Competition

IJCAI 2024poster

Training an agent to adapt to specific tasks through co-optimization of morphology and control has widely attracted attention. However, whether there exists an optimal configuration and tactics for agents in a multiagent competition scenario is still an issue that is challenging to definitively conc…

2024

Demonstrating HumanTHOR: A Simulation Platform and Benchmark for Human-Robot Collaboration in a Shared Workspace

RSS 2024poster

Human-robot collaboration (HRC) in a shared workspace has become a common pattern in real-world robot applications and has garnered significant research interest. However, most existing studies for human-in-the-loop (HITL) collaboration with robots in a shared workspace evaluate in either simplified…

Cited by 1SourcePDFScholar
2024

Efficient Multi-scale Network with Learnable Discrete Wavelet Transform for Blind Motion Deblurring

CVPR 2024poster

Coarse-to-fine schemes are widely used in traditional single-image motion deblur; however in the context of deep learning existing multi-scale algorithms not only require the use of complex modules for feature fusion of low-scale RGB images and deep semantics but also manually generate low-resolutio…

2024

EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models

CVPR 2024highlight

Vision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities with the majority focusing on the third-person perspective and only a few addressing specific tasks from the first-person perspective. Howeve…

2024

GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

CVPR 2024poster

3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods which rely on representations like meshes and point clouds often fall short in realistically depicting complex scenes. On the other hand methods based on implicit 3D representations like…

2024

Learning a Distributed Hierarchical Locomotion Controller for Embodied Cooperation

CoRL 2024poster

In this work, we propose a distributed hierarchical locomotion control strategy for whole-body cooperation and demonstrate the potential for migration into large numbers of agents. Our method utilizes a hierarchical structure to break down complex tasks into smaller, manageable sub-tasks. By incorpo…

Cited by 3SourceScholar
2024

Leveraging Large Language Model for Heterogeneous Ad Hoc Teamwork Collaboration

RSS 2024poster

Compared with the widely investigated homogeneous multi-robot collaboration, heterogeneous robots with different capabilities can provide a more efficient and flexible collaboration for more complex tasks. In this paper, we consider a more challenging heterogeneous ad hoc teamwork collaboration prob…

Cited by 8SourcePDFScholar
2024

Rethinking Causal Relationships Learning in Graph Neural Networks

AAAI 2024technical

Graph Neural Networks (GNNs) demonstrate their significance by effectively modeling complex interrelationships within graph-structured data. To enhance the credibility and robustness of GNNs, it becomes exceptionally crucial to bolster their ability to capture causal relationships. However, despite…

2024

Smooth Computation without Input Delay: Robust Tube-Based Model Predictive Control for Robot Manipulator Planning

ICRA 2024poster

Model Predictive Control (MPC) has exhibited remarkable capabilities in optimizing objectives and meeting constraints. However, the substantial computational burden associated with solving the Optimal Control Problem (OCP) at each triggering instant introduces significant delays between state sampli…

Cited by 2SourceScholar
2024

Towards Objectively Benchmarking Social Intelligence of Language Agents at the Action Level

ACL 2024findings

Prominent large language models have exhibited human-level performance in many domains, even enabling the derived agents to simulate human and social interactions. While practical works have substantiated the practicability of grounding language agents in sandbox simulation or embodied simulators, c…

2024

USD-SLAM: A Universal Visual SLAM Based on Large Segmentation Model in Dynamic Environments

RA-L 2024

Visual Simultaneous Localization and Mapping (SLAM) has been widely adopted in autonomous driving and robotics. While most SLAM systems operate effectively in static or low-dynamic environments, achieving precise pose estimation in diverse unknown dynamic environments continues to pose a significant

Cited by 9SourceScholar
2024

Vision-Language Foundation Models as Effective Robot Imitators

ICLR 2024spotlight

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on…

Cited by 133SourcePDFScholar
2023

Adaptive Optimal Electrical Resistance Tomography for Large-Area Tactile Sensing

ICRA 2023poster

It is critical to perceive physical contact for intelligent robots to safely interact in dynamic, unstructured environments. As physical contacts can occur at any location, a well-performing tactile sensing system should be able to deploy a large area on robotic surface. Some researchers have implem…

Cited by 8SourceScholar
2023

Embodied Referring Expression for Manipulation Question Answering in Interactive Environment

ICRA 2023poster

Embodied agents are expected to perform more complicated tasks in an interactive environment, with the progress of Embodied AI in recent years. Existing embodied tasks including Embodied Referring Expression (ERE) and other QA-form tasks mainly focuses on interaction in term of linguistic instructio…

Cited by 7SourceScholar
2023

Extracting Dynamic Navigation Goal from Natural Language Dialogue

IROS 2023poster

Effective access to relevant environmental changes in large human environments is critical for service robots to perform tasks. Since the position of a dynamic goal such as a human is variable, it will be difficult for the robot to locate him accurately. It is worth noting that humans can obtain inf…

Cited by 8SourceScholar
2023

Masked Space-Time Hash Encoding for Efficient Dynamic Scene Reconstruction

NeurIPS 2023spotlight

In this paper, we propose the Masked Space-Time Hash encoding (MSTH), a novel method for efficiently reconstructing dynamic 3D scenes from multi-view or monocular videos. Based on the observation that dynamic scenes often contain substantial static areas that result in redundancy in storage and comp…

Cited by 30SourcePDFScholar
2023

Mixed Neural Voxels for Fast Multi-view Video Synthesis

ICCV 2023oral

Synthesizing high-fidelity videos from real-world multiview input is challenging due to the complexities of real-world environments and high-dynamic movements. Previous works based on neural radiance fields have demonstrated high-quality reconstructions of dynamic scenes. However, training such mode…

Cited by 72PDFcodeScholar
2023

Natural Language Instruction Understanding for Robotic Manipulation: a Multisensory Perception Approach

ICRA 2023poster

It has always been expected that the robot can understand the natural language instruction and thus a more natural human-robot interaction is achieved. Currently, the robot usually interprets the instruction by visually grounding the textual information to its surroundings, while it may be not enoug…

Cited by 8SourceScholar
2023

Tg-Critic: A Timbre-Guided Model For Reference-Independent Singing Evaluation

ICASSP 2023accepted

Automatic singing evaluation independent of reference melody is a challenging task due to its subjective and multi-dimensional nature. As an essential attribute of singing voices, vocal timbre has a non-negligible effect and influence on human perception of singing quality. However, no research has…

Cited by 0SourceScholar
2023

TrOMR:Transformer-Based Polyphonic Optical Music Recognition

ICASSP 2023accepted

Optical Music Recognition (OMR) is an important technology in music and has been researched for a long time. Previous approaches for OMR are usually based on CNN for image understanding and RNN for music symbol classification. In this paper, we propose a transformer-based approach with excellent glo…

Cited by 0SourceScholar
2022

Audio-Visual Grounding Referring Expression for Robotic Manipulation

ICRA 2022poster

Referring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both the audio and visual information to understand the referrin…

Cited by 19SourceScholar
2022

C-Shaped Bidirectional Stiffness Joint Design For Anthropomorphic Hand

RA-L 2022

In this letter, we propose a C-shaped bidirectional stiffness joint for an anthropomorphic hand. The spring steel piece is used to connect the knuckles with inner concave C-shape as a rotational joint. The finger is bent by the tendon driven with low stiffness and reset by its own elasticity. The re

Cited by 11SourceScholar
2022

Depth-Aware Vision-and-Language Navigation using Scene Query Attention Network

ICRA 2022poster

Vision-and-language navigation (VLN) has been an important task in the field of Robotics and Computer Vision. However, most existing vision-and-language navigation models only use features extracted from RGB observation as input, while robots can utilize depth sensors in the real world. Existing res…

Cited by 4SourceScholar
2022

Embodied Multi-Agent Task Planning from Ambiguous Instruction

RSS 2022poster

In human-robots collaboration scenarios, a human would give robots an instruction that is intuitive for the human himself to accomplish. However, the instruction given to robots is likely ambiguous for them to understand as some information is implicit in the instruction. Therefore, it is necessary…

Cited by 26SourcePDFScholar
2022

InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection

IROS 2022poster

Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggest…

Cited by 26SourceScholar
2022

Sim2Real Object-Centric Keypoint Detection and Description

AAAI 2022technical

Keypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, r…

Cited by 9SourcePDFScholar
2021

Evaluations of the Gap between Supervised and Reinforcement Lifelong Learning on Robotic Manipulation Tasks

CoRL 2021poster

Overcoming catastrophic forgetting is of great importance for deep learning and robotics. Recent lifelong learning research has great advances in supervised learning. However, little work focuses on reinforcement learning(RL). We focus on evaluating the performances of state-of-the-art lifelong lear…

Cited by 13SourceScholar
2021

Line-based Automatic Extrinsic Calibration of LiDAR and Camera

ICRA 2021poster

Reliable real-time extrinsic parameters of 3D Light Detection and Ranging (LiDAR) and camera are a key component of multi-modal perception systems. However, extrinsic transformation may drift gradually during operation, which can result in decreased accuracy of perception system. To solve this probl…

Cited by 54SourceScholar
2021

Neighborhood Spatial Aggregation based Efficient Uncertainty Estimation for Point Cloud Semantic Segmentation

ICRA 2021poster

Uncertainty estimation for point cloud semantic segmentation is to quantify the confidence degree for the predicted label of points, which is essential for decision-making tasks. This paper proposes a neighborhood spatial aggregation based method, NSA-MC dropout, to achieve efficient uncertainty est…

Cited by 4SourcecodeScholar
2020

Multi-Agent Embodied Question Answering in Interactive Environments

ECCV 2020poster

We investigate a new AI task --- Multi-Agent Interactive Question Answering --- where several agents explore the scene jointly in interactive environments to answer a question. To cooperate efficiently and answer accurately, agents must be well-organized to have balanced work division and share know…

Cited by 38SourcePDFScholar
2020

Self-Supervised Learning for Alignment of Objects and Sound

ICRA 2020poster

The sound source separation problem has many useful applications in the field of robotics, such as human-robot interaction, scene understanding, etc. However, it remains a very challenging problem. In this paper, we utilize both visual and audio information of videos to perform the sound source sepa…

Cited by 5SourceScholar
2020

Unsupervised Representation Learning by Invariance Propagation

NeurIPS 2020spotlight

Unsupervised learning methods based on contrastive learning have drawn increasing attention and achieved promising results. Most of them aim to learn representations invariant to instance-level variations, which are provided by different views of the same instance. In this paper, we propose Invarian…

2019

An Object Attribute Guided Framework for Robot Learning Manipulations from Human Demonstration Videos

IROS 2019poster

Learning manipulations from videos is an inspiriting way for robots to acquire new skills. In this paper, we propose a framework that can generate robotic manipulation plans by observing human demonstration videos without special marks or unnatural demonstrated behaviors. More specifically, the fram…

Cited by 4SourceScholar
2019

Deep Reinforcement Learning for Robotic Pushing and Picking in Cluttered Environment

IROS 2019poster

In this paper, a novel robotic grasping system is established to automatically pick up objects in cluttered scenes. A composite robotic hand composed of a suction cup and a gripper is designed for grasping the object stably. The suction cup is used for lifting the object from the clutter first and t…

Cited by 103SourceScholar
2019

Imitation Learning from Observations by Minimizing Inverse Dynamics Disagreement

NeurIPS 2019spotlight

This paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervisions, LfO is more practical in leveraging previously inapplicable resources (e.g., videos)…

Cited by 90SourcePDFScholar
2019

Toa Source Node Self-positioning with Unknown Clock Skew in Wireless Sensor Networks

ICASSP 2019accepted

This paper investigates time-of-arrival (TOA) source node self-positioning with unknown clock skews in wireless sensor networks. For the source-to-anchor direction, source node clock skew does not affect the localization performance. When synchronized anchor nodes simultaneously transmit signals to…

Cited by 0SourceScholar
2018

A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand Grasping

ICRA 2018poster

Humans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on…

Cited by 64SourceScholar
2018

Deep Feature Pyramid Reconfiguration for Object Detection

ECCV 2018poster

State-of-the-art object detectors usually learn multi-scale representations to get better results by employing feature pyramids. However, the current designs for feature pyramids are still inefficient to integrate the semantic information over different scales. In this paper, we begin by investigati…

2018

Semidefinite Programming for Tdoa Localization with Locally Synchronized Anchor Nodes

ICASSP 2018accepted

The most state-of-art time-difference-of-arrival (TDOA) localization algorithms are performed under the assumption that all the nodes are synchronized. However, for a widely distributed wireless sensor networks (WSNs), time synchronization between all the nodes is not a trival problem. In this paper…

Cited by 0SourceScholar
2017

RON: Reverse Connection With Objectness Prior Networks for Object Detection

CVPR 2017poster

We present RON, an efficient and effective framework for generic object detection. Our motivation is to smartly associate the best of the region-based (e.g., Faster R-CNN) and region-free (e.g., SSD) methodologies. Under fully convolutional architecture, RON mainly focuses on two fundamental problem…

Cited by 539PDFScholar
2016

Sparse Coding and Dictionary Learning With Linear Dynamical Systems

CVPR 2016oral

Linear Dynamical Systems (LDSs) are the fundamental tools for encoding spatio-temporal data in various disciplines. To enhance the performance of LDSs, in this paper, we address the challenging issue of performing sparse coding on the space of LDSs, where both data and dictionary atoms are LDSs. Rat…

Cited by 38PDFScholar