← Search

Liangjun Zhang

42 accepted papers

2025

A Joint Learning of Force Feedback of Robotic Manipulation and Textual Cues for Granular Materials Classification

RA-L 2025

Granular materials (GMs) are formed by a collection of particles. Even if their visual representation is straightforward, it can be seriously affected in the visually constrained environment. Based on frequency features observed in force signals, this paper proposes a non-visual classifier, <bold xm

Cited by 22SourceScholar
2025

Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning

RA-L 2025

Granular materials (GMs) are ubiquitous in daily life. Understanding their properties is also important, especially in agriculture and industry. However, existing works require dedicated measurement equipment and also need large human efforts to handle a large number of particles. In this paper, we

Cited by 3SourceScholar
2024

3D Human Pose Estimation via Non-Causal Retentive Networks

ECCV 2024poster

"Temporal dependencies are essential in 3D human pose estimation to mitigate depth ambiguity. Previous methods typically use a fixed-length sliding window to capture these dependencies. However, they treat past and future frames equally, ignoring the fact that relying on too many future frames incre…

2024

HO-Gaussian: Hybrid Optimization of 3D Gaussian Splatting for Urban Scenes

ECCV 2024poster

"The rapid growth of 3D Gaussian Splatting (3DGS) has revolutionized neural rendering, enabling real-time production of high-quality renderings. However, the previous 3DGS-based methods have limitations in urban scenes due to reliance on initial Structure-from-Motion (SfM) points and difficulties in…

Cited by 15SourcePDFScholar
2024

LiDAR-CS Dataset: LiDAR Point Cloud Dataset with Cross-Sensors for 3D Object Detection

ICRA 2024poster

Over the past few years, there has been remarkable progress in research on 3D point clouds and their use in autonomous driving scenarios has become widespread. However, deep learning methods heavily rely on annotated data and often face domain generalization issues. Unlike 2D images whose domains us…

Cited by 21SourcecodeScholar
2024

RLingua: Improving Reinforcement Learning Sample Efficiency in Robotic Manipulations With Large Language Models

RA-L 2024

Reinforcement learning (RL) has demonstrated its capability in solving various tasks but is notorious for its low sample efficiency. In this paper, we propose RLingua, a framework that can leverage the internal knowledge of large language models (LLMs) to reduce the sample complexity of RL in roboti

Cited by 31SourcecodeScholar
2024

RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language Model

IROS 2024poster

Recent advances in Large Language Models (LLMs) have showcased their remarkable reasoning capabilities, making them influential across various fields. However, in robotics, their use has primarily been limited to manipulation planning tasks due to their inherent textual output. This paper addresses…

Cited by 6SourceScholar
2024

Safety-Critical Scenario Generation Via Reinforcement Learning Based Editing

ICRA 2024poster

Generating safety-critical scenarios is essential for testing and verifying the safety of autonomous vehicles. Traditional optimization techniques suffer from the curse of dimensionality and limit the search space to fixed parameter spaces. To address these challenges, we propose a deep reinforcemen…

Cited by 9SourceScholar
2024

VIHE: Virtual In-Hand Eye Transformer for 3D Robotic Manipulation

IROS 2024poster

In this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the…

Cited by 3SourcecodeScholar
2023

Boosting Feedback Efficiency of Interactive Reinforcement Learning by Adaptive Learning from Scores

IROS 2023poster

Interactive reinforcement learning has shown promise in learning complex robotic tasks. However, the process can be human-intensive due to the requirement of a large amount of interactive feedback. This paper presents a new method that uses scores provided by humans instead of pairwise preferences t…

Cited by 0SourcecodeScholar
2023

FloorplanNet: Learning Topometric Floorplan Matching for Robot Localization

ICRA 2023poster

Given a building floorplan, humans can localize themselves by matching the observation of the environment with the floorplan using geometric, semantic, and topological clues. Inspired by this insight, this paper proposes a learning- based topometric robot localization method FloorplanNet, which impl…

Cited by 9SourcecodeScholar
2023

GOATS: Goal Sampling Adaptation for Scooping with Curriculum Reinforcement Learning

IROS 2023poster

In this work, we first formulate the problem of robotic water scooping using goal-conditioned reinforcement learning. This task is particularly challenging due to the complex dynamics of fluid and the need to achieve multi-modal goals. The policy is required to successfully reach both position goals…

Cited by 10SourceScholar
2023

Interpretable and Flexible Target-Conditioned Neural Planners For Autonomous Vehicles

ICRA 2023poster

Learning-based approaches to autonomous vehicle planners have the potential to scale to many complicated real-world driving scenarios by leveraging huge amounts of driver demonstrations. However, prior work only learns to estimate a single planning trajectory, while there may be multiple acceptable…

Cited by 3SourceScholar
2023

MFF-Net: Towards Efficient Monocular Depth Completion With Multi-Modal Feature Fusion

RA-L 2023

Remarkable progress has been achieved by current depth completion approaches, which produce dense depth maps from sparse depth maps and corresponding color images. However, the performances of these approaches are limited due to the insufficient feature extractions and fusions. In this work, we prop

Cited by 38SourceScholar
2023

MapNeRF: Incorporating Map Priors into Neural Radiance Fields for Driving View Simulation

IROS 2023poster

Simulating camera sensors is a crucial task in autonomous driving. Although neural radiance fields are exceptional at synthesizing photorealistic views in driving simulations, they still fail to generate extrapolated views. This paper proposes to incorporate map priors into neural radiance fields to…

Cited by 13SourceScholar
2023

NeRF-Loc: Transformer-Based Object Localization Within Neural Radiance Fields

RA-L 2023

Neural Radiance Fields (NeRFs) have become a widely-applied scene representation technique in recent years, showing advantages for robot navigation and manipulation tasks. To further advance the utility of NeRFs for robotics, we propose a transformer-based framework, <monospace xmlns:mml="http://www

Cited by 14SourceScholar
2023

VINet: Visual and Inertial-based Terrain Classification and Adaptive Navigation over Unknown Terrain

ICRA 2023poster

We present a visual and inertial-based terrain classification network (VINet) for robotic navigation over different traversable surfaces. We use a novel navigation-based labeling scheme for terrain classification and generalization on unknown surfaces. Our proposed perception method and adaptive sch…

Cited by 12SourceScholar
2022

A Generalized Continuous Collision Detection Framework of Polynomial Trajectory for Mobile Robots in Cluttered Environments

RA-L 2022

In this letter, we introduce a generalized continuous collision detection (CCD) framework for the mobile robot along the polynomial trajectory in cluttered environments including various static obstacle models. Specifically, we find that the collision conditions between robots and obstacles could be

Cited by 16SourceScholar
2022

Excavation of Fragmented Rocks with Multi-modal Model-based Reinforcement Learning

IROS 2022poster

This paper presents a multi-modal model-based reinforcement learning (MBRL) approach to the excavation of fragmented rocks, which are very challenging to model due to their highly variable sizes and geometries, and visual occlusions. A multi-modal recurrent neural network (RNN) learns the dynamics o…

Cited by 9SourceScholar
2022

Imitation Learning and Model Integrated Excavator Trajectory Planning

IROS 2022poster

Automated excavation is promising to improve the safety and efficiency of excavators, and trajectory planning is one of the most important techniques. In this paper, we propose a two-stage method that integrates data-driven imitation learning and model-based trajectory optimization to generate optim…

Cited by 12SourceScholar
2022

PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching

ECCV 2022poster

"Existing deep learning based stereo matching methods either focus on achieving optimal performances on the target dataset while with poor generalization for other datasets or focus on handling the cross-domain generalization by suppressing the domain sensitive features which results in a significan…

Cited by 98SourcePDFScholar
2022

ProposalContrast: Unsupervised Pre-training for LiDAR-Based 3D Object Detection

ECCV 2022poster

"Existing approaches for unsupervised point cloud pre-training are constrained to either scene-level or point/voxel-level instance discrimination. Scene-level methods tend to lose local details that are crucial for recognizing the road objects, while point/voxel-level methods inherently suffer from…

2022

Semi-Supervised 3D Object Detection with Proficient Teachers

ECCV 2022poster

"Dominated point cloud-based 3D object detectors in autonomous driving scenarios rely heavily on the huge amount of accurately labeled samples, however, 3D annotation in the point cloud is extremely tedious, expensive and time-consuming. To reduce the dependence on large supervision, semi-supervised…

2022

TNS: Terrain Traversability Mapping and Navigation System for Autonomous Excavators

RSS 2022poster

We present a terrain traversability mapping and navigation system (TNS) for autonomous excavator applications in an unstructured environment. We use an efficient approach to extract terrain features from RGB images and 3D point clouds and incorporate them into a global map for planning and navigatio…

2022

Text2video: Text-Driven Talking-Head Video Synthesis with Personalized Phoneme - Pose Dictionary

ICASSP 2022accepted

With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds a phoneme-pose dictionary and trains a generative adversari…

Cited by 0SourceScholar
2021

AutoShape: Real-Time Shape-Aware Monocular 3D Object Detection

ICCV 2021poster

Existing deep learning-based approaches for monocular 3D object detection in autonomous driving often model the object as a rotated 3D cuboid while the object's geometric shape has been ignored. In this work, we propose an approach for incorporating the shape-aware 2D/3D constraints into the 3D dete…

Cited by 165PDFcodeScholar
2021

FCFR-Net: Feature Fusion based Coarse-to-Fine Residual Learning for Depth Completion

AAAI 2021technical

Depth completion aims to recover a dense depth map from a sparse depth map with the corresponding color image as input. Recent approaches mainly formulate the depth completion as a one-stage end-to-end learning task, which outputs dense depth maps directly. However, the feature extraction and superv…

Cited by 138SourcePDFScholar
2021

LiDAR-Aug: A General Rendering-Based Augmentation Framework for 3D Object Detection

CVPR 2021poster

Annotating the LiDAR point cloud is crucial for deep learning-based 3D object detection tasks. Due to expensive labeling costs, data augmentation has been taken as a necessary module and plays an important role in training the neural network. "Copy" and "paste" (i.e., GT-Aug) is the most commonly us…

Cited by 75PDFScholar
2021

Look Before You Act: Boosting Pseudo-LiDAR with Online Semantic Embedding

IROS 2021poster

Vision-based 3D object detection is a research focus in the field of autonomous driving system. While recently proposed pseudo-LiDAR is a promising solution, its performance is severely restricted by the image-based depth estimator, leading to a considerable performance gap against the LiDAR-based c…

Cited by 0SourceScholar
2021

Robust 2D/3D Vehicle Parsing in Arbitrary Camera Views for CVIS

ICCV 2021poster

We present a novel approach to robustly detect and perceive vehicles in different camera views as part of a cooperative vehicle-infrastructure system (CVIS). Our formulation is designed for arbitrary camera views and makes no assumptions about intrinsic or extrinsic parameters. First, to deal with m…

Cited by 3PDFcodeScholar
2021

Self-Supervised Monocular Depth Estimation for All Day Images Using Domain Separation

ICCV 2021poster

Remarkable results have been achieved by DCNN based self-supervised depth estimation approaches. However, most of these approaches can only handle either day-time or night-time images, while their performance degrades for all-day images due to large domain shift and the variation of illumination bet…

Cited by 86PDFcodeScholar
2020

3D Part Guided Image Editing for Fine-Grained Object Understanding

CVPR 2020poster

Holistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving ve…

Cited by 14PDFcodeScholar
2019

Compact Reachability Map for Excavator Motion Planning

IROS 2019poster

In this paper, we propose a novel compact reachability map representation for excavator motion planning. The constructed reachability map can concisely encode the bucket’s reachable pose and the translation capability limited by excavator’s kinematic structure. By explicitly exploiting the property…

Cited by 24SourceScholar