← Search

Lujia Wang

30 accepted papers

2026

LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer

ICLR 2026poster

Universal image restoration (UIR) aims to recover images degraded by unknown mixtures while preserving semantics—conditions under which discriminative restorers and UNet-based diffusion priors often oversmooth, hallucinate, or drift. We present LucidFlux, a caption-free UIR framework that adapts a l…

Cited by 0SourcecodeScholar
2026

PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design Generation

CVPR 2026

Recent advancements in the text-rendering capabilities of image generation models have made the end-to-end creation of graphic design content, such as posters, increasingly feasible. However, existing reward models fall short of accurately assessing design quality, as they primarily focus on global

Cited by 0SourcecodeScholar
2026

SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting

ICRA 2026poster

Vision-Language-Action (VLA) models, such as OpenVLA, hold the promise of generalist robots, yet their performance is often impaired by distracted attention, which we identify as a manifestation of shortcut learning. We posit that the solution lies not in architectural modifications, but in a new tr…

Cited by 0Scholar
2025

PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

IROS 2025

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in

Cited by 60SourceScholar
2024

MGCBS: An Optimal and Efficient Algorithm for Solving Multi-Goal Multi-Agent Path Finding Problem

IJCAI 2024poster

With the expansion of the scale of robotics applications, the multi-goal multi-agent pathfinding (MG-MAPF) problem began to gain widespread attention. This problem requires each agent to visit pre-assigned multiple goal points at least once without conflict. Some previous methods have been proposed…

2023

CenterLineDet: CenterLine Graph Detection for Road Lanes with Vehicle-mounted Sensors by Transformer for HD Map Generation

ICRA 2023poster

With the fast development of autonomous driving technologies, there is an increasing demand for high-definition (HD) maps, which provide reliable and robust prior information about the static part of the traffic environments. As one of the important elements in HD maps, road lane centerline is criti…

Cited by 18SourcecodeScholar
2023

Completely Rational $\text{SO}(n)$ Orthonormalization

ICRA 2023poster

The rotation orthonormalization on the special orthogonal group \text{SO}(n)\text{SO}(n), also known as the high dimensional nearest rotation problem, has been revisited. A new generalized simple iterative formula has been proposed that solves this problem in a completely rational manner. Rational o…

Cited by 0SourceScholar
2023

PBACalib: Targetless Extrinsic Calibration for High-Resolution LiDAR-Camera System Based on Plane-Constrained Bundle Adjustment

RA-L 2023

The strategy of fusing multi-model data especially from cameras, light detection and ranging sensors (LiDAR), is frequently considered in robotics to enhance the performance of the perception and navigation tasks. Extrinsic calibration, which spatially aligns different sources into a unified coordin

Cited by 20SourceScholar
2023

RNGDet++: Road Network Graph Detection by Transformer With Instance Segmentation and Multi-Scale Features Enhancement

RA-L 2023

The road network graph is a critical component for downstream tasks in autonomous driving, such as global route planning and navigation. In the past years, road network graphs are usually annotated by human experts manually, which is time-consuming and labor-intensive. To annotate road network graph

Cited by 50SourceScholar
2023

Real-Time Neural Dense Elevation Mapping for Urban Terrain With Uncertainty Estimations

RA-L 2023

Having good knowledge of terrain information is essential for improving the performance of various downstream tasks on complex terrains, especially for the locomotion and navigation of legged robots. We present a novel framework for neural urban terrain reconstruction with uncertainty estimations. I

Cited by 21SourceScholar
2022

An Online Interactive Approach for Crowd Navigation of Quadrupedal Robots

IROS 2022poster

Robot navigation in human crowds remains the challenge of understanding human behaviors in different scenarios. We present an approach for interactive and human-friendly crowd navigation in complex static environments. The planner models the online interactions among the robot, humans, and the stati…

Cited by 7SourceScholar
2022

Efficient Speed Planning for Autonomous Driving in Dynamic Environment With Interaction Point Model

RA-L 2022

Safely interacting with other traffic participants is one of the core requirements for autonomous driving, especially in intersections and occlusions. Most existing approaches are designed for particular scenarios and require significant human labor in parameter tuning to be applied to different sit

Cited by 16SourcecodeScholar
2022

FusionPortable: A Multi-Sensor Campus-Scene Dataset for Evaluation of Localization and Mapping Accuracy on Diverse Platforms

IROS 2022poster

Combining multiple sensors enables a robot to maximize its perceptual awareness of environments and enhance its robustness to external disturbance, crucial to robotic navigation. This paper proposes the FusionPortable benchmark, a complete multi-sensor dataset with a diverse set of sequences for mob…

Cited by 39SourceScholar
2022

MMFN: Multi-Modal-Fusion-Net for End-to-End Driving

IROS 2022poster

Inspired by the fact that humans use diverse sensory organs to perceive the world, sensors with different modalities are deployed in end-to-end driving to obtain the global context of the 3D scene. In previous works, camera and LiDAR inputs are fused through transformers for better driving performan…

Cited by 35SourcecodeScholar
2022

R-PCC: A Baseline for Range Image-based Point Cloud Compression

ICRA 2022poster

In autonomous vehicles or robots, point clouds from LiDAR can provide accurate depth information of objects compared with 2D images, but they also suffer a large volume of data, which is inconvenient for data storage or transmission. In this paper, we propose a Range image-based Point Cloud Compress…

Cited by 26SourcecodeScholar
2022

UnDAF: A General Unsupervised Domain Adaptation Framework for Disparity or Optical Flow Estimation

ICRA 2022poster

Disparity and optical flow estimation are respectively 1D and 2D dense correspondence matching (DCM) tasks in nature. Unsupervised domain adaptation (UDA) is crucial for their success in new and unseen scenarios, enabling networks to draw inferences across different domains without manually-labeled…

Cited by 8SourceScholar
2022

csBoundary: City-Scale Road-Boundary Detection in Aerial Images for High-Definition Maps

RA-L 2022

High-Definition (HD) maps can provide precise geometric and semantic information of static traffic environments for autonomous driving. Road-boundary is one important information presented in HD maps since it distinguishes between road areas and off-road areas, which can guide vehicles to drive with

Cited by 37SourceScholar
2021

CP-loss: Connectivity-preserving Loss for Road Curb Detection in Autonomous Driving with Aerial Images

IROS 2021poster

Road curb detection is important for autonomous driving. It can be used to determine road boundaries to constrain vehicles on roads, so that potential accidents could be avoided. Most of the current methods detect road curbs online using vehicle-mounted sensors, such as cameras or 3-D Lidars. Howeve…

Cited by 13SourceScholar
2021

DiTNet: End-to-End 3D Object Detection and Track ID Assignment in Spatio-Temporal World

RA-L 2021

End-to-end 3D object detection and tracking based on point clouds is receiving more and more attention in many robotics applications, such as autonomous driving. Compared with 2D images, 3D point clouds do not have enough texture information for data association. Thus, we propose an end-to-end point

Cited by 22SourceScholar
2021

Greedy-Based Feature Selection for Efficient LiDAR SLAM

ICRA 2021poster

Modern LiDAR-SLAM (L-SLAM) systems have shown excellent results in large-scale, real-world scenarios. However, they commonly have a high latency due to the expensive data association and nonlinear optimization. This paper demonstrates that actively selecting a subset of features significantly improv…

Cited by 50SourceScholar
2021

Learning Interpretable End-to-End Vision-Based Motion Planning for Autonomous Driving with Optical Flow Distillation

ICRA 2021poster

Recently, deep-learning based approaches have achieved impressive performance for autonomous driving. However, end-to-end vision-based methods typically have limited interpretability, making the behaviors of the deep networks difficult to explain. Hence, their potential applications could be limited…

Cited by 57SourceScholar
2021

On Bundle Adjustment for Multiview Point Cloud Registration

RA-L 2021

Multiview registration is used to estimate Rigid Body Transformations (RBTs) from multiple frames and reconstruct a scene with corresponding scans. Despite the success of pairwise registration and pose synchronization, the concept of Bundle Adjustment (BA) has been proven to better maintain global c

Cited by 24SourcecodeScholar
2021

Peer-Assisted Robotic Learning: A Data-Driven Collaborative Learning Approach for Cloud Robotic Systems

ICRA 2021poster

A technological revolution is occurring in the field of robotics with the data-driven deep learning technology. However, building datasets for each local robot is laborious. Meanwhile, data islands between local robots make data unable to be utilized collaboratively. To address this issue, the work…

Cited by 28SourceScholar
2020

Federated Imitation Learning: A Novel Framework for Cloud Robotic Systems With Heterogeneous Sensor Data

RA-L 2020

Humans are capable of learning a new behavior by observing others to perform the skill. Similarly, robots can also implement this by imitation learning. Furthermore, if with external guidance, humans can master the new behavior more efficiently. So, how can robots achieve this? To address the issue,

Cited by 80SourceScholar
2019

Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems

IROS 2019poster

This paper was motivated by the problem of how to make robots fuse and transfer their experience so that they can effectively use prior knowledge and quickly adapt to new environments. To address the problem, we present a learning architecture for navigation in cloud robotic systems: Lifelong Federa…

Cited by 325SourceScholar
2019

Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems

RA-L 2019

This letter was motivated by the problem of how to make robots fuse and transfer their experience so that they can effectively use prior knowledge and quickly adapt to new environments. To address the problem, we present a learning architecture for navigation in cloud robotic systems: Lifelong Feder

Cited by 272SourceScholar
2018

Indoor Mapping and Localization for Pedestrians using Opportunistic Sensing with Smartphones

IROS 2018poster

Indoor localization for pedestrians has gained increasing popularity among the rich body of literature for the last decade. In this paper, a low-cost indoor mapping and localization solution is proposed using the opportunistic signals from ambient indoor environments with a smartphone. It is compose…

Cited by 11SourceScholar
2018

Plugo: A Scalable Visible Light Communication System Towards Low-Cost Indoor Localization

IROS 2018poster

Indoor localization is critical to many location-aware applications, however, a low-cost solution with guaranteed accuracies has not yet come. Visible Light Communication (VLC-) based localization techniques are very promising to fill this gap. In this paper, we propose Plugo, a novel VLC system wit…

Cited by 14SourceScholar