← Search

Zhe Liu

125 accepted papers

2026

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in Unified Multimodal Models via Decompositional Verifiable Reward

ICML 2026poster

In this paper, we propose **AlphaGRPO**, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models (UMMs) to enhance multimodal generation capabilities without relying on external knowledge injection. Our approach unlocks the model's intrinsic…

Cited by 0SourceScholar
2026

DIAL-GS: Dynamic Instance Aware Reconstruction for Label-Free Street Scenes with 4D Gaussian Splatting

ICRA 2026poster

Urban scene reconstruction is critical for autonomous driving, enabling structured 3D representations for data synthesis and closed-loop testing. Supervised approaches rely on costly human annotations and lack scalability, while current self-supervised methods often confuse static and dynamic elemen…

2026

Delving into Non-Exchangeability for Conformal Prediction in Graph-Structured Multivariate Time Series

ICML 2026poster

Point forecasting for graph-structured multivariate time series is a fundamental problem, but rigorous uncertainty quantification for such predictions is still underexplored. Conformal prediction (CP) offers uncertainty estimation with a solid coverage guarantee under the exchangeability assumption,…

Cited by 0SourceScholar
2026

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

CVPR 2026

Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction outputs in autonomous driving remains underexplored. In this paper, we propose DrivePI, a novel spatial-aware 4D MLLM th

Cited by 0SourcecodeScholar
2026

Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) often hallucinate when visual evidence conflicts with world knowledge, i.e., in counterfactual scenarios. We propose Envision-Attend-Respond (EnAR), a training-free framework that leverages visual priors to steer the model's attention toward counterfactual elemen

Cited by 0SourcecodeScholar
2026

Fair in Mind, Fair in Action? A Synchronous Benchmark for Understanding and Generation in UMLLMs

ICLR 2026poster

As artificial intelligence (AI) is increasingly deployed across domains, ensuring fairness has become a core challenge. However, the field faces a "Tower of Babel'' dilemma: fairness metrics abound, yet their underlying philosophical assumptions often conflict, hindering unified paradigms—particular…

Cited by 0SourceScholar
2026

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

CVPR 2026

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos, which makes learning difficult and leads to physically incons

Cited by 0SourceScholar
2026

GeoTeacher: Geometry-Guided Semi-Supervised 3D Object Detection

ICRA 2026poster

Semi-supervised 3D object detection (SS3D), aiming to explore unlabeled data for boosting 3D object detectors, has emerged as an active research area in recent years. Some previous methods have shown substantial improvements by either employing heterogeneous teacher models to provide high-quality ps…

2026

Many Minds, One Path: LLM-Augmented Consensus Decision for Distributed Control in Multi-Agent Collaborative Stable Scenarios

AAAI 2026technical

Distributed multi-agent systems are increasingly deployed in dynamic and high-stakes environments such as power grids, intelligent traffic systems, and collaborative robotics. In these systems, long-term stability, the ability to maintain coherent and safe system behavior over time, is critical but

Cited by 0SourcePDFScholar
2026

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

ICML 2026poster

Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, this capability introduces a new safety attack surface: harmful outputs may arise from tool orchestration, where individually benign steps combine into u…

Cited by 0SourceScholar
2026

SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

ICML 2026poster

With the rapid evolution of foundation models, Large Language Model (LLM) agents have demonstrated increasingly powerful tool-use capabilities. However, this proficiency introduces significant security risks, as malicious actors can manipulate agents into executing tools to generate harmful content.…

Cited by 0SourceScholar
2026

SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

ICML 2026poster

With the rapid evolution of foundation models, Large Language Model (LLM) agents have demonstrated increasingly powerful tool-use capabilities. However, this proficiency introduces significant security risks, as malicious actors can manipulate agents into executing tools to generate harmful content.…

Cited by 0SourceScholar
2026

Towards a Universally Transferable Acceleration Method for Density Functional Theory

ICLR 2026poster

Recently, sophisticated deep learning-based approaches have been developed for generating efficient initial guesses to accelerate the convergence of density functional theory (DFT) calculations. While the actual initial guesses are often density matrices (DM), quantities that can convert into densit…

Cited by 0SourceScholar
2026

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

CVPR 2026

Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. However, existing reconstruction methods are primarily limited to replicating observed scenes and lack the capability for diverse weather simulation. While i

Cited by 0SourcecodeScholar
2025

An Optimized GPU-based Acceleration of CRYSTALS-Dilithium

ICASSP 2025accepted

CRYSTALS-Dilithium has recently been selected as one of the next generation post-quantum signature algorithm standards. However, due to the extensive volume of data elements and the high complexity of operations, post-quantum cryptographic algorithms commonly face significant performance challenges,…

Cited by 0SourceScholar
2025

DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately

AAAI 2025technical

The emergence of the large language model (LLM) has shown its superiority in a wide range of disciplines, including language understanding and translation, relational logic reasoning, and even partial differential equations solving. The transformer is the pervasive backbone architecture for the foun…

2025

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

ICCV 2025poster

Open-set 3D object retrieval (3DOR) is an emerging task aiming to retrieve 3D objects of unseen categories beyond the training set. Existing methods typically utilize all modalities (i.e., voxels, point clouds, multi-view images) and train specific backbones before fusion. However, they still strugg…

2025

Differential Private Stochastic Optimization with Heavy-tailed Data: Towards Optimal Rates

AAAI 2025technical

We study convex optimization problems under differential privacy (DP). With heavy-tailed gradients, existing works achieve suboptimal rates. The main obstacle is that existing gradient estimators have suboptimal tail property, resulting in a superfluous factor of d in the union bound. In this paper,…

Cited by 4SourcePDFScholar
2025

Enhancing Learning with Label Differential Privacy by Vector Approximation

ICLR 2025spotlight

Label differential privacy (DP) is a framework that protects the privacy of labels in training datasets, while the feature vectors are public. Existing approaches protect the privacy of labels by flipping them randomly, and then train a model to make the output approximate the privatized label. Howe…

Cited by 2SourcePDFScholar
2025

FLAME: Learning to Navigate with Multimodal LLM in Urban Environments

AAAI 2025technical

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specializ…

2025

Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic Warehousing

ICRA 2025

In warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing

Cited by 1SourceScholar
2025

GIPD: Global Intent Prediction and Decomposition of Cooperative Multi-Robot System in Non-Communication Environments

IROS 2025

In complex multi-robot application scenarios, particularly in dynamically adversarial, hazardous, or disaster environments, traditional cooperation paradigms face significant challenges due to unreliable or absent communication links. Achieving efficient cooperation in the absence of communication h

Cited by 0SourceScholar
2025

Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement

CVPR 2025poster

Vision Transformers (ViTs) have been widely applied in various computer vision and vision-language tasks. To gain insights into their robustness in practical scenarios, transferable adversarial examples on ViTs have been extensively studied. A typical approach to improving adversarial transferabilit…

2025

Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path

AAAI 2025technical

Transferable adversarial examples are known to cause threats in practical, black-box attack scenarios. A notable approach to improving transferability is using integrated gradients (IG), originally developed for model interpretability. In this paper, we find that existing IG-based attacks have limit…

2025

IoU-Aware Clustering for Anchor Configuration Determination in Efficient Defect Detection

IROS 2025

Deep-learning-based object detection has gained widespread application in surface defect inspection, with anchor-based detectors achieving remarkable success by utilizing dense anchors to align with defects. Determining the optimal anchor configuration, i.e., sizes and aspect ratios of anchor boxes,

Cited by 0SourceScholar
2025

LONG3R: Long Sequence Streaming 3D Reconstruction

ICCV 2025poster

Recent advancements in multi-view scene reconstruction have been significant, yet existing methods face limitations when processing streams of input images. These methods either rely on time-consuming offline optimization or are restricted to shorter sequences, hindering their applicability in real-…

2025

MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices

ICCV 2025poster

Recent advancements in deep neural networks have driven significant progress in image enhancement (IE). However, deploying deep learning models on resource-constrained platforms, such as mobile devices, remains challenging due to high computation and memory demands. To address these challenges and f…

2025

MovSAM: A Single-image Moving Object Segmentation Framework Based on Deep Thinking

IROS 2025

Moving object segmentation plays a vital role in understanding dynamic visual environments. While existing methods rely on multi-frame image sequences to identify moving objects, single-image MOS is critical for applications like motion intention prediction and handling camera frame drops. However,

Cited by 1SourcecodeScholar
2025

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

AAAI 2025technical

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of…

Cited by 0SourcePDFScholar
2025

Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation

NeurIPS 2025poster

Visual navigation is a fundamental problem in embodied AI, yet practical deployments demand long-horizon planning capabilities to address multi-objective tasks. A major bottleneck is data scarcity: policies learned from limited data often overfit and fail to generalize OOD. Existing neural network-b…

Cited by 0SourceScholar
2025

TWIST: Text-encoder Weight-editing for Inserting Secret Trojans in Text-to-Image Models

ACL 2025long

Text-to-image (T2I) models excel at generating high-quality images from text via powerful text encoders but training these encoders demands substantial computational resources. Consequently, many users seek pre-trained text encoders from model plugin-sharing platforms like Civitai and Hugging Face,…

Cited by 0SourcePDFScholar
2025

Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification

AAAI 2025technical

Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly…

2025

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

IJCAI 2025

Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range and temporal resolution but face difficulties with sparse data. Combining event

Cited by 0SourcePDFScholar
2025

Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration

NeurIPS 2025poster

Autonomous exploration in complex multi-agent reinforcement learning (MARL) with sparse rewards critically depends on providing agents with effective intrinsic motivation. While artificial curiosity offers a powerful self-supervised signal, it often confuses environmental stochasticity with meaningf…

Cited by 0SourceScholar
2024

A Huber Loss Minimization Approach to Mean Estimation under User-level Differential Privacy

NeurIPS 2024poster

Privacy protection of users' entire contribution of samples is important in distributed systems. The most effective approach is the two-stage scheme, which finds a small interval first and then gets a refined estimate by clipping samples into the interval. However, the clipping operation induces bia…

Cited by 7SourcePDFScholar
2024

Attribute-Missing Graph Clustering Network

AAAI 2024technical

Deep clustering with attribute-missing graphs, where only a subset of nodes possesses complete attributes while those of others are missing, is an important yet challenging topic in various practical applications. It has become a prevalent learning paradigm in existing studies to perform data imputa…

2024

Context-based and Diversity-driven Specificity in Compositional Zero-Shot Learning

CVPR 2024poster

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen attribute-object pairs based on a limited set of observed examples. Current CZSL methodologies despite their advancements tend to neglect the distinct specificity levels present in attributes. For instance given images of sliced strawb…

Cited by 16SourcePDFScholar
2024

Contextual Biasing of Named-Entities with Large Language Models

ICASSP 2024accepted

We explore contextual biasing with Large Language Models (LLMs) to enhance Automatic Speech Recognition (ASR) in second-pass rescoring. Our approach introduces the utilization of prompts for LLMs during rescoring without the need for fine-tuning. These prompts incorporate a biasing list and a set of…

Cited by 0SourceScholar
2024

Cooperative Path Planning for Four-Way Shuttle Vehicles in Storage and Retrieval Systems: A Hierarchically Dynamic Graph-Based Approach

IROS 2024poster

Recently, Shuttle-based Storage and Retrieval Systems (SBS/RSs) have garnered significant attention from both academia and industry, owing to their high spatial utilization and rapid response speed. However, the weak connectivity of roadmaps in densely stored environments increases the likelihood of…

Cited by 0SourceScholar
2024

DOC-RAG: ASR Language Model Personalization with Domain-Distributed Co-occurrence Retrieval Augmentation

COLING 2024main

We propose DOC-RAG - Domain-distributed Co-occurrence Retrieval Augmentation for ASR language model personalization aiming to improve the automatic speech recognition of rare word patterns in unseen domains. Our approach involves contrastively training a document retrieval module to rank external kn…

Cited by 2SourcePDFScholar
2024

DVLO: Deep Visual-LiDAR Odometry with Local-to-Global Feature Fusion and Bi-Directional Structure Alignment

ECCV 2024oral

"Information inside visual and LiDAR data is well complementary derived from the fine-grained texture of images and massive geometric information in point clouds. However, it remains challenging to explore effective visual-LiDAR fusion, mainly due to the intrinsic data structure inconsistency betwee…

2024

DVSAI: Diverse View-Shared Anchors Based Incomplete Multi-View Clustering

AAAI 2024technical

In numerous real-world applications, it is quite common that sample information is partially available for some views due to machine breakdown or sensor failure, causing the problem of incomplete multi-view clustering (IMVC). While several IMVC approaches using view-shared anchors have successfully…

Cited by 17SourcePDFScholar
2024

DifFlow3D: Toward Robust Uncertainty-Aware Scene Flow Estimation with Iterative Diffusion-Based Refinement

CVPR 2024poster

Scene flow estimation which aims to predict per-point 3D displacements of dynamic scenes is a fundamental task in the computer vision field. However previous works commonly suffer from unreliable correlation caused by locally constrained searching ranges and struggle with accumulated inaccuracy aris…

2024

Enhancing Exploratory Capability of Visual Navigation Using Uncertainty of Implicit Scene Representation

IROS 2024

In the context of visual navigation in unknown scenes, both “exploration” and “exploitation” are equally crucial. Robots must first establish environmental cognition through exploration and then utilize the cognitive information to accomplish target searches. However, most existing methods for image

Cited by 2SourcecodeScholar
2024

Frontier-enhanced Topological Memory with Improved Exploration Awareness for Embodied Visual Navigation

ECCV 2024poster

"We present a novel graph memory structure for navigation, called Frontier-enhanced Topological Memory (FTM). Most prior research primarily focuses on maintaining memory representations for explored areas. In contrast, our approach incorporates ghost nodes into the topological map to characterize un…

2024

Hawkes-Enhanced Spatial-Temporal Hypergraph Contrastive Learning Based on Criminal Correlations

AAAI 2024technical

Crime prediction is a crucial yet challenging task within urban computing, which benefits public safety and resource optimization. Over the years, various models have been proposed, and spatial-temporal hypergraph learning models have recently shown outstanding performances. However, three correlati…

Cited by 7SourcePDFScholar
2024

LION: Linear Group RNN for 3D Object Detection in Point Clouds

NeurIPS 2024poster

The benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward…

2024

Learning Hierarchical Graph-Based Policy for Goal-Reaching in Unknown Environments

RA-L 2024

Goal-reaching in unknown environments is one of the essential tasks in robot applications. Large-scale perception and long-horizon decision-making are the keys to solving this task as the operation scope expands or complexity rises. Existing navigation methods may suffer from degraded performance in

Cited by 6SourceScholar
2024

OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection

ECCV 2024poster

"Accurate depth information is crucial for enhancing the performance of multi-view 3D object detection. Despite the success of some existing multi-view 3D detectors utilizing pixel-wise depth supervision, they overlook two significant phenomena: 1) the depth supervision obtained from LiDAR points is…

2024

RCBEVDet: Radar-camera Fusion in Bird's Eye View for 3D Object Detection

CVPR 2024poster

Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However relying solely on cameras is difficult to achieve highly accurate and robust…

2024

Recovering from Privacy-Preserving Masking with Large Language Models

ICASSP 2024accepted

Model adaptation is crucial to handle the discrepancy between proxy training data and actual users’ data received. To effectively perform adaptation, textual data of users is typically stored on servers or their local devices, where downstream natural language processing (NLP) models can be directly…

Cited by 0SourceScholar
2024

SEED: A Simple and Effective 3D DETR in Point Clouds

ECCV 2024poster

"Recently, detection transformers (DETRs) have gradually taken a dominant position in 2D detection thanks to their elegant framework. However, DETR-based detectors for 3D point clouds are still difficult to achieve satisfactory performance. We argue that the main challenges are twofold: 1) How to ob…

2024

Toward Universal and Scalable Road Graph Partitioning for Efficient Multi-Robot Path Planning

IROS 2024

To date, multi-robot path planning has primarily been addressed by centralized solvers, typically aiming to maintain optimality. However, given its NP-hard nature, directly applying existing solvers in large and complex scenarios proves inefficient. A promising alternative lies in adopting a divide-

Cited by 1SourceScholar
2024

Traffic Flow Learning Enhanced Large-Scale Multi-Robot Cooperative Path Planning Under Uncertainties

ICRA 2024poster

Robotic systems with hundreds or even thousands of robots are widely implemented in logistic and industrial applications. In such systems, cooperative path planning is of great importance, as local congestion and motion conflict may greatly degrade system performance, especially in the presence of u…

Cited by 3SourceScholar
2023

A Simple Vision Transformer for Weakly Semi-supervised 3D Object Detection

ICCV 2023poster

Advanced 3D object detection methods usually rely on large-scale, elaborately labeled datasets to achieve good performance. However, labeling the bounding boxes for the 3D objects is difficult and expensive. Although semi-supervised (SS3D) and weakly-supervised 3D object detection (WS3D) methods can…

Cited by 29PDFScholar
2023

Auto-Weighted Multi-View Clustering for Large-Scale Data

AAAI 2023technical

Multi-view clustering has gained broad attention owing to its capacity to exploit complementary information across multiple data views. Although existing methods demonstrate delightful clustering performance, most of them are of high time complexity and cannot handle large-scale data. Matrix factori…

2023

Consistency of Multiple Kernel Clustering

ICML 2023poster

Consistency plays an important role in learning theory. However, in multiple kernel clustering (MKC), the consistency of kernel weights has not been sufficiently investigated. In this work, we fill this gap with a non-asymptotic analysis on the consistency of kernel weights of a novel method termed…

Cited by 9SourcePDFScholar
2023

DDS3D: Dense Pseudo-Labels with Dynamic Threshold for Semi-Supervised 3D Object Detection

ICRA 2023poster

In this paper, we present a simple yet effective semi-supervised 3D object detector named DDS3D. Our main contributions have two-fold. On the one hand, different from previous works using Non-Maximal Suppression (NMS) or its variants for obtaining the sparse pseudo labels, we propose a dense pseudo-…

Cited by 19SourcecodeScholar
2023

E-NER: Evidential Deep Learning for Trustworthy Named Entity Recognition

ACL 2023findings

Most named entity recognition (NER) systems focus on improving model performance, ignoring the need to quantify model uncertainty, which is critical to the reliability of NER systems in open environments. Evidential deep learning (EDL) has recently been proposed as a promising solution to explicitly…

2023

Gyro-Net: IMU Gyroscopes Random Errors Compensation Method Based on Deep Learning

RA-L 2023

To solve the problem of inaccurate orientation estimation after long-term operations of the Inertial Measurement Unit (IMU), we present a learning-based method (called Gyro-Net) to estimate and compensate for IMU gyroscope random errors. We firstly introduce a semi-dense network structure, which ext

Cited by 23SourceScholar
2023

Improved Event-Based Dense Depth Estimation via Optical Flow Compensation

ICRA 2023poster

Event cameras have the potential to overcome the limitations of classical computer vision in real-world applications. Depth estimation is a crucial step for high-level robotics tasks and has attracted much attention from the community. In this paper, we propose an event-based dense depth estimation…

Cited by 7SourceScholar
2023

Let the Data Choose: Flexible and Diverse Anchor Graph Fusion for Scalable Multi-View Clustering

AAAI 2023technical

In the past few years, numerous multi-view graph clustering algorithms have been proposed to enhance the clustering performance by exploring information from multiple views. Despite the superior performance, the high time and space expenditures limit their scalability. Accordingly, anchor graph lear…

2023

Lyapunov Constrained Safe Reinforcement Learning for Multicopter Visual Servoing

IROS 2023poster

Traditional methods based on Lyapunov analysis and learning-based approaches such as reinforcement learning (RL) are two powerful tools in visual servo tasks. Traditional methods are interpretable and their stability can be guar-anteed by Lyapunov analysis. However, they tend to have a high dependen…

Cited by 0SourceScholar
2023

Multi-Layer Feature Division Transferable Adversarial Attack

ICASSP 2023accepted

Improving the transferability of adversarial examples for the purpose of attacking unknown black-box models has been intensively studied. In particular, feature-level transfer-based attacks, which destroy the intermediate feature outputs of source models, are proven to generate more transferable adv…

Cited by 0SourceScholar
2023

PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval Augmentation

EMNLP 2023long findings

We introduce PersonaLM - Domain-distributed Span-Aggregated K-nearest N-gram retrieval augmentation to improve language modeling for Automatic Speech Recognition (ASR) personalization. PersonaLM leverages contextually similar n-gram word frequencies for recognizing rare word patterns associated with…

Cited by 0SourceScholar
2023

Query-based Temporal Fusion with Explicit Motion for 3D Object Detection

NeurIPS 2023poster

Effectively utilizing temporal information to improve 3D detection performance is vital for autonomous driving vehicles. Existing methods either conduct temporal fusion based on the dense BEV features or sparse 3D proposal features. However, the former does not pay more attention to foreground objec…

2023

RLSAC: Reinforcement Learning Enhanced Sample Consensus for End-to-End Robust Estimation

ICCV 2023poster

Robust estimation is a crucial and still challenging task, which involves estimating model parameters in noisy environments. Although conventional sampling consensus-based algorithms sample several times to achieve robustness, these algorithms cannot use data features and historical information effe…

Cited by 6PDFcodeScholar
2023

RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud Registration

ICCV 2023poster

Although point cloud registration has achieved remarkable advances in object-level and indoor scenes, large-scale registration methods are rarely explored. Challenges mainly arise from the huge point number, complex distribution, and outliers of outdoor LiDAR scans. In addition, most existing regist…

Cited by 58PDFcodeScholar
2023

See What the Robot Can't See: Learning Cooperative Perception for Visual Navigation

IROS 2023poster

We consider the problem of navigating a mobile robot towards a target in an unknown environment that is endowed with visual sensors, where neither the robot nor the sensors have access to global positioning information and only use first-person- view images. In order to overcome the need for positio…

Cited by 4SourcecodeScholar
2023

Self-supervised Multi-frame Monocular Depth Estimation with Pseudo-LiDAR Pose Enhancement

ICRA 2023poster

Depth estimation is one of the most important tasks in scene understanding. In the existing joint self-supervised learning approaches of depth-pose estimation, depth estimation and pose estimation networks are independent of each other. They only use the adjacent image frames for pose estimation and…

Cited by 5SourceScholar
2023

StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object Detection

AAAI 2023technical

In this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the response level, which is usually overlooked in 3D object detection distillation. Th…

Cited by 10SourcePDFScholar
2023

TransLO: A Window-Based Masked Point Transformer Framework for Large-Scale LiDAR Odometry

AAAI 2023technical

Recently, transformer architecture has gained great success in the computer vision community, such as image classification, object detection, etc. Nonetheless, its application for 3D vision remains to be explored, given that point cloud is inherently sparse, irregular, and unordered. Furthermore, ex…

2023

You Only Look Bottom-Up for Monocular 3D Object Detection

RA-L 2023

Monocular 3D Object Detection is an essential task for autonomous driving. Meanwhile, accurate 3D object detection from pure images is very challenging due to the loss of depth information. Most existing image-based methods infer objects' location in 3D space based on their 2D sizes on the image pla

Cited by 5SourceScholar
2022

Hilbert Distillation for Cross-Dimensionality Networks

NeurIPS 2022accept

3D convolutional neural networks have revealed superior performance in processing volumetric data such as video and medical imaging. However, the competitive performance by leveraging 3D networks results in huge computational costs, which are far beyond that of 2D networks. In this paper, we propose…

2022

Neural-FST Class Language Model for End-to-End Speech Recognition

ICASSP 2022accepted

We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transducers (FSTs) in a mathematically consistent framework. Our method utilizes a background NNLM which models generic backgroun…

Cited by 0SourceScholar
2022

What Matters for 3D Scene Flow Network

ECCV 2022poster

"3D scene flow estimation from point clouds is a low-level 3D motion perception task in computer vision. Flow embedding is a commonly used technique in scene flow estimation, and it encodes the point motion between two consecutive frames. Thus, it is critical for the flow embeddings to capture the c…

2021

A Registration-aided Domain Adaptation Network for 3D Point Cloud Based Place Recognition

IROS 2021poster

In the field of large-scale SLAM for autonomous driving and mobile robotics, 3D point cloud based place recognition has aroused significant research interest due to its robustness to changing environments with drastic daytime and weather variance. However, it is time-consuming and effort-costly to o…

Cited by 11SourceScholar
2021

Compact Pneumatic Clutch With Integrated Stiffness Variation and Position Feedback

RA-L 2021

Stiffness variation and real-time position feedback are critical for any robotic system but most importantly for active and wearable devices to interact with the user and environment. Currently, for compact sizes, there is a lack of solutions bringing high-fidelity feedback and maintaining design an

Cited by 9SourceScholar
2021

Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point Estimation

ICCV 2021poster

Existing vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts…

Cited by 9PDFScholar
2021

Learning To Identify Correct 2D-2D Line Correspondences on Sphere

CVPR 2021poster

Given a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both stru…

Cited by 4PDFScholar
2021

Message-Aware Graph Attention Networks for Large-Scale Multi-Robot Path Planning

RA-L 2021

The domains of transport and logistics are increasingly relying on autonomous mobile robots for the handling and distribution of passengers or resources. At large system scales, finding decentralized path planning and coordination solutions is key to efficient system performance. Recently, Graph Neu

Cited by 184SourcecodeScholar
2021

Multifunctional Robotic Glove with Active-Passive Training Modes for Hand Rehabilitation and Assistance

IROS 2021poster

Soft robotic gloves have shown great advantages in assisting individuals with hand pathologies to perform continuous exercises to restore their hand functions, which could considerably accelerate the rehabilitation process and reduce the costs. However, single rehabilitation mode, difficulty in achi…

Cited by 5SourceScholar
2021

PWCLO-Net: Deep LiDAR Odometry in 3D Point Clouds Using Hierarchical Embedding Mask Optimization

CVPR 2021poster

A novel 3D point cloud learning model for deep LiDAR odometry, named PWCLO-Net, using hierarchical embedding mask optimization is proposed in this paper. In this model, the Pyramid, Warping, and Cost volume (PWC) structure for the LiDAR odometry task is built to refine the estimated pose in a coarse…

Cited by 81PDFcodeScholar
2021

Task Aligned Generative Meta-learning for Zero-shot Learning

AAAI 2021technical

Zero-shot learning (ZSL) refers to the problem of learning to classify instances from novel classes (unseen) that are absent in the training set (seen). Most ZSL methods infer the correlation between visual features and attributes to train the classifier for unseen classes. They may have a strong bi…

Cited by 47SourcePDFScholar
2021

Task-Space Trajectory Tracking Control for Coordinated Manipulation Using Sampled Coupling Data

RA-L 2021

This letter studies the task-space synchronization control problem of networked manipulators which are commanded to track the desired task-space trajectories to achieve the coordinated manipulation transportation tasks in industrial and logistic applications. To guarantee the practical applicability

Cited by 8SourceScholar
2021

When and Why a Model Fails? A Human-in-the-loop Error Detection Framework for Sentiment Analysis

NAACL 2021industry

Although deep neural networks have been widely employed and proven effective in sentiment analysis tasks, it remains challenging for model developers to assess their models for erroneous predictions that might exist prior to deployment. Once deployed, emergent errors can be hard to identify in predi…

Cited by 14SourcePDFScholar
2020

A Synchronization Approach for Achieving Cooperative Adaptive Cruise Control Based Non-Stop Intersection Passing

ICRA 2020poster

Cooperative adaptive cruise control (CACC) of intelligent vehicles contributes to improving cruise control performance, reducing traffic congestion, saving energy and increasing traffic flow capacity. In this paper, we resolve the CACC problem from the viewpoint of synchronization control, our main…

Cited by 7SourceScholar
2020

An Empirical Study of Transformer-Based Neural Language Model Adaptation

ICASSP 2020accepted

We explore two adaptation approaches of deep Transformer based neural language models (LMs) for automatic speech recognition. The first approach is a pretrain-finetune framework, where we first pretrain a Transformer LM on a large-scale text corpus from scratch and then adapt it to relatively small…

Cited by 32SourceScholar
2020

CUHK-AHU Dataset: Promoting Practical Self-Driving Applications in the Complex Airport Logistics, Hill and Urban Environments

IROS 2020poster

This paper presents a novel dataset targeting three types of challenging environments for autonomous driving, i.e., the industrial logistics environment, the undulating hill environment and the mixed complex urban environment. To the best of the author’s knowledge, similar dataset has not been publi…

Cited by 5SourceScholar
2020

EPNet: Enhancing Point Features with Image Semantics for 3D Object Detection

ECCV 2020poster

In this paper, we aim at addressing two critical issues in the 3D detection task, including the exploitation of multiple sensors (namely LiDAR point cloud and camera image), as well as the inconsistency between the localization and classification confidence. To this end, we propose a novel fusion mo…

2020

End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondences

IROS 2020poster

3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learn…

Cited by 26SourcecodeScholar
2020

Global Context-enhanced Graph Convolutional Networks for Document-level Relation Extraction

COLING 2020main

Document-level Relation Extraction (RE) is particularly challenging due to complex semantic interactions among multiple entities in a document. Among exiting approaches, Graph Convolutional Networks (GCN) is one of the most effective approaches for document-level RE. However, traditional GCN simply…

2020

Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World

ECCV 2020poster

Atlanta world holds for the scenes composed of a vertical dominant direction and several horizontal dominant directions. Vanishing point (VP) is the intersection of the image lines projected from parallel 3D lines. In Atlanta world, given a set of image lines, we aim to cluster them by the unknown-b…

Cited by 17SourcePDFScholar
2020

Mobile Robot Path Planning in Dynamic Environments Through Globally Guided Reinforcement Learning

RA-L 2020

Path planning for mobile robots in large dynamic environments is a challenging problem, as the robots are required to efficiently reach their given goals while simultaneously avoiding potential conflicts with other robots or dynamic objects. In the presence of dynamic obstacles, traditional solution

Cited by 321SourceScholar
2020

Online Trajectory Planning for an Industrial Tractor Towing Multiple Full Trailers

ICRA 2020poster

This paper presents a novel solution for online trajectory planning of a full-size tractor-trailers vehicle composed of a car-like tractor and arbitrary number of passive full trailers. The motion planning problem for such systems was rarely addressed due to the complex nonlinear dynamics. A simulat…

Cited by 16SourceScholar
2020

Robust Dynamic State Estimation for Lateral Control of an Industrial Tractor Towing Multiple Passive Trailers

IROS 2020poster

In this paper, we propose a dynamic state estimation framework for lateral control of a heavy tractor-trailers system using only mass-produced low-cost sensors. This issue is challenging since the lateral velocity of the lead tractor is difficult to measure directly. The performance of existing dyna…

Cited by 0SourceScholar
2020

Robust Path Following of the Tractor-Trailers System in GPS-Denied Environments

RA-L 2020

This letter reports a general path following framework for the tractor-trailers system in Global Positioning System (GPS)-denied environments. Compared to existing methods, this approach prioritizes a robust, cost-optimized, and easy-to-implement solution. First, to achieve accurate path following,

Cited by 27SourceScholar
2020

Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual Odometry

ICRA 2020poster

Given a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In c…

Cited by 5SourceScholar
2019

A Hierarchical Framework for Coordinating Large-Scale Robot Networks

ICRA 2019poster

In this paper, we study the cooperative path planning and motion coordination problems of the multi-robot system with large number of robots, aiming for practical applications in robotic warehouses and automated transportation systems. Particularly, we solve the life-long planning problem and guaran…

Cited by 15SourceScholar
2019

LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment Analysis

ICCV 2019poster

Point cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named L…

Cited by 345PDFScholar
2019

Leveraging Structural Regularity of Atlanta World for Monocular SLAM

ICRA 2019poster

A wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocul…

Cited by 48SourceScholar
2019

Modelling and Dynamic Tracking Control of Industrial Vehicles with Tractor-trailer Structure

IROS 2019poster

Existing works on control of tractor-trailers systems only consider the kinematics model without taking dynamics into account. Also, most of them treat the issue as a pure control theory problem whose solutions are difficult to implement. This paper presents a trajectory tracking control approach fo…

Cited by 21SourceScholar
2019

Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan World

ICCV 2019oral

The image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The…

Cited by 37PDFScholar
2019

Retrieval-based Localization Based on Domain-invariant Feature Learning under Changing Environments

IROS 2019poster

Visual localization is a crucial problem in mobile robotics and autonomous driving. One solution is to retrieve images with known pose from a database for the localization of query images. However, in environments with drastically varying conditions (e.g. illumination changes, seasons, occlusion, dy…

Cited by 30SourcecodeScholar
2019

SeqLPD: Sequence Matching Enhanced Loop-Closure Detection Based on Large-Scale Point Cloud Description for Self-Driving Vehicles

IROS 2019poster

Place recognition and loop-closure detection are main challenges in the localization, mapping and navigation tasks of self-driving vehicles. In this paper, we solve the loop-closure detection problem by incorporating the deep-learning based point cloud description method and the coarse-to-fine seque…

Cited by 70SourceScholar
2019

Vision-Based Dynamic Control of Car-Like Mobile Robots

ICRA 2019poster

Most existing controllers for Car-Like Mobile Robots (CLMR) are designed to handle dynamic effects by decoupling speed and steering controls, also assume that full states are accessible, which are unrealistic for real-world applications. This paper presents a combined speed and steering control syst…

Cited by 9SourceScholar
2018

A Failure-Tolerant Approach to Synchronous Formation Control of Mobile Robots Under Communication Delays

ICRA 2018poster

Robot malfunction is inevitable in practical applications of the robot formation control due to uncontrolled crashing, system malfunction or communication loss. In this paper, we study the synchronous formation control problem in the presence of robot malfunctions. Our main idea is to improve the ne…

Cited by 7SourceScholar
2018

Stabilize an Unsupervised Feature Learning for LiDAR-based Place Recognition

IROS 2018poster

Place recognition is one of the major challenges for the LiDAR-based effective localization and mapping task. Traditional methods are usually relying on geometry matching to achieve place recognition, where a global geometry map need to be restored. In this paper, we accomplish the place recognition…

Cited by 21SourceScholar
2018

Vision-Based State Estimation and Trajectory Tracking Control of Car-Like Mobile Robots with Wheel Skidding and Slipping

IROS 2018poster

Most existing trajectory tracking controllers are based on non-skidding and non-slipping assumptions, also assume that full states are accessible, which is unrealistic for real-world applications due to tire-road interaction. This paper presents a novel vision-based approach to achieve high performa…

Cited by 12SourceScholar
2015

A gradient-based self-healing algorithm for mobile robot formation

IROS 2015poster

In this paper, we investigate the self-healing problem of mobile robot formation after some robots have been damaged, and present a gradient-based algorithm which enables mobile robots to restore the topology of the formation through local interactions among neighboring robots. Firstly, in order to…

Cited by 10SourceScholar