← Search

Zheng Fang

62 accepted papers

2026

4DRaL: Bridging 4D Radar with LiDAR for Place Recognition Using Knowledge Distillation

ICRA 2026poster

Place recognition is crucial for loop closure detection and global localization in robotics. Although mainstream algorithms typically rely on cameras and LiDAR, these sensors are susceptible to adverse weather conditions. Fortunately, the recently developed 4D millimeter-wave radar (4D radar) offers…

2026

DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image Generation

CVPR 2026

Text-to-image diffusion models have achieved remarkable progress, generating visually realistic and semantically coherent images from textual prompts. However, natural language alone lacks the precision required for design-centric applications that demand strict spatial and structural fidelity--part

Cited by 0SourceScholar
2026

Large Language Model Unlearning for Source Code

AAAI 2026technical

While Large Language Models (LLMs) excel at code generation, their inherent tendency toward verbatim memorization of training data introduces critical risks like copyright infringement, insecurity emission, and deprecated API utilization, etc. A straightforward yet promising defense is unlearning, i

Cited by 0SourcePDFScholar
2026

MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework

ICML 2026poster

Inferring molecular structures from multimodal spectroscopic measurements requires integrating complementary yet highly heterogeneous signals. However, the common paradigm of directly concatenating multispectral sequences can exhibit anomalous performance degradation, primarily due to pronounced het…

Cited by 0SourceScholar
2026

MetaToolAgent: Towards Generalizable Tool Usage in LLMs through Meta-Learning

ICASSP 2026oral

Tool learning is increasingly important for large language models (LLMs) to effectively coordinate and utilize a diverse set of tools in order to solve complex real-world tasks. By selecting and integrating appropriate tools, LLMs extend their capabilities beyond pure language understanding to perfo…

Cited by 0SourcePDFScholar
2026

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

ICML 2026poster

Jailbreak attacks on audio language models (ALMs) optimize audio perturbations to elicit unsafe generations, and they typically update the entire waveform densely throughout optimization. In this work, we investigate the necessity of such dense optimization by analyzing the structure of token-aligne…

Cited by 0SourceScholar
2026

Spike-EVPR: Deep Spiking Residual Networks With SNN-Tailored Representations for Event-Based Visual Place Recognition

RA-L 2026

Event cameras are ideal for visual place recognition (VPR) in challenging environments due to their high temporal resolution and high dynamic range. However, existing methods convert sparse events into dense frame-like representations for Artificial Neural Networks (ANNs), ignoring event sparsity an

Cited by 0SourceScholar
2026

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

ICASSP 2026poster

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehension, necessitating extensive processing of all video frames before any queries ca…

Cited by 0SourcePDFScholar
2026

Towards 3D Object-Centric Feature Learning for Semantic Scene Completion

AAAI 2026technical

Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details,

Cited by 0SourcePDFScholar
2026

Transferring Policy of Offline Reinforcement Learning From Hybrid Dataset to Real World via Progressive Neural Network

RA-L 2026

Offline reinforcement learning (Offline RL) provides a compelling solution for applying RL in high-risk or resourceconstrained real-world domains such as healthcare, autonomous driving, and robotic manipulation, where online exploration can be unsafe or impractical. However, Offline RL faces critica

Cited by 0SourceScholar
2026

Transferring Policy of Offline Reinforcement Learning from Hybrid Dataset to Real World Via Progressive Neural Network

ICRA 2026poster

Offline reinforcement learning (Offline RL) provides a compelling solution for applying RL in high-risk or resource-constrained real-world domains such as healthcare, autonomous driving, and robotic manipulation. However, Offline RL faces critical challenges arising from limited data coverage and po…

Cited by 0SourceScholar
2025

A Coarse-to-Fine Event-based Framework for Camera Pose Relocalization with Spatio-Temporal Retrieval and Refinement Network

ICRA 2025

Most existing event-based camera pose relocalization (CPR) learning methods implicitly encode environmental information into network parameters to achieve end-to-end mapping from event stream to pose. However, these end-to-end CPR methods fail to utilize prior environmental information effectively.

Cited by 0SourceScholar
2025

Beyond Completion: A Foundation Model for General Knowledge Graph Reasoning

ACL 2025finding

In natural language processing (NLP) and computer vision (CV), the successful application of foundation models across diverse tasks has demonstrated their remarkable potential. However, despite the rich structural and textual information embedded in knowledge graphs (KGs), existing research of found…

2025

CAO-RONet: A Robust 4D Radar Odometry with Exploring More Information from Low-Quality Points

ICRA 2025

Recently, 4D millimetre-wave radar exhibits more stable perception ability than LiDAR and camera under adverse conditions (e.g. rain and fog). However, low-quality radar points hinder its application, especially the odometry task that requires a dense and accurate matching. To fully explore the pote

Cited by 3SourcecodeScholar
2025

EDE-Distill: Boosting Event-Based Monocular Depth Estimation Performance via Knowledge Distillation

RA-L 2025

Monocular depth estimation based on event cameras has attracted widespread attention of researchers as event-cameras, with their high dynamic range and temporal resolution, can offer enhanced environmental perception ability under challenging lighting conditions. However, due to the inherent texture

Cited by 1SourceScholar
2025

FlexControl: Computation-Aware Conditional Control with Differentiable Router for Text-to-Image Generation

ICML 2025poster

Spatial conditioning control offers a powerful way to guide diffusion‐based generative models. Yet, most implementations (e.g., ControlNet) rely on ad-hoc heuristics to choose which network blocks to control — an approach that varies unpredictably with different tasks. To address this gap, we propos…

2025

LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba

AAAI 2025technical

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information through the attention mechanism, enabling the 3D semantic o…

Cited by 1SourcePDFScholar
2025

Learning Null Geodesics for Gravitational Lensing Rendering in General Relativity

ICCV 2025poster

We present GravlensX, an innovative method for rendering black holes with gravitational lensing effects using neural networks. The methodology involves training neural networks to fit the spacetime around black holes and then employing these trained models to generate the path of light rays affected…

Cited by 0SourcePDFScholar
2025

Multi-Agent Generative Adversarial Interactive Self-Imitation Learning for AUV Formation Control and Obstacle Avoidance

RA-L 2025

Multiple autonomous underwater vehicles (multi-AUVs) can cooperatively accomplish tasks that a single AUV cannot complete. Recently, multi-agent reinforcement learning has been introduced to control of multi-AUV. However, designing efficient reward functions for various tasks of multi-AUV control is

Cited by 8SourceScholar
2025

Nonlinear Motion-Guided and Spatio-Temporal Aware Network for Unsupervised Event-Based Optical Flow

ICRA 2025

Event cameras have the potential to capture continuous motion information over time and space, making them well-suited for optical flow estimation. However, most existing learning-based methods for event-based optical flow adopt frame-based techniques, ignoring the spatio-temporal characteristics of

Cited by 0SourceScholar
2025

Optimize Battery Control: A Multi-Objective Evolutionary Ensemble Reinforcement Learning Approach

IJCAI 2025

The Dynamically Reconfigurable Battery (DRB) systems, which use high-speed power electronic switches to dynamically adjust battery interconnections in real-time, are critical to the performance of the battery pack. Traditional battery management strategies often fail to address multi-objective optim

Cited by 0SourcePDFScholar
2025

Robust Model-Free Path Tracking Algorithm for Hydraulic Center-Articulated Scooptrams

IROS 2025

This paper proposes a model-free steering control method to address the path tracking challenges of Hydraulic Center-articulated Scooptrams (HCS) in narrow underground mining environments. Due to the nonlinear and time-delay characteristics of the hydraulic steering system, the HCS exhibits response

Cited by 0SourceScholar
2025

Shape-Adaptive Planning and Control for a Deformable Quadrotor

IROS 2025

Drones have become essential in various applications, but conventional quadrotors face limitations in confined spaces and complex tasks. Deformable drones, which can adapt their shape in real-time, offer a promising solution to overcome these challenges, while also enhancing maneuverability and enab

Cited by 1SourceScholar
2025

StreamMOS: Streaming Moving Object Segmentation With Multi-View Perception and Dual-Span Memory

RA-L 2025

Moving object segmentation based on LiDAR is a crucial and challenging task for autonomous driving and mobile robotics. Most approaches explore spatio-temporal information from LiDAR sequences to predict moving objects in the current frame. However, they often focus on transferring temporal cues in

Cited by 5SourcecodeScholar
2025

Target-Aware Viewpoint Generation for Active Robotic Exploration in Unknown Environments

ICRA 2025

When entering an unfamiliar environment, animals usually sweep off their surroundings to identify points of interest. In search and rescue robotics, autonomous exploration requires both coarse mapping of unknown areas and detailed target detection, which poses a significant challenge in balancing th

Cited by 0SourceScholar
2025

Visual Localization with Offline Google Satellite Map-Assisted for Ground Vehicles in GNSS-Denied Environment

IROS 2025

Vehicle localization is a critical component in the planning and navigation of autonomous driving system. Generally, traditional vehicle localization methods rely on the Global Navigation Satellite System (GNSS) for self-localization. Unfortunately, GNSS can become unreliable and may fail in urban c

Cited by 0SourcecodeScholar
2024

ASML-VDIO: Visual-Depth-Inertial Odometry using Selected Accurate and Stable Multi-Modal Landmarks in Structural Environments

IROS 2024poster

In complex indoor structural scenes such as shopping centers and malls, camera pose estimation using pure point features is easy to fail due to the difficulty in extracting sufficient and stable point features from weak textures or dynamic environments. Recent works have attempted to address these c…

Cited by 0SourceScholar
2024

CWTM: Leveraging Contextualized Word Embeddings from BERT for Neural Topic Modeling

COLING 2024main

Most existing topic models rely on bag-of-words (BOW) representation, which limits their ability to capture word order information and leads to challenges with out-of-vocabulary (OOV) words in new documents. Contextualized word embeddings, however, show superiority in word sense disambiguation and e…

2024

DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

ACL 2024findings

How to evaluate the coding abilities of Large Language Models (LLMs) remains an open question. We find that existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of LLMs.To address the knowledge gap, we propose a new benchmark…

2024

EventRPG: Event Data Augmentation with Relevance Propagation Guidance

ICLR 2024poster

Event camera, a novel bio-inspired vision sensor, has drawn a lot of attention for its low latency, low power consumption, and high dynamic range. Currently, overfitting remains a critical problem in event-based classification tasks for Spiking Neural Network (SNN) due to its relatively weak spatial…

2024

EverySync: An Open Hardware Time Synchronization Sensor Suite for Common Sensors in SLAM

IROS 2024poster

Multi-sensor fusion systems have been widely applied in various fields, including mobile robot, simultaneous localization and mapping (SLAM), and autonomous driving. For a tightly coupled multi-sensor fusion system, strict time synchronization between sensors will improve the accuracy of the system.…

Cited by 1SourceScholar
2024

LM-Mapping: Large-Scale and Multi-Session Point Cloud Consistent Mapping

RA-L 2024

In the field of autonomous driving and mobile robotics, constructing high-precision prior maps is a significant problem. For large outdoor scenes, maps often need to be segmented or collected repeatedly. Issues such as sensor degradation and measurement error can result in map inconsistencies or loc

Cited by 4SourceScholar
2024

Model Composition for Multimodal Large Language Models

ACL 2024long

Recent developments in Multimodal Large Language Models (MLLMs) have shown rapid progress, moving towards the goal of creating versatile MLLMs that understand inputs from various modalities. However, existing methods typically rely on joint training with paired multimodal instruction data, which is…

2024

Observation Time Difference: an Online Dynamic Objects Removal Method for Ground Vehicles

ICRA 2024poster

In the process of urban environment mapping, the sequential accumulations of dynamic objects will leave a large number of traces in the map. These traces will usually have bad influences on the localization accuracy and navigation performance of the robot. Therefore, dynamic objects removal plays an…

Cited by 1SourcecodeScholar
2024

PARE: A Plane-Assisted Autonomous Robot Exploration Framework in Unknown and Uneven Terrain

IROS 2024poster

Identifying traversable areas is a critical task for unmanned vehicles exploring safely through unstructured environments. In practice, the ambiguity in perceiving terrain traversability usually brings great challenges for autonomous exploration in unknown and uneven terrain, which often leads to co…

Cited by 0SourceScholar
2024

SeqTrack3D: Exploring Sequence Information for Robust 3D Point Cloud Tracking

ICRA 2024poster

3D single object tracking (SOT) is an important and challenging task for the autonomous driving and mobile robotics. Most existing methods perform tracking between two consecutive frames while ignoring the motion patterns of the target over a series of frames, which would cause performance degradati…

Cited by 1SourcecodeScholar
2024

Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

NeurIPS 2024poster

In current deep learning tasks, Adam-style optimizers—such as Adam, Adagrad, RMSprop, Adafactor, and Lion—have been widely used as alternatives to SGD-style optimizers. These optimizers typically update model parameters using the sign of gradients, resulting in more stable convergence curves. The l…

Cited by 6SourcePDFScholar
2024

Trans-Rotor: An Active Omnidirectional Aerial-Ground Vehicle With Differential Gear Joint Transformation Mechanism

IROS 2024poster

Aerial-ground vehicles have shown great potential in various fields due to their superior mobility and outstanding endurance. However, most of morphing aerial-ground vehicles consider little about controllability and traversability in ground mode. We present a novel aerial-ground vehicle called Tran…

Cited by 0SourceScholar
2023

FE-Fusion-VPR: Attention-Based Multi-Scale Network Architecture for Visual Place Recognition by Fusing Frames and Events

RA-L 2023

Traditional visual place recognition (VPR), usually using standard cameras, is easy to fail due to glare or high-speed motion. By contrast, event cameras have the advantages of low latency, high temporal resolution, and high dynamic range, which can deal with the above issues. Nevertheless, event ca

Cited by 28SourceScholar
2023

FastRecon: Few-shot Industrial Anomaly Detection via Fast Feature Reconstruction

ICCV 2023poster

In industrial anomaly detection, data efficiency and the ability for fast migration across products become the main concerns when developing detection algorithms. Existing methods tend to be data-hungry and work in the one-model-one-category way, which hinders their effectiveness in real-world indus…

Cited by 58PDFScholar
2022

Adaptive Face Forgery Detection in Cross Domain

ECCV 2022poster

"It is necessary to develop effective face forgery detection methods with constantly evolving technologies in synthesizing realistic faces which raises serious risks on malicious face tampering. A large and growing body of literature has investigated deep learning-based approaches, especially those…

2022

Disentangled Learning of Stance and Aspect Topics for Vaccine Attitude Detection in Social Media

NAACL 2022long

Building models to detect vaccine attitudes on social media is challenging because of the composite, often intricate aspects involved, and the limited availability of annotated data. Existing approaches have relied heavily on supervised training that requires abundant annotations and pre-defined asp…

2022

Exploiting More Information in Sparse Point Cloud for 3D Single Object Tracking

RA-L 2022

3D single object tracking is a key task in 3D computer vision. However, the sparsity of point clouds makes it difficult to compute the similarity and locate the object, posing big challenges to the 3D tracker. Previous works tried to solve the problem and improved the tracking performance in some co

Cited by 28SourcecodeScholar
2022

Non-Autoregressive Chinese ASR Error Correction with Phonological Training

NAACL 2022long

Automatic Speech Recognition (ASR) is an efficient and widely used input method that transcribes speech signals into text. As the errors introduced by ASR systems will impair the performance of downstream tasks, we introduce a post-processing error correction method, PhVEC, to correct errors in text…

2021

Deep Differential Amplifier for Extractive Summarization

ACL 2021long

For sentence-level extractive summarization, there is a disproportionate ratio of selected and unselected sentences, leading to flatting the summary features when maximizing the accuracy. The imbalanced classification of summarization is inherent, which can’t be addressed by common algorithms easily…

2021

PTT: Point-Track-Transformer Module for 3D Single Object Tracking in Point Clouds

IROS 2021poster

3D single object tracking is a key issue for robotics. In this paper, we propose a transformer module called Point-Track-Transformer (PTT) for point cloud-based 3D single object tracking. PTT module contains three blocks for feature embedding, position encoding, and self-attention feature computatio…

Cited by 96SourcecodeScholar
2021

TEBNER: Domain Specific Named Entity Recognition with Type Expanded Boundary-aware Network

EMNLP 2021main

To alleviate label scarcity in Named Entity Recognition (NER) task, distantly supervised NER methods are widely applied to automatically label data and identify entities. Although the human effort is reduced, the generated incomplete and noisy annotations pose new challenges for learning effective n…

Cited by 15SourcePDFScholar
2021

Vanishing Point Aided LiDAR-Visual-Inertial Estimator

ICRA 2021poster

In this paper, we propose a vanishing point aided LiDAR-Visual-Inertial estimator to achieve real-time, low-drift and robust pose estimation. The proposed method is mainly composed of 3 sequential modules, namely IMU-aided vanishing point (VP) detection module, voxel-map based feature depth associat…

Cited by 19SourceScholar
2020

Multi-Way Multi-View Deep Autoencoder for Image Feature Learning with Multi-Level Graph Regularization

ICASSP 2020accepted

Multi-view feature learning has garnered much attention recently since many real world data are comprised of different representations or views. How to explore the consensus structure and eliminate the inconsistency noise in different views remains a challenging problem in multi-view feature learnin…

Cited by 0SourceScholar
2020

Neural Coding Strategies for Event-Based Vision Data

ICASSP 2020accepted

Neural coding schemes are powerful tools used within neuroscience. This paper introduces three different neural coding scheme formations for event-based vision data which are designed to emulate the neural behaviour exhibited by neurons under stimuli. Presented are phase-of-firing and two sparse neu…

Cited by 0SourceScholar
2020

TP-TIO: A Robust Thermal-Inertial Odometry with Deep ThermalPoint

IROS 2020poster

To achieve robust motion estimation in visually degraded environments, thermal odometry has been an attraction in the robotics community. However, most thermal odometry methods are purely based on classical feature extractors, which is difficult to establish robust correspondences in successive fram…

Cited by 49SourceScholar
2019

A Robust Laser-Inertial Odometry and Mapping Method for Large-Scale Highway Environments

IROS 2019poster

In this paper, we propose a novel laser-inertial odometry and mapping method to achieve real-time, low-drift and robust pose estimation in large-scale highway environments. The proposed method is mainly composed of four sequential modules, namely scan pre-processing module, dynamic object detection…

Cited by 98SourceScholar
2015

Real-time onboard 6DoF localization of an indoor MAV in degraded visual environments using a RGB-D camera

ICRA 2015poster

Real-time and reliable localization is a prerequisite for autonomously performing high-level tasks with micro aerial vehicles(MAVs). Nowadays, most existing methods use vision system for 6DoF pose estimation, which can not work in degraded visual environments. This paper presents an onboard 6DoF pos…

Cited by 38SourceScholar