← Search

Yanyong Zhang

54 accepted papers

2026

C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognition

CVPR 2026

LiDAR-based place recognition is highly sensitive to rain, snow, and fog, where scattering and attenuation distort geometric structure and intensity. We tackle this problem with Conditional Latent Velocity Field (C-LaV) denoising, which restores weather-robust representations before retrieval. Singl

Cited by 0SourceScholar
2026

EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

ICML 2026spotlight

While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity…

Cited by 0SourceScholar
2026

Learning Surgical Robotic Manipulation with 3D Spatial Priors

CVPR 2026

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the d

Cited by 0SourceScholar
2026

On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models

ICML 2026poster

Entropy serves as a critical metric for measuring the diversity of outputs generated by large language models (LLMs), providing valuable insights into their exploration capabilities. While recent studies increasingly focus on monitoring and adjusting entropy to better balance exploration and exploit…

Cited by 0SourceScholar
2026

RaCFusion: Improving Camera-Based 3D Object Detection via Radar-Assisted Hierarchical Refinement

RA-L 2026

Cameras and radar sensors are complementary in 3D object detection in that cameras specialize in capturing an object's visual information while radar provides spatial information and velocity hints. Existing radar-camera fusion methods often employ a symmetrical architecture that processes inputs fr

Cited by 0SourcecodeScholar
2026

SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs

CVPR 2026

3D Large Vision-Language Models (3D LVLMs) built upon Large Language Models (LLMs) have achieved remarkable progress across various multimodal tasks. However, their inherited position-dependent modeling mechanism, Rotary Position Embedding (RoPE), remains suboptimal for 3D multimodal understanding.

Cited by 0SourceScholar
2026

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models

CVPR 2026

The recent success of reinforcement learning (RL) in large reasoning models has inspired the growing adoption of RL for post-training Multimodal Large Language Models (MLLMs) to enhance their visual reasoning capabilities. Although many studies have reported improved performance, it remains unclear

Cited by 0SourceScholar
2026

Versatile Vision-Language Model for 3D Computed Tomography

AAAI 2026technical

Representation learning serves as a foundational component of medical vision-language models (MVLMs), enabling cross-modal alignment, semantic consistency, and enhanced generalization capabilities for downstream tasks. As generalist models rapidly evolve, there is a pressing need to unify diverse do

Cited by 0SourcePDFScholar
2026

VistaDepth: Improving Far-Range Depth Estimation With Spectral Modulation and Adaptive Reweighting

RA-L 2026

Monocular depth estimation infers per-pixel depth from a single RGB image. It remains particularly challenging in far-range regions, where sparse observations and long-tailed depth distributions bias learning toward near-range content. Diffusion models offer a promising alternative to discriminative

Cited by 0SourceScholar
2025

CAFE-AD: Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving

ICRA 2025

Imitation learning based planning tasks on the nuPlan dataset have gained great interest due to their potential to generate human-like driving behaviors. However, open-loop training on the nuPlan dataset tends to cause causal confusion during closed-loop testing, and the dataset also presents a long

Cited by 2SourcecodeScholar
2025

CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map Representation

ICRA 2025

SLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks

Cited by 2SourceScholar
2025

FACE: A General Framework for Mapping Collaborative Filtering Embeddings into LLM Tokens

NeurIPS 2025poster

Recently, large language models (LLMs) have been explored for integration with collaborative filtering (CF)-based recommendation systems, which are crucial for personalizing user experiences. However, a key challenge is that LLMs struggle to interpret the latent, non-semantic embeddings produced by…

Cited by 0SourcecodeScholar
2025

GARD: A Geometry-Informed and Uncertainty-Aware Baseline Method for Zero-Shot Roadside Monocular Object Detection

RA-L 2025

Roadside camera-based perception methods are in high demand for developing efficient vehicle-infrastructure collaborative perception systems. By focusing on object-level depth prediction, we explore the potential benefits of integrating environmental priors into such systems and propose a geometry-b

Cited by 1SourceScholar
2025

GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions

ICCV 2025poster

Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities of the large language models (LLMs) to establish mappings between expressions and targets, allowing robots to comprehend…

2025

Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots

ICML 2025poster

Autoregressive models have emerged as a powerful generative paradigm for visual generation. The current de-facto standard of next token prediction commonly operates over a single-scale sequence of dense image tokens, and is incapable of utilizing global context especially for early tokens prediction…

2025

Improving Efficiency of Answer Set Planning with Rough Solutions from Large Language Models for Robotic Task Planning

IJCAI 2025

Answer Set Programming (ASP) planning can be used to refine the rough solutions generated by Large Language Models (LLMs) to handle specific restrictions of actions, i.e., reconstruct the rough solutions to be executable, for robotic task planning. However, it is still challenging to efficiently sol

2025

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

ICCV 2025poster

Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos for tasks like perception and planning. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level temporal consistency…

Cited by 0SourcePDFScholar
2025

Language Adaptation of Large Language Models: An Empirical Study on LLaMA2

COLING 2025main

There has been a surge of interest regarding language adaptation of Large Language Models (LLMs) to enhance the processing of texts in low-resource languages. While traditional language models have seen extensive research on language transfer, modern LLMs still necessitate further explorations in la…

2025

MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited Overlap

ICRA 2025

Large-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) c

Cited by 0SourceScholar
2025

Modalities Contribute Unequally: Enhancing Medical Multi-modal Learning through Adaptive Modality Token Re-balancing

ICML 2025poster

Medical multi-modal learning requires an effective fusion capability of various heterogeneous modalities. One vital challenge is how to effectively fuse modalities when their data quality varies across different modalities and patients. For example, in the TCGA benchmark, the performance of the same…

Cited by 0SourcePDFScholar
2025

OG-Gaussian: Occupancy Based Street Gaussians for Autonomous Driving

ICRA 2025

Accurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiD

Cited by 5SourceScholar
2025

OccMamba: Semantic Occupancy Prediction with State Space Models

CVPR 2025poster

Training deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability…

2025

Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures

RA-L 2025

When planning for autonomous driving, it is crucial to consider essential traffic elements such as lanes, intersections, traffic regulations, and dynamic agents. However, they are often overlooked by the traditional end-to-end planning methods, likely leading to inefficiencies and non-compliance wit

Cited by 1SourceScholar
2025

RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion

CVPR 2025poster

We propose Radar-Camera fusion transformer (RaCFormer) to boost the accuracy of 3D object detection by the following insight. The Radar-Camera fusion in outdoor 3D scene perception is capped by the image-to-BEV transformation-if the depth of pixels is not accurately estimated, the naive combination…

2025

S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene Reconstruction

ICCV 2025poster

Recently, 3D Gaussian Splatting (3DGS) has reshaped the field of photorealistic 3D reconstruction, achieving impressive rendering quality and speed. However, when applied to large-scale street scenes, existing methods suffer from rapidly escalating per-viewpoint reconstruction costs as scene size in…

2025

STDArm: Transfer Visuomotor Policy From Static Data Training to Dynamic Robot Manipulation

RSS 2025poster

Learning visuomotor policy from human demonstrations serves as an effective method for robots to acquire complex tasks. However, data collection on mobile platforms such as drones is extremely challenging, resulting in most research being conducted with robots in stationary conditions for data colle…

Cited by 0PDFScholar
2025

SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images

ICCV 2025poster

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods…

Cited by 0SourcePDFScholar
2024

BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation

ICASSP 2024accepted

Recent research has demonstrated the advantages of Bird’s-eye-view (BEV) representation in the field of 3D perception. However, due to the lack of height information, BEV representation alone is insufficient to accurately reconstruct the complete surrounding 3D scene. On the other hand, voxel repres…

Cited by 0SourceScholar
2024

CRPlace: Camera-Radar Fusion with BEV Representation for Place Recognition

IROS 2024poster

The integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies o…

Cited by 4SourceScholar
2024

CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network

ICRA 2024poster

The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different…

Cited by 14SourceScholar
2024

DGR: A General Graph Desmoothing Framework for Recommendation via Global and Local Perspectives

IJCAI 2024poster

Graph Convolutional Networks (GCNs) have become pivotal in recommendation systems for learning user and item embeddings by leveraging the user-item interaction graph's node information and topology. However, these models often face the famous over-smoothing issue, leading to indistinct user and item…

2024

EdgeCalib: Multi-Frame Weighted Edge Features for Automatic Targetless LiDAR-Camera Calibration

RA-L 2024

In multimodal perception systems, achieving precise extrinsic calibration between LiDAR and camera is of critical importance. However, the pre-calibrated extrinsic parameters may gradually drift during operation, leading to a decrease in the accuracy of the perception system. It is challenging to ad

Cited by 20SourceScholar
2024

FARFusion: A Practical Roadside Radar-Camera Fusion System for Far-Range Perception

RA-L 2024

Far-range perception through roadside sensors is crucial to the effectiveness of intelligent transportation systems. The main challenge of far-range perception is due to the difficulty of performing accurate object detection and tracking under far distances <italic xmlns:mml="http://www.w3.org/1998/

Cited by 24SourceScholar
2024

Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization

ICASSP 2024accepted

In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adapti…

Cited by 0SourceScholar
2024

LDP: A Local Diffusion Planner for Efficient Robot Navigation and Collision Avoidance

IROS 2024poster

The conditional diffusion model has been demonstrated as an efficient tool for learning robot policies, owing to its advancement to accurately model the conditional distribution of policies. The intricate nature of real-world scenarios, characterized by dynamic obstacles and maze-like structures, un…

Cited by 13SourceScholar
2024

MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes

IROS 2024poster

Localization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this work, we present MM-Gaussian, a LiDAR-camera multimodal fusion s…

Cited by 12SourceScholar
2024

OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving

ICRA 2024poster

Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupa…

Cited by 8SourcecodeScholar
2024

Profiling Power Consumption in Low-Speed Autonomous Guided Vehicles

RA-L 2024

The increasing demand for automation has led to a rise in the use of low-speed Autonomous guided vehicles (AGVs). However, AGVs rely on batteries for their power source, which limits their operational time and affects their overall performance. To optimize their energy usage and enhance their batter

Cited by 7SourceScholar
2024

SDAC: A Multimodal Synthetic Dataset for Anomaly and Corner Case Detection in Autonomous Driving

AAAI 2024technical

Nowadays, closed-set perception methods for autonomous driving perform well on datasets containing normal scenes. However, they still struggle to handle anomalies in the real world, such as unknown objects that have never been seen while training. The lack of public datasets to evaluate the model pe…

Cited by 3SourcePDFScholar
2024

mmPlace: Robust Place Recognition With Intermediate Frequency Signal of Low-Cost Single-Chip Millimeter Wave Radar

RA-L 2024

Place recognition is crucial for tasks like loop-closure detection and re-localization. Single-chip millimeter wave radar (single-chip radar in short) emerges as a low-cost sensor option for place recognition, with the advantage of insensitivity to degraded visual environments. However, it encounter

Cited by 9SourceScholar
2023

Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection

CVPR 2023poster

LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear.…

2023

CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection

NeurIPS 2023poster

Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution l…

Cited by 6SourcePDFScholar
2023

Reinforcement Learning for Robot Navigation with Adaptive Forward Simulation Time (AFST) in a Semi-Markov Model

IROS 2023poster

Deep reinforcement learning (DRL) algorithms have proven effective in robot navigation, especially in unknown environments, by directly mapping perception inputs into robot control commands. However, most existing methods ignore the local minimum problem in navigation and thereby cannot handle compl…

Cited by 0SourcecodeScholar
2022

${\mathsf{EZFusion}}$: A Close Look at the Integration of LiDAR, Millimeter-Wave Radar, and Camera for Accurate 3D Object Detection and Tracking

RA-L 2022

A recent trend is to combine multiple sensors ( <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e.</i> , cameras, LiDARs and millimeter-wave Radars) to achieve robust multi-modal perception for autonomous systems such as self-driving vehicles. Alth

Cited by 11SourceScholar
2022

PF-MOT: Probability Fusion Based 3D Multi-Object Tracking for Autonomous Vehicles

ICRA 2022poster

3D Multi-Object Tracking (MOT) plays a crucial role in efficient and safe operation of automatic driving, especially in scenarios of occlusion or poor visibility. Most 3D MOT methods leverage only positional distance, which is insufficient for scenes with high density of objects or drastic changes i…

Cited by 13SourceScholar
2022

PFilter: Building Persistent Maps through Feature Filtering for Fast and Accurate LiDAR-based SLAM

IROS 2022poster

Simultaneous localization and mapping (SLAM) based on laser sensors has been widely adopted by mobile robots and autonomous vehicles. These SLAM systems are required to support accurate localization with limited computational resources. In particular, point cloud registration, i.e., the process of m…

Cited by 21SourceScholar
2021

DRQN-based 3D Obstacle Avoidance with a Limited Field of View

IROS 2021poster

In this paper, we propose a map-based end-to-end DRL approach for three-dimensional (3D) obstacle avoidance in a partially observed environment, which is applied to achieve autonomous navigation for an indoor mobile robot using a depth camera with a narrow field of view. We first train a neural netw…

Cited by 10SourceScholar
2021

Towards an Online RRT-based Path Planning Algorithm for Ackermann-steering Vehicles

ICRA 2021poster

It is challenging to develop an online path planning algorithm for Ackermann-steering vehicles to find collision-free and kinematically-feasible paths, that is efficient for dense environments, adaptable to various environments, and suitable for environments with narrow passages. In this paper, we p…

Cited by 10SourcecodeScholar
2021

Voxel R-CNN: Towards High Performance Voxel-based 3D Object Detection

AAAI 2021technical

Recent advances on 3D object detection heavily rely on how the 3D data are represented, i.e., voxel-based or point-based representation. Many existing high performance 3D detectors are point-based because this structure can better retain precise point positions. Nevertheless, point-level features le…