← Search

Chenglu Wen

48 accepted papers

2026

FlashCap: Millisecond-Accurate Human Motion Capture via Flashing LEDs and Event-Based Vision

CVPR 2026

Precise motion timing (PMT) is crucial for swift motion analysis. A millisecond difference may determine victory or defeat in sports competitions. Despite substantial progress in human pose estimation (HPE), PMT remains largely overlooked by the HPE community due to the limited availability of high-

Cited by 0SourceScholar
2026

GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting

AAAI 2026technical

3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features into 3D Gaussian splatting, enabling open-vocabulary querie

Cited by 0SourcePDFScholar
2026

LEADER: Learning Reliable Local-to-Global Correspondences for LiDAR Relocalization

CVPR 2026

LiDAR relocalization has attracted increasing attention as it can deliver accurate 6-DoF pose estimation in complex 3D environments. Recent learning-based regression methods offer efficient solutions by directly predicting global poses without the need for explicit map storage. However, these method

Cited by 0SourcecodeScholar
2026

MAC-NeRF: Motion-Aware Curriculum Learning for Dynamic LiDAR NeRFs

ICML 2026poster

While LiDAR NeRFs excel in static environments, synthesizing dynamic scenes remains challenging as moving objects break multi-view consistency, causing conflicting supervision and ghosting artifacts across frames. Existing methods typically suffer from optimization difficulty from the start, struggl…

Cited by 0SourceScholar
2026

MAGE: Multi-scale Autoregressive Generation for Offline Reinforcement Learning

ICLR 2026poster

Generative models have gained significant traction in offline reinforcement learning (RL) due to their ability to model complex trajectory distributions. However, existing generation-based approaches still struggle with long-horizon tasks characterized by sparse rewards. Some hierarchical generation…

Cited by 0SourcecodeScholar
2026

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

CVPR 2026

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training. However, existing zero-shot methods that build explicit 3D scene

Cited by 0SourceScholar
2026

MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations

ICML 2026poster

Human motion inherently exhibits a sophisticated temporal hierarchical architecture, spanning from global low-frequency trajectories to local high-frequency dynamics. Inspired by this intrinsic property and the success of multi-scale autoregressive modeling in vision, we propose MotionMAR, a novel f…

Cited by 0SourceScholar
2026

OWL: Unsupervised 3D Object Detection by Occupancy Guided Warm-up and Large Model Priors Reasoning

AAAI 2026technical

Unsupervised 3D object detection leverages heuristic algorithms to discover potential objects, offering a promising route to reduce annotation costs in autonomous driving. Existing approaches mainly generate pseudo labels and refine them through self-training iterations. However, these pseudo-label

Cited by 0SourcePDFScholar
2026

TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection

CVPR 2026

Reliable navigation and decision-making of autonomous vehicles require both accurate localization and object detection. Traditionally, these two tasks are handled separately, leading to redundant computation and limited cross-task knowledge transfer. This paper proposes TACO, the first Task-Aware CO

Cited by 0SourcecodeScholar
2026

Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robots

CVPR 2026

Humanoid robots have achieved significant progress in motion generation and control, exhibiting movements that appear increasingly natural and human-like. Inspired by the Turing Test, we propose the Motion Turing Test, a framework that evaluates whether human observers can discriminate between human

Cited by 0SourceScholar
2026

V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perception

CVPR 2026

Modern autonomous vehicle perception systems are often constrained by occlusions, blind spots, and limited sensing range. While existing cooperative perception paradigms, such as Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure (V2I), have demonstrated their effectiveness in mitigating these c

Cited by 0SourcecodeScholar
2026

V2VLoc: Robust GNSS-Free Collaborative Perception via LiDAR Localization

AAAI 2026technical

Multi-agents rely on accurate poses to share and align observations, enabling a collaborative perception of the environment. However, traditional GNSS-based localization often fails in GNSS-denied environments, making consistent feature alignment difficult in collaboration. To tackle this challenge,

Cited by 0SourcePDFScholar
2026

WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object Detection

AAAI 2026technical

3D object detection is critical for autonomous driving, yet it remains fundamentally challenging to simultaneously maximize computational efficiency and capture long-range spatial dependencies.We observed that Mamba-based models, with their linear state-space design, capture long-range dependencies

Cited by 0SourcePDFScholar
2026

X-MoGe: A Cross-Modal Adaptation Framework with Mixture-of-Experts and Geometry Guidance for Heterogeneous Collaborative Perception

ICML 2026poster

Multi-agent collaborative perception improves perception range and robustness in autonomous driving. However, most existing methods assume homogeneous sensors and perception networks, which is unrealistic in real-world heterogeneous systems. Differences in sensing modalities and independently traine…

Cited by 0SourceScholar
2025

AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label Correction

AAAI 2025technical

Recently, Visual Foundation Models (VFMs) have shown a remarkable generalization performance in 3D perception tasks. However, their effectiveness in large-scale outdoor datasets remains constrained by the scarcity of accurate supervision signals, the extensive noise caused by variable outdoor cond…

2025

Boosting Adversarial Transferability through Augmentation in Hypothesis Space

CVPR 2025poster

Adversarial examples can mislead deep neural networks with subtle perturbations, causing them to make incorrect predictions. Notably, adversarial examples crafted for one model can also deceive other models, a phenomenon known as the transferability of adversarial examples. To improve transferabilit…

2025

ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World Coordinate

CVPR 2025highlight

Human Motion Recovery (HMR) research mainly focuses on ground-based motions such as running. The study on capturing climbing motion, an off-ground motion, is sparse. This is partly due to the limited availability of climbing motion datasets, especially large-scale and challenging 3D labeled datasets…

Cited by 0SourcePDFScholar
2025

DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based Refinement

CVPR 2025poster

LiDAR odometry is a critical module in autonomous driving systems, responsible for accurate localization by estimating the relative pose transformation between consecutive point cloud frames. However, existing studies frequently encounter challenges with unreliable pose estimation, due to the lack o…

2025

DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning

ICLR 2025poster

Diffusion models have been widely adopted in image and language generation and are now being applied to reinforcement learning. However, the application of diffusion models in offline cooperative Multi-Agent Reinforcement Learning (MARL) remains limited. Although existing studies explore this direct…

2025

L4DR: LiDAR-4DRadar Fusion for Weather-Robust 3D Object Detection

AAAI 2025technical

LiDAR-based 3D object detection is crucial for autonomous driving. However, due to the quality deterioration of LiDAR point clouds, it suffers from performance degradation in adverse weather conditions. Fusing LiDAR with the weatherrobust 4D radar sensor is expected to solve this problem; however, i…

2025

Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual Labels

CVPR 2025poster

Unsupervised 3D object detection serves as an important solution for offline 3D object annotation. However, due to the data sparsity and limited views, the clustering-based label fitting in unsupervised object detection often generates low-quality pseudo-labels. Multi-agent collaborative dataset, wh…

2025

LightLoc: Learning Outdoor LiDAR Localization at Light Speed

CVPR 2025poster

Scene coordinate regression achieves impressive results in outdoor LiDAR localization but requires days of training. Since training needs to be repeated for each new scene, long training times make these impractical for applications requiring time-sensitive system upgrades, such as autonomous drivin…

2025

Motal: Unsupervised 3D Object Detection by Modality and Task-specific Knowledge Transfer

ICCV 2025poster

The performance of unsupervised 3D object classification and bounding box regression relies heavily on the quality of initial pseudo-labels. Traditionally, the labels of classification and regression are represented by a single set of candidate boxes generated by motion or geometry heuristics. Howev…

Cited by 0SourcePDFScholar
2025

Pretend Benign: A Stealthy Adversarial Attack by Exploiting Vulnerabilities in Cooperative Perception

ICCV 2025poster

Recently, learning-based multi-agent cooperative perception has garnered widespread attention. However, the inherent vulnerabilities of neural networks, combined with the risks posed by cooperative communication as a wide-open backdoor, render these systems highly susceptible to adversarial attacks.…

Cited by 0SourcePDFScholar
2025

SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts

CVPR 2025highlight

Recently, sparsely-supervised 3D object detection has gained great attention, achieving performance close to fully-supervised 3D detectors while requiring only a few annotated instances. Nevertheless, these methods suffer challenges when accurate labels are extremely absent. In this paper, we propos…

2025

Seg2Box: 3D Object Detection by Point-Wise Semantics Supervision

AAAI 2025technical

LIDAR-based 3D object detection and semantic segmentation are critical tasks in 3D scene understanding. Traditional detection and segmentation methods supervise their models through bounding box labels and semantic mask labels. However, these two independent labels inherently contain significant red…

Cited by 0SourcePDFScholar
2024

Commonsense Prototype for Outdoor Unsupervised 3D Object Detection

CVPR 2024poster

The prevalent approaches of unsupervised 3D object detection follow cluster-based pseudo-label generation and iterative self-training processes. However the challenge arises due to the sparsity of LiDAR scans which leads to pseudo-labels with erroneous size and position resulting in subpar detection…

2024

DiffLoc: Diffusion Model for Outdoor LiDAR Localization

CVPR 2024poster

Absolute pose regression (APR) estimates global pose in an end-to-end manner achieving impressive results in learn-based LiDAR localization. However compared to the top-performing methods reliant on 3D-3D correspondence matching APR's accuracy still has room for improvement. We recognize APR's lack…

2024

FedPFT: Federated Proxy Fine-Tuning of Foundation Models

IJCAI 2024poster

Adapting Foundation Models (FMs) for down- stream tasks through Federated Learning (FL) emerges a promising strategy for protecting data privacy and valuable FMs. Existing methods fine- tune FM by allocating sub-FM to clients in FL, however, leading to suboptimal performance due to insufficient tuni…

2024

Federated Graph Learning for Cross-Domain Recommendation

NeurIPS 2024poster

Cross-domain recommendation (CDR) offers a promising solution to the data sparsity problem by enabling knowledge transfer across source and target domains. However, many recent CDR models overlook crucial issues such as privacy as well as the risk of negative transfer (which negatively impact model…

Cited by 2SourcePDFScholar
2024

HINTED: Hard Instance Enhanced Detector with Mixed-Density Feature Fusion for Sparsely-Supervised 3D Object Detection

CVPR 2024poster

Current sparsely-supervised object detection methods largely depend on high threshold settings to derive high-quality pseudo labels from detector predictions. However hard instances within point clouds frequently display incomplete structures causing decreased confidence scores in their assigned pse…

2024

LiSA: LiDAR Localization with Semantic Awareness

CVPR 2024highlight

LiDAR localization is a fundamental task in robotics and computer vision which estimates the pose of a LiDAR point cloud within a global map. Scene Coordinate Regression (SCR) has demonstrated state-of-the-art performance in this task. In SCR a scene is represented as a neural network which outputs…

2024

Mining and Transferring Feature-Geometry Coherence for Unsupervised Point Cloud Registration

NeurIPS 2024poster

Point cloud registration, a fundamental task in 3D vision, has achieved remarkable success with learning-based methods in outdoor environments. Unsupervised outdoor point cloud registration methods have recently emerged to circumvent the need for costly pose annotations. However, they fail to estab…

2024

RELI11D: A Comprehensive Multimodal Human Motion Dataset and Method

CVPR 2024poster

Comprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB LiDAR or IMU data. However solely using these modalities or a combination of them may not be adequ…

Cited by 8SourcePDFScholar
2024

SPEAL: Skeletal Prior Embedded Attention Learning for Cross-Source Point Cloud Registration

AAAI 2024technical

Point cloud registration, a fundamental task in 3D computer vision, has remained largely unexplored in cross-source point clouds and unstructured scenes. The primary challenges arise from noise, outliers, and variations in scale and density. However, neglected geometric natures of point clouds restr…

2024

Sunshine to Rainstorm: Cross-Weather Knowledge Distillation for Robust 3D Object Detection

AAAI 2024technical

LiDAR-based 3D object detection models inevitably struggle under rainy conditions due to the degraded and noisy scanning signals. Previous research has attempted to address this by simulating the noise from rain to improve the robustness of detection models. However, significant disparities exist be…

Cited by 18SourcePDFScholar
2023

CIMI4D: A Large Multimodal Climbing Motion Dataset Under Human-Scene Interactions

CVPR 2023poster

Motion capture is a long-standing research problem. Although it has been studied for decades, the majority of research focus on ground-based movements such as walking, sitting, dancing, etc. Off-grounded actions such as climbing are largely overlooked. As an important type of action in sports and fi…

Cited by 30SourcePDFScholar
2023

CoIn: Contrastive Instance Feature Mining for Outdoor 3D Object Detection with Very Limited Annotations

ICCV 2023poster

Recently, 3D object detection with sparse annotations has received great attention. However, current detectors usually perform poorly under very limited annotations. To address this problem, we propose a novel Contrastive Instance feature mining method, named CoIn. To better identify indistinguishab…

Cited by 27PDFcodeScholar
2023

SGLoc: Scene Geometry Encoding for Outdoor LiDAR Localization

CVPR 2023poster

LiDAR-based absolute pose regression estimates the global pose through a deep network in an end-to-end manner, achieving impressive results in learning-based localization. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding the scene geom…

Cited by 16SourcePDFScholar
2023

SLOPER4D: A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments

CVPR 2023poster

We present SLOPER4D, a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild. Employing a head-mounted device integrated with a LiDAR and camera, we record 12 human subjects' activiti…

2023

Transformation-Equivariant 3D Object Detection for Autonomous Driving

AAAI 2023technical

3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augm…

2023

Virtual Sparse Convolution for Multimodal 3D Object Detection

CVPR 2023poster

Recently, virtual/pseudo-point-based 3D object detection that seamlessly fuses RGB images and LiDAR data by depth completion has gained great attention. However, virtual points generated from an image are very dense, introducing a huge amount of redundant computation during detection. Meanwhile, noi…

2022

HSC4D: Human-Centered 4D Scene Capture in Large-Scale Indoor-Outdoor Space Using Wearable IMUs and LiDAR

CVPR 2022poster

We propose Human-centered 4D Scene Capture (HSC4D) to accurately and efficiently create a dynamic digital world, containing large-scale indoor-outdoor scenes, diverse human motions, and rich interactions between humans and environments. Using only body-mounted IMUs and LiDAR, HSC4D is space-free wit…

Cited by 36PDFcodeScholar
2022

LiDARCap: Long-Range Marker-Less 3D Human Motion Capture With LiDAR Point Clouds

CVPR 2022poster

Existing motion capture datasets are largely short-range and cannot yet fit the need of long-range applications. We propose LiDARHuman26M, a new human motion capture dataset captured by LiDAR at a much longer range to overcome this limitation. Our dataset also includes the ground truth human motions…

Cited by 62PDFScholar
2021

Federated Learning with Fair Averaging

IJCAI 2021poster

Fairness has emerged as a critical problem in federated learning (FL). In this work, we identify a cause of unfairness in FL -- conflicting gradients with large differences in the magnitudes. To address this issue, we propose the federated fair averaging (FedFV) algorithm to mitigate potential confl…

2021

Tracklet Proposal Network for Multi-Object Tracking on Point Clouds

IJCAI 2021poster

This paper proposes the first tracklet proposal network, named PC-TCNN, for Multi-Object Tracking (MOT) on point clouds. Our pipeline first generates tracklet proposals, then refines these tracklets and associates them to generate long trajectories. Specifically, object proposal generation and moti…

Cited by 53SourcePDFScholar
2019

RF-Net: An End-To-End Image Matching Network Based on Receptive Field

CVPR 2019poster

This paper proposes a new end-to-end trainable matching network based on receptive field, RF-Net, to compute sparse correspondence between images. Building end-to-end trainable matching framework is desirable and challenging. The very recent approach, LF-Net, successfully embeds the entire feature e…

Cited by 128PDFScholar