← Search

Yuchen Wu

28 accepted papers

2026

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

CVPR 2026

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cros

Cited by 0SourcecodeScholar
2026

CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual Localization

CVPR 2026

Scene Coordinate Regression (SCR) has emerged as a memory-efficient paradigm for visual localization. While SCR has demonstrated performance comparable to classic feature matching based approaches in small-scale scenes, it has consistently underperformed in large-scale environments. Large-scale loca

Cited by 0SourceScholar
2026

FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM

AAAI 2026technical

We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance f

Cited by 0SourcePDFScholar
2026

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning

ICML 2026spotlight

Solving complex geometric problems inherently requires \textit{interleaved reasoning}: a tight alternation between constructing diagrams and performing logical deductions. Although recent Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities in visual generation and plotting…

Cited by 0SourceScholar
2026

PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

CVPR 2026

This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward solution is to employ a pseudo-labeling strateg

Cited by 0SourceScholar
2026

PressTrack-HMR: Pressure-Based Top-Down Multi-Person Global Human Mesh Recovery

AAAI 2026technical

Multi-person global human mesh recovery (HMR) is crucial for understanding crowd dynamics and interactions. Traditional vision-based HMR methods sometimes face limitations in real-world scenarios due to mutual occlusions, insufficient lighting, and privacy concerns. Human-floor tactile interactions

Cited by 0SourcePDFScholar
2026

SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings

CVPR 2026

Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual divergence of estimated scale over long sequences. Existing frame-to-frame methods achieve real-time performance through lo

Cited by 0SourceScholar
2025

AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models

NeurIPS 2025poster

Effective human-agent collaboration in physical environments requires understanding not only what to act upon, but also where the actionable elements are and how to interact with them. Existing approaches often operate at the object level or disjointedly handle fine-grained affordance reasoning, lac…

Cited by 0SourceScholar
2025

Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs

ACL 2025finding

Knowledge editing allows for efficient adaptation of large language models (LLMs) to new information or corrections without requiring full retraining. However, prior methods typically focus on either single-language editing or basic multilingual editing, failing to achieve true cross-linguistic know…

2025

Efficient Multi-Robot Task and Path Planning in Large-Scale Cluttered Environments

RA-L 2025

As the potential of multi-robot systems continues to be explored and validated across various real-world applications, such as package delivery, search and rescue, and autonomous exploration, the need to improve the efficiency and quality of task and path planning has become increasingly urgent, par

Cited by 4SourceScholar
2025

MARF: Cooperative Multi-Agent Path Finding with Reinforcement Learning and Frenet Lattice in Dynamic Environments

ICRA 2025

Multi-agent path finding (MAPF) in dynamic and complex environments is a highly challenging task. Recent research has focused on the scalability of agent numbers or the complexity of the environment. Usually, they disregard the agents' physical constraints or use a differential-driven model. However

Cited by 1SourceScholar
2025

Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach

NeurIPS 2025poster

Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely. Existing safety evaluations primarily rely on context-independent metrics—such as fact…

Cited by 0SourcecodeScholar
2025

Revisiting Continual Ultra-fine-grained Visual Recognition with Pre-trained Models

IJCAI 2025

Continual ultra-fine-grained visual recognition (C-UFG) aims to continuously learn to categorize the increasing number of cultivates (VC-UFG) and consistently recognize crops across reproductive stages (HC-UFG), which is a fundamental goal of intelligent agriculture. Despite the progress made in gen

2025

Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA

EMNLP 2025

Large language models (LLMs) encode vast amounts of world knowledge but remain static once trained, making timely integration of emerging facts prohibitively expensive via full retraining. Knowledge-editing techniques have thus emerged to inject or overwrite specific facts into LLMs, yet they either

2024

A Distributed Pipeline for Collaborative Pursuit in the Target Guarding Problem

RA-L 2024

The target guarding problem (TGP) is a classical combat game where pursuers aim to capture evaders to protect a territory from intrusion. This paper proposes a distributed pipeline for multi-pursuer multi-evader TGP with the capability to accommodate varying numbers of evaders and criteria for succe

Cited by 8SourceScholar
2024

Failures and Successes of Cross-Validation for Early-Stopped Gradient Descent

AISTATS 2024poster

We analyze the statistical properties of generalized cross-validation (GCV) and leave-one-out cross-validation (LOOCV) applied to early-stopped gradient descent (GD) in high-dimensional least squares regression. We prove that GCV is generically inconsistent as an estimator of the prediction risk of…

Cited by 5SourcePDFScholar
2024

Hierarchical Search-Based Cooperative Motion Planning

IROS 2024poster

Cooperative path planning, a crucial aspect of multi-agent systems research, serves a variety of sectors, including military, agriculture, and industry. Many existing algorithms, however, come with certain limitations, such as simplified kinematic models and inadequate support for multiple group sce…

Cited by 0SourcecodeScholar
2024

Optimizing Multi-Touch Textile and Tactile Skin Sensing Through Circuit Parameter Estimation

ICRA 2024poster

Tactile and textile skin technologies have become increasingly important for enhancing human-robot interaction and allowing robots to adapt to different environments. Despite notable advancements, there are ongoing challenges in skin signal processing, particularly in achieving both accuracy and spe…

Cited by 1SourceScholar
2024

PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizes

ICASSP 2024accepted

3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target datas…

Cited by 0SourceScholar
2024

Theoretical insights for diffusion guidance: A case study for Gaussian mixture models

ICML 2024poster

Diffusion models benefit from instillation of task-specific information into the score function to steer the sample generation towards desired properties. Such information is coined as guidance. For example, in text-to-image synthesis, text input is encoded as guidance to generate semantically align…

Cited by 29SourcePDFScholar
2023

ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

NeurIPS 2023poster

We present a comprehensive solution to learn and improve text-to-image models from human preference feedback. To begin with, we build ImageReward---the first general-purpose text-to-image human preference reward model---to effectively encode human preferences. Its training is based on our systematic…

2023

PateGail: A Privacy-Preserving Mobility Trajectory Generator with Imitation Learning

AAAI 2023technical

Generating human mobility trajectories is of great importance to solve the lack of large-scale trajectory data in numerous applications, which is caused by privacy concerns. However, existing mobility trajectory generation methods still require real-world human trajectories centrally collected as th…

2023

Picking up Speed: Continuous-Time Lidar-Only Odometry Using Doppler Velocity Measurements

RA-L 2023

Frequency-Modulated Continuous-Wave (FMCW) lidar is a recently emerging technology that additionally enables per-return instantaneous relative radial velocity measurements via the Doppler effect. In this letter, we present the first continuous-time lidar-only odometry algorithm using these Doppler v

Cited by 40SourcecodeScholar
2023

Toward Closed-Loop Additive Manufacturing: Paradigm Shift in Fabrication, Inspection, and Repair

IROS 2023poster

Increased usage of additive manufacturing (AM) in various industries has solidified its role as an advanced manufacturing technique. However, there is an inherent lack of reliability in AM processes, particularly common in extrusion or deposition-based methods due to the stochastic nature of ma-teri…

Cited by 4SourceScholar
2022

Are We Ready for Radar to Replace Lidar in All-Weather Mapping and Localization?

RA-L 2022

We present an extensive comparison between three topometric localization systems: radar-only, lidar-only, and a cross-modal radar-to-lidar system across varying seasonal and weather conditions using the Boreas dataset. Contrary to our expectations, our experiments showed that our lidar-only pipeline

Cited by 84SourcecodeScholar
2021

Shaping Rewards for Reinforcement Learning with Imperfect Demonstrations using Generative Models

ICRA 2021poster

The potential benefits of model-free reinforcement learning to real robotics systems are limited by its uninformed exploration that leads to slow convergence, lack of data-efficiency, and unnecessary interactions with the environment. To address these drawbacks we propose a method that combines rein…

Cited by 35SourceScholar
2021

Streaming Belief Propagation for Community Detection

NeurIPS 2021poster

The community detection problem requires to cluster the nodes of a network into a small number of well-connected ‘communities’. There has been substantial recent progress in characterizing the fundamental statistical limits of community detection under simple stochastic block models. However, in re…

Cited by 6SourcePDFScholar