← Search

Pengyu Yin

13 accepted papers

2026

Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment

ICML 2026poster

Physics-informed diffusion models typically impose PDE constraints only on the final output, leaving intermediate features unconstrained. This can enable shortcut solutions that fit training statistics yet generalize poorly under shifted boundary conditions. We introduce \textbf{REPA-P}, a \emph{tea…

Cited by 0SourceScholar
2025

Large-Scale UWB Anchor Calibration and One-Shot Localization Using Gaussian Process

ICRA 2025

Ultra-wideband (UWB) is gaining popularity with devices like AirTags for precise home item localization but faces significant challenges when scaled to large environments like seaports. The main challenges are calibration and localization under obstructed conditions, which are common in logistics en

Cited by 15SourceScholar
2024

An Image Acquisition Scheme for Visual Odometry based on Image Bracketing and Online Attribute Control

ICRA 2024poster

Visual odometry (VO) system is challenged by complex illumination environments. Image quality and its consistency in the time domain directly determine feature detection and tracking performance, which further affect the robustness and accuracy of the entire system. In this paper, an image acquisiti…

Cited by 2SourceScholar
2024

LIO-GVM: An Accurate, Tightly-Coupled Lidar-Inertial Odometry With Gaussian Voxel Map

RA-L 2024

This letter presents a probabilistic voxel-based LiDAR Inertial Odometry framework for accurate and robust pose estimation. The framework addresses the correspondence mismatching issue by representing the LiDAR points as a set of Gaussian distributions and evaluating the divergence in variance for o

Cited by 21SourcecodeScholar
2024

MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception

CVPR 2024highlight

Perception plays a crucial role in various robot applications. However existing well-annotated datasets are biased towards autonomous driving scenarios while unlabelled SLAM datasets are quickly over-fitted and often lack environment and domain variations. To expand the frontier of these fields we i…

Cited by 36SourcePDFScholar
2024

MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation

ICRA 2024poster

Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-i…

Cited by 19SourcecodeScholar
2024

Multi-Robot Active Graph Exploration with Reduced Pose-SLAM Uncertainty via Submodular Optimization

IROS 2024poster

This paper considers the multi-robot active graph exploration problem, where robots need to collaboratively cover a graph environment while maintaining reliable pose estimation in collaborative Simultaneous Localization and Mapping (SLAM). Considering both objectives presents challenges for multi-ro…

Cited by 2SourcecodeScholar
2024

Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning

ICRA 2024poster

One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been extensively studied as a global descriptor retrieval (i.e., loop…

Cited by 19SourcecodeScholar
2024

Reliable Spatial-Temporal Voxels For Multi-Modal Test-Time Adaptation

ECCV 2024poster

"Multi-modal test-time adaptation (MM-TTA) is proposed to adapt models to an unlabeled target domain by leveraging the complementary multi-modal inputs in an online manner. Previous MM-TTA methods for 3D segmentation rely on predictions of cross-modal information in each input frame, while they igno…

2024

SGBA: Semantic Gaussian Mixture Model-Based LiDAR Bundle Adjustment

RA-L 2024

LiDAR bundle adjustment (BA) is an effective approach to reduce the drifts in pose estimation from the front-end. Existing works on LiDAR BA usually rely on predefined geometric features for landmark representation. This reliance restricts generalizability, as the system will inevitably deteriorate

Cited by 8SourceScholar
2023

Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation

ICCV 2023poster

Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmenta…

Cited by 20PDFScholar
2023

Segregator: Global Point Cloud Registration with Semantic and Geometric Cues

ICRA 2023poster

This paper presents Segregator, a global point cloud registration framework that exploits both semantic information and geometric distribution to efficiently build up outlier-robust correspondences and search for inliers. Current state-of-the-art algorithms rely on point features to set up putative…

Cited by 28SourcecodeScholar
2020

CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence

IROS 2020poster

In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective functi…

Cited by 15SourcecodeScholar