← Search

Wenshan Wang

36 accepted papers

2026

AnyThermal: Towards Learning Universal Representations for Thermal Perception

ICRA 2026poster

We present AnyThermal, a thermal backbone that captures robust task-agnostic thermal features suitable for a variety of tasks such as cross-modal place recognition, thermal segmentation, and monocular depth estimation using thermal images. Existing thermal backbones that follow task-specific trainin…

2026

RAVEN: Resilient Aerial Navigation Via Open-Set Semantic Memory and Behavior Adaptation

ICRA 2026poster

Aerial outdoor semantic navigation requires robots to explore large, unstructured environments to locate target objects. Recent advances in semantic navigation have demonstrated open-set object-goal navigation in indoor settings, but these methods remain limited by constrained spatial ranges and str…

2026

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

ICRA 2026poster

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs to navigation presents two key challenges: effectively parsing and structuring complex environment information and determ…

2026

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

RSS 2026poster

Robotic navigation in human environments requires a spatio-temporal semantic representation that can reconcile open-vocabulary perception with long-term environmental changes. While foundation models provide strong zero-shot recognition, their predictions are intermittent and view-dependent, and nai…

Cited by 0SourceScholar
2026

TravSUITE: Traversability via Self-Supervised, Uncertainty-Aware IRL and Terrain Estimation

RSS 2026poster

Traversability analysis in off-road settings remains a fundamental challenge for mobile robots. Key difficulties include constructing an accurate, expressive local map from multi-modal sensor data and using the map to design traversability rules that yield desirable navigation behavior. Importantly,…

Cited by 0SourceScholar
2025

IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes

ICRA 2025

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can operate in diverse environments given natural language input. One such application is indoor navigation using natural la

Cited by 3SourcecodeScholar
2025

MAC-VO: Metrics-Aware Covariance for Learning-Based Stereo Visual Odometry mac-vo.github.io

ICRA 2025

We propose MAC-VO, a novel learning-based stereo visual odometry (VO) framework that trains a metrics-aware uncertainty model to serve two critical functions: selecting keypoints and weighting residuals in pose graph optimization. Unlike traditional geometric methods that favor texture-rich features

Cited by 12SourceScholar
2025

RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration

IROS 2025

Open-set semantic mapping is crucial for openworld robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settings, where overall they fail to combine within-range and beyond-range observations. Furthermore, these methods make a trad

Cited by 22SourceScholar
2025

SALON: Self-supervised Adaptive Learning for Off-road Navigation

ICRA 2025

Autonomous robot navigation in off-road environments presents a number of challenges due to its lack of structure, making it difficult to handcraft robust heuristics for diverse scenarios. While learned methods using hand labels or self-supervised data improve generalizability, they often require a

Cited by 10SourceScholar
2025

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

IROS 2025

Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is often challenging due to the diversity of scenes, large number of fine-grained objects, and complex free-form nature of

Cited by 8SourcecodeScholar
2025

Tartan IMU: A Light Foundation Model for Inertial Positioning in Robotics

CVPR 2025poster

Despite recent advances in deep learning, most existing learning IMU odometry methods are trained on specific datasets, lack generalization, and are prone to overfitting, which limits their real-world application. To address these challenges, we present Tartan IMU, a foundation model designed for ge…

Cited by 0SourcePDFScholar
2025

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

IROS 2025

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environments includes multiple RGB stereo cameras for 360-degree coverage, along with de

Cited by 17SourceScholar
2025

UFM: A Simple Path towards Unified Dense Correspondence with Flow

NeurIPS 2025poster

Dense image correspondence is central to many applications, such as visual odometry, 3D reconstruction, object association, and re-identification. Historically, dense correspondence has been tackled separately for wide-baseline scenarios and optical flow estimation, despite the common goal of matchi…

Cited by 0SourceScholar
2024

BEVRender: Vision-based Cross-view Vehicle Registration in Off-road GNSS-denied Environment

IROS 2024poster

We introduce BEVRender, a novel learning-based approach for the localization of ground vehicles in Global Navigation Satellite System (GNSS)-denied off-road scenarios. These environments are typically challenging for conventional vision-based state estimation due to the lack of distinct visual landm…

Cited by 1SourceScholar
2024

Interactive-FAR:Interactive, Fast and Adaptable Routing for Navigation Among Movable Obstacles in Complex Unknown Environments

IROS 2024poster

This paper introduces a real-time algorithm for navigating complex unknown environments cluttered with movable obstacles. Our algorithm achieves fast, adaptable routing by actively attempting to manipulate obstacles during path planning and adjusting the global plan from sensor feedback. The main co…

Cited by 3SourceScholar
2024

LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

NeurIPS 2024poster

Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, t…

2024

TartanDrive 2.0: More Modalities and Better Infrastructure to Further Self-Supervised Learning Research in Off-Road Driving Tasks

ICRA 2024poster

We present TartanDrive 2.0, a large-scale off-road driving dataset for self-supervised learning tasks. In 2021 we released TartanDrive 1.0, which is one of the largest datasets for off-road terrain. As a follow-up to our original dataset, we collected seven hours of data at speeds of up to 15m/s wit…

Cited by 21SourceScholar
2024

Velociraptor: Leveraging Visual Foundation Models for Label-Free, Risk-Aware Off-Road Navigation

CoRL 2024poster

Traversability analysis in off-road regimes is a challenging task that requires understanding of multi-modal inputs such as camera and LiDAR. These measurements are often sparse, noisy, and difficult to interpret, particularly in the off-road setting. Existing systems are very engineering-intensive,…

Cited by 2SourceScholar
2023

DytanVO: Joint Refinement of Visual Odometry and Motion Segmentation in Dynamic Environments

ICRA 2023poster

Learning-based visual odometry (VO) algorithms achieve remarkable performance on common static scenes, benefiting from high-capacity models and massive annotated data, but tend to fail in dynamic, populated environments. Semantic segmentation is largely used to discard dynamic associations before es…

Cited by 55SourcecodeScholar
2023

How Does It Feel? Self-Supervised Costmap Learning for Off-Road Vehicle Traversability

ICRA 2023poster

Estimating terrain traversability in off-road environments requires reasoning about complex interaction dynamics between the robot and these terrains. However, it is challenging to create informative labels to learn a model in a supervised manner for these interactions. We propose a method that lear…

Cited by 72SourcecodeScholar
2023

Learning Risk-Aware Costmaps via Inverse Reinforcement Learning for Off-Road Navigation

ICRA 2023poster

The process of designing costmaps for off-road driving tasks is often a challenging and engineering-intensive task. Recent work in costmap design for off-road driving focuses on training deep neural networks to predict costmaps from sensory observations using corpora of expert driving data. However,…

Cited by 31SourceScholar
2023

PyPose: A Library for Robot Learning With Physics-Based Optimization

CVPR 2023poster

Deep learning has had remarkable success in robotic perception, but its data-centric nature suffers when it comes to generalizing to ever-changing environments. By contrast, physics-based optimization generalizes better, but it does not perform as well in complicated tasks due to the lack of high-le…

2022

AirDOS: Dynamic SLAM benefits from Articulated Objects

ICRA 2022poster

Dynamic Object-aware SLAM (DOS) exploits object-level information to enable robust motion estimation in dynamic environments. Existing methods mainly focus on identifying and excluding dynamic objects from the optimization. In this paper, we show that feature-based visual SLAM systems can also benef…

Cited by 64SourcecodeScholar
2022

COMPASS: Contrastive Multimodal Pretraining for Autonomous Systems

IROS 2022poster

Learning representations that generalize across tasks and domains is challenging yet necessary for autonomous systems. Although task-driven approaches are appealing, de-signing models specific to each application can be difficult in the face of limited data, especially when dealing with highly varia…

Cited by 10SourcecodeScholar
2022

TartanDrive: A Large-Scale Dataset for Learning Off-Road Dynamics Models

ICRA 2022poster

We present TartanDrive, a large scale dataset for learning dynamics models for off-road driving. We collected a dataset of roughly 200,000 off-road driving interactions on a modified Yamaha Viking ATV with seven unique sensing modalities in diverse terrains. To the authors' knowledge, this is the la…

Cited by 62SourcecodeScholar
2021

Improving Off-road Planning Techniques with Learned Costs from Physical Interactions

ICRA 2021poster

Autonomous ground vehicles have improved greatly over the past decades, but they still have their limitations when it comes to off-road environments. There is still a need for planning techniques that effectively handle physical interactions between a vehicle and its surroundings. We present a metho…

Cited by 20SourceScholar
2021

ORStereo: Occlusion-Aware Recurrent Stereo Matching for 4K-Resolution Images

IROS 2021poster

Stereo reconstruction models trained on small images do not generalize well to high-resolution data. Training a model on high-resolution image size faces difficulties of data availability and is often infeasible due to limited computing resources. In this work, we present the Occlusion-aware Recurre…

Cited by 11SourceScholar
2021

Rough Terrain Navigation Using Divergence Constrained Model-Based Reinforcement Learning

CoRL 2021poster

Autonomous navigation of wheeled robots in rough terrain environments has been a long standing challenge. In these environments, predicting the robot's trajectory can be challenging due to the complexity of terrain interactions, as well as the divergent dynamics that cause model uncertainty to compo…

Cited by 17SourceScholar
2020

TartanAir: A Dataset to Push the Limits of Visual SLAM

IROS 2020poster

We present a challenging dataset, the TartanAir, for robot navigation tasks and more. The data is collected in photo-realistic simulation environments with the presence of moving objects, changing light and various weather conditions. By collecting data in simulations, we are able to obtain multi-mo…

Cited by 406SourcecodeScholar
2020

Visual Memorability for Robotic Interestingness via Unsupervised Online Learning

ECCV 2020poster

In this paper, we explore the problem of interesting scene prediction for mobile robots. This area is currently underexplored but is crucial for many practical applications such as autonomous exploration and decision making. Inspired by industrial demands, we first propose a novel translation-invari…

2019

Can a Robot Become a Movie Director? Learning Artistic Principles for Aerial Cinematography

IROS 2019poster

Aerial filming is constantly gaining importance due to the recent advances in drone technology. It invites many intriguing, unsolved problems at the intersection of aesthetical and scientific challenges. In this work, we propose a deep reinforcement learning agent which supervises motion planning of…

Cited by 72SourceScholar
2019

Improved Generalization of Heading Direction Estimation for Aerial Filming Using Semi-Supervised Regression

ICRA 2019poster

In the task of Autonomous aerial filming of a moving actor (e.g. a person or a vehicle), it is crucial to have a good heading direction estimation for the actor from the visual input. However, the models obtained in other similar tasks, such as pedestrian collision risk analysis and human-robot inte…

Cited by 8SourceScholar
2019

Improving Learning-based Ego-motion Estimation with Homomorphism-based Losses and Drift Correction

IROS 2019poster

Visual odometry is an essential problem for mobile robots. Traditional methods for solving VO mostly utilize geometric optimization. While capable of achieving high accuracy, these methods require accurate sensor calibration and complicated parameter tuning to work well in practice. With the rise of…

Cited by 16SourceScholar
2019

Towards a Robust Aerial Cinematography Platform: Localizing and Tracking Moving Targets in Unstructured Environments

IROS 2019poster

The use of drones for aerial cinematography has revolutionized several applications and industries that require live and dynamic camera viewpoints such as entertainment, sports, and security. However, safely controlling a drone while filming a moving target usually requires multiple expert human ope…

Cited by 108SourceScholar
2018

Integrating kinematics and environment context into deep inverse reinforcement learning for predicting off-road vehicle trajectories

CoRL 2018

Predicting the motion of a mobile agent from a third-person perspective is an important component for many robotics applications, such as autonomous navigation and tracking. With accurate motion prediction of other agents, robots can plan for more intelligent behaviors to achieve specified objective