← Search

Yifan Bai

12 accepted papers

2026

An Adaptive Inspection Planning Approach towards Routine Monitoring in Uncertain Environments

ICRA 2026poster

In this work, we present a hierarchical framework designed to support robotic inspection under environment uncertainty. By leveraging a known environment model, existing methods plan and safely track inspection routes to visit points of interest. However, discrepancies between the model and actual s…

2026

Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees

ICML 2026poster

Statement autoformalization acts as a critical bridge between human mathematics and formal mathematics by translating natural language problems into formal language. While prior works have focused on data synthesis and diverse training paradigms to optimize end-to-end Large Language Models (LLMs), t…

Cited by 0SourceScholar
2026

FARTrack: Fast Autoregressive Visual Tracking with High Performance

ICLR 2026poster

Inference speed and tracking performance are two critical evaluation metrics in the field of visual tracking. However, high-performance trackers often suffer from slow processing speeds, making them impractical for deployment on resource-constrained devices. To alleviate this issue, we propose $\tex…

Cited by 0SourcecodeScholar
2026

Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving

AAAI 2026technical

Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture th

Cited by 0SourcePDFScholar
2026

ReMoT: Reinforcement Learning with Motion Contrast Triplets

CVPR 2026

We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency--a critical failure point in navigation, robotics, and autonomous driving. ReMoT integrates two core components: (i) A rule-based automatic framework that gener

Cited by 0SourceScholar
2026

TACOcc: Target-Adaptive Cross-Modal Fusion with Sequential Volume Rendering for 3D Semantic Occupancy Prediction

ICRA 2026poster

Multi-modal 3D semantic occupancy prediction remains challenged by two fundamental issues: (i) geometric--semantic misalignment introduced by fixed-neighborhood fusion under heterogeneous sensing distributions, and (ii) feature degradation with prediction inconsistency in dynamic scenes caused by sp…

Cited by 0Scholar
2025

Collaborative Task Assignment, Sequencing and Multi-agent Path-finding

IROS 2025

In this article, we address the problem of collaborative task assignment, sequencing, and multi-agent pathfinding (TSPF), where a team of agents must visit a set of task locations without collisions while minimizing flowtime. TSPF incorporates agent-task compatibility constraints and ensures that al

Cited by 0SourceScholar
2025

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

NeurIPS 2025spotlight

Vision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability. Most prior work, however, inserts textual chains-of-thought (CoT) as intermediate steps tailored to the current scene. Such symbolic compressions can blur spatio-tempo…

Cited by 0SourcecodeScholar
2025

Multi-Agent Path Finding Using Conflict-Based Search and Structural-Semantic Topometric Maps

ICRA 2025

As industries increasingly adopt large robotic fleets, there is a pressing need for computationally efficient, practical, and optimal conflict-free path planning for multiple robots. Conflict-Based Search (CBS) is a popular method for multi-agent path finding (MAPF) due to its completeness and optim

Cited by 2SourceScholar
2024

Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation

ECCV 2024oral

"This paper introduces the point-axis representation for oriented object detection, as depicted in aerial images in Figure ??, emphasizing its flexibility and geometrically intuitive nature with two key components: points and axes. 1) Points delineate the spatial extent and contours of objects, prov…

Cited by 5SourcePDFScholar