← Search

Mu Xu

27 accepted papers

2026

AMap: Distilling Future Priors for Ahead-Aware Online HD Map Construction

CVPR 2026

Online High-Definition (HD) map construction is pivotal for autonomous driving. While recent approaches leverage historical temporal fusion to improve performance, we identify a critical safety flaw in this paradigm: it is inherently "spatially backward-looking." These methods predominantly enhance

Cited by 0SourceScholar
2026

AstraNav-Memory: Contexts Compression for Long Memory

CVPR 2026

Lifelong embodied navigation requires agents to accumulate, retain, and exploit spatial-semantic experience across tasks, enabling efficient exploration in novel environments and rapid goal reaching in familiar ones. While object-centric memory is interpretable, it depends on detection and reconstru

Cited by 0SourcecodeScholar
2026

CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local Navigation

ICLR 2026poster

Generalizing local navigation policies across diverse robot morphologies is a critical challenge. Progress is often hindered by the need for costly and embodiment-specific data, the tight coupling of planning and control, and the "disastrous averaging" problem where deterministic models fail to capt…

Cited by 0SourcecodeScholar
2026

CLoD-GS: Continuous Level-of-Detail via 3D Gaussian Splatting

ICLR 2026poster

Level of Detail (LoD) is a fundamental technique in real-time computer graphics for managing the rendering costs of complex scenes while preserving visual fidelity. Traditionally, LoD is implemented using discrete levels (DLoD), where multiple, distinct versions of a model are swapped out at differe…

Cited by 0SourcecodeScholar
2026

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

ICML 2026poster

End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptation…

Cited by 0SourceScholar
2026

EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence

CVPR 2026

Video-based spatial reasoning -- such as estimating distances, judging directions, or understanding layouts from multiple views -- requires selecting informative frames and, when needed, actively seeking additional viewpoints during inference. Existing multimodal large language models (MLLMs) consum

Cited by 0SourceScholar
2026

FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework

AAAI 2026technical

Human-Scene Interaction (HSI) seeks to generate realistic human behaviors within complex environments, yet it faces significant challenges in handling long-horizon, high-level tasks and generalizing to unseen scenes. To address these limitations, we introduce FantasyHSI, a novel HSI framework cente

Cited by 0SourcePDFScholar
2026

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

AAAI 2026technical

Recent advances in audio-driven portrait animation have demonstrated impressive capabilities. However, existing methods struggle to align with fine-grained human preferences across multiple dimensions, such as motion naturalness, lip-sync accuracy, and visual quality. This is due to the difficulty

Cited by 0SourcePDFScholar
2026

FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-and-Language Navigation

CVPR 2026

Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to understand textual instructions, perceive visual observations, and reason over long action sequences. Recent works, such as NavCoT and NavGPT-2, demonstrate the potential of Chain-of-Thought (CoT)

Cited by 0SourcecodeScholar
2026

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

ICLR 2026poster

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative priors, current video foundation models lack explicit 3D groundin…

Cited by 0SourcecodeScholar
2026

JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation

ICLR 2026poster

Vision-and-Language Navigation (VLN) requires an embodied agent to navigate through unseen environments, guided by natural language instructions and a continuous video stream. Recent advances in VLN have been driven by the powerful semantic understanding of Multimodal Large Language Models (MLLMs).…

Cited by 0SourcecodeScholar
2026

MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving

CVPR 2026

Vision-Language Models (VLM) exhibit strong reasoning capabilities, showing promise for end-to-end autonomous driving systems. Chain-of-Thought (CoT), as VLM's widely used reasoning strategy, is facing critical challenges. Existing textual CoT has a large gap between text semantic space and trajecto

Cited by 0SourcecodeScholar
2026

NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction

CVPR 2026

Embodied navigation for long-horizon tasks, guided by complex natural language instructions, remains a formidable challenge in artificial intelligence. Existing agents often struggle with robust long-term planning about unseen environments, leading to high failure rates. To address these limitations

Cited by 0SourceScholar
2026

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

ICML 2026poster

Despite the rapid progress of Vision-Language-Action (VLA) models, the prevailing paradigm of predicting discrete waypoints remains fundamentally misaligned with the intrinsic continuity of physical motion. This discretization imposes rigid sampling rates, lacks high-order differentiability, and int…

Cited by 0SourceScholar
2026

OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation

ICLR 2026poster

Embodied navigation is a foundational challenge for intelligent robots, demanding the ability to comprehend visual environments, follow natural language instructions, and explore autonomously. However, existing models struggle to provide a unified solution across heterogeneous navigation paradigms,…

Cited by 0SourceScholar
2026

Online Navigation Refinement: Achieving Lane-Level Guidance by Associating Standard-Definition and Online Perception Maps

ICLR 2026poster

Lane-level navigation is critical for geographic information systems and navigation-based tasks, offering finer-grained guidance than road-level navigation by standard definition (SD) maps. However, it currently relies on expansive global HD maps that cannot adapt to dynamic road conditions. Recentl…

Cited by 0SourcecodeScholar
2026

Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving

AAAI 2026technical

Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture th

Cited by 0SourcePDFScholar
2026

PriorDrive: Enhancing Online HD Mapping with Unified Vector Priors

AAAI 2026technical

High-Definition Maps (HD maps) are essential for the precise navigation and decision-making of autonomous vehicles, yet their creation and upkeep present significant cost and timeliness challenges. The online construction of HD maps using on-board sensors has emerged as a promising solution; however

Cited by 0SourcePDFScholar
2026

RehearseVLA: Simulated Post-Training for VLAs with Physically-Consistent World Model

CVPR 2026

Vision-Language-Action (VLA) models trained via imitation learning suffer from significant performance degradation in data-scarce scenarios due to their reliance on large-scale demonstration datasets. Although reinforcement learning (RL)-based post-training has proven effective in addressing data sc

Cited by 0SourcecodeScholar
2026

Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models

ICRA 2026poster

Latent Action Models (LAMs) enable Vision-Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottlenecks of LAMs: 1) the commonly adopted end-to-end trained image encoder suffers from poor spatial understanding; 2) LAMs can …

2026

SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation

CVPR 2026

Embodied navigation that adheres to social norms remains an open research challenge. Our SocialNav is a foundational model for socially-aware navigation with a hierarchical "brain-action" architecture, capable of understanding high-level social norms and generating low-level, socially compliant traj

Cited by 0SourcecodeScholar
2026

UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data

AAAI 2026technical

Large-scale map construction is foundational for critical applications such as autonomous driving and navigation systems. Traditional large-scale map construction approaches mainly rely on costly and inefficient special data collection vehicles and labor-intensive annotation processes. While existin

Cited by 0SourcePDFScholar
2025

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

NeurIPS 2025spotlight

Vision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability. Most prior work, however, inserts textual chains-of-thought (CoT) as intermediate steps tailored to the current scene. Such symbolic compressions can blur spatio-tempo…

Cited by 0SourcecodeScholar
2025

G3PT: Unleash the Power of Autoregressive Modeling in 3D Generation via Cross-Scale Querying Transformer

IJCAI 2025

Autoregressive transformers have revolutionized generative models in language processing and shown substantial promise in image and video generation. However, these models face significant challenges when extended to 3D generation tasks due to their reliance on next-token prediction to learn token s

Cited by 0SourcePDFScholar
2025

HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset

CVPR 2025highlight

With the rapid evolution of 3D generation algorithms, the cost of producing 3D humanoid character models has plummeted, yet the field is impeded by the lack of a comprehensive dataset for automatic rigging--a pivotal step in character animation. Addressing this gap, we present HumanRig, the first la…

2025

SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions

ICCV 2025poster

Accurate lane topology is essential for autonomous driving, yet traditional methods struggle to model the complex, non-linear structures--such as loops and bidirectional lanes--prevalent in real-world road structure. We present SeqGrowGraph, a novel framework that learns lane topology as a chain of…

Cited by 0SourcePDFScholar
2015

Effect of vibrotactile cues for guiding simultaneous procedural motion of two joints on upper limbs

IROS 2015poster

Simultaneous motion control of multiple joints has many potential applications such as Tai Chi, Yoga etc. The capability of vibrotactile cues to assist this kind of motor task has not been well explored. In this paper, we studied the effect of vibrotactile cues for guiding procedural motion of two j…

Cited by 5SourceScholar