← Search

Xudong Cai

7 accepted papers

2026

Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction

AAAI 2026technical

Reconstructing dense geometry for dynamic scenes from a monocular video is a critical yet challenging task. Recent memory-based methods enable efficient online reconstruction, but they fundamentally suffer from a Memory Demand Dilemma: The memory representation faces an inherent conflict be

Cited by 0SourcePDFScholar
2026

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

AAAI 2026technical

Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results w

Cited by 0SourcePDFScholar
2026

VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models

ICLR 2026poster

Large reasoning models such as OpenAI o1 and DeepSeek-R1 have demonstrated remarkable performance in complex reasoning tasks. A critical component of their training is the incorporation of reference-based reward systems within reinforcement learning (RL), where model outputs are evaluated against gr…

Cited by 0SourcecodeScholar
2025

Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation

NeurIPS 2025poster

Vision-Language Navigation is a critical task for developing embodied agents that can follow natural language instructions to navigate in complex real-world environments. Recent advances by finetuning large pretrained models have significantly improved generalization and instruction grounding compa…

Cited by 0SourceScholar
2025

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing

CVPR 2025poster

Deep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization,…

Cited by 0SourcePDFScholar
2024

VOLoc: Visual Place Recognition by Querying Compressed Lidar Map

ICRA 2024poster

The availability of city-scale Lidar maps enables the potential of city-scale place recognition using mobile cameras. However, the city-scale Lidar maps generally need to be compressed for storage efficiency, which increases the difficulty of direct visual place recognition in compressed Lidar maps.…

Cited by 7SourcecodeScholar
2023

ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding

ICRA 2023poster

Recently, a growing number of work design unsupervised paradigms for point cloud processing to alleviate the limitation of expensive manual annotation and poor transferability of supervised methods. Among them, CrossPoint follows the contrastive learning framework and exploits image and point cloud…

Cited by 11SourcecodeScholar