← Search

Deying Li

18 accepted papers

2026

EIMC: Efficient Instance-Aware Multi-Modal Collaborative Perception

ICRA 2026poster

Multi-modal collaborative perception calls for great attention to enhancing the safety of autonomous driving. However, current multi-modal approaches remain a ``local fusion to communication” sequence, which fuses multi-modal data locally and needs high bandwidth to transmit an individual's feature …

2026

MapDream: Task-Driven Map Learning for Vision-Language Navigation

ICML 2026poster

Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception. However, most existing approaches rely on hand-crafted maps constructed independently…

Cited by 0SourceScholar
2026

Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction

AAAI 2026technical

Reconstructing dense geometry for dynamic scenes from a monocular video is a critical yet challenging task. Recent memory-based methods enable efficient online reconstruction, but they fundamentally suffer from a Memory Demand Dilemma: The memory representation faces an inherent conflict be

Cited by 0SourcePDFScholar
2026

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

AAAI 2026technical

Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results w

Cited by 0SourcePDFScholar
2026

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

CVPR 2026

Vision-Language Navigation requires agents to act coherently over long horizons by understanding not only local visual context but also how far they have advanced within a multi-step instruction.However, recent Vision-Language-Action models focus on direct action prediction and earlier progress meth

Cited by 0SourceScholar
2026

SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything

CVPR 2026

Online 3D instance segmentation is a critical capability for embodied agents navigating in dynamic environments. However, a fundamental challenge remains in adapting powerful 2D foundation models, like SAM, to 3D online segmentation. Naively lifting SAM's 2D masks to 3D results in severe spatial fra

Cited by 0SourceScholar
2026

SupGS-SLAM: Gaussian Splatting SLAM with Efficient Keyframe Strategy and Supplementary Mapping

ICRA 2026poster

Gaussian Splatting SLAM methods have exhibited impressive high-fidelity rendering performance. Existing methods maintain high rendering quality around the current camera viewpoint, but the rendering quality degrades in previously observed regions as the camera moves away, particularly in real-world …

Cited by 0codeScholar
2025

Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation

NeurIPS 2025poster

Vision-Language Navigation is a critical task for developing embodied agents that can follow natural language instructions to navigate in complex real-world environments. Recent advances by finetuning large pretrained models have significantly improved generalization and instruction grounding compa…

Cited by 0SourceScholar
2025

Is Discretization Fusion All You Need for Collaborative Perception?

ICRA 2025

Collaborative perception in multi-agent system enhances overall perceptual capabilities by facilitating the exchange of complementary information among agents. Current mainstream collaborative perception methods rely on discretized feature maps to conduct fusion, which however, lacks flexibility in

Cited by 3SourcecodeScholar
2025

LA-MOTR: End-to-End Multi-Object Tracking by Learnable Association

ICCV 2025poster

This paper proposes LA-MOTR, a novel Tracking-by-Learnable-Association framework that resolves the competing optimization objectives between detection and association in end-to-end Tracking-by-Attention (TbA) Multi-Object Tracking. Current TbA methods employ shared decoders for simultaneous object d…

2025

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing

CVPR 2025poster

Deep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization,…

Cited by 0SourcePDFScholar
2025

Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis

CVPR 2025poster

This paper proposes a general solution to enable point cloud recognition models to handle distribution shifts at test time. Unlike prior methods, which rely heavily on training data (often inaccessible during online inference) and are limited to recognizing a fixed set of point cloud classes predefi…

2025

STAR: Spatial-Temporal Tracklet Matching for Multi-Object Tracking

NeurIPS 2025poster

Existing tracking-by-detection Multi-Object Tracking methods mainly rely on associating objects with tracklets using motion and appearance features. However, variations in viewpoint and occlusions can result in discrepancies between the features of current objects and those of historical tracklets.…

Cited by 0SourceScholar
2024

DroneMOT: Drone-based Multi-Object Tracking Considering Detection Difficulties and Simultaneous Moving of Drones and Objects

ICRA 2024poster

Multi-object tracking (MOT) on static platforms, such as by surveillance cameras, has achieved significant progress, with various paradigms providing attractive performances. However, the effectiveness of traditional MOT methods is significantly reduced when it comes to dynamic platforms like drones…

Cited by 8SourcecodeScholar
2024

Parameter-efficient Prompt Learning for 3D Point Cloud Understanding

ICRA 2024poster

This paper presents a parameter-efficient prompt tuning method, named PPT, to adapt a large multi-modal model for 3D point cloud understanding. Existing strategies are quite expensive in computation and storage, and depend on timeconsuming prompt engineering. We address the problems from three aspec…

Cited by 7SourcecodeScholar
2024

Point-PRC: A Prompt Learning Based Regulation Framework for Generalizable Point Cloud Analysis

NeurIPS 2024poster

This paper investigates the 3D domain generalization (3DDG) ability of large 3D models based on prevalent prompt learning. Recent works demonstrate the performances of 3D point cloud recognition can be boosted remarkably by parameter-efficient prompt tuning. However, we observe that the improvement…

2024

VOLoc: Visual Place Recognition by Querying Compressed Lidar Map

ICRA 2024poster

The availability of city-scale Lidar maps enables the potential of city-scale place recognition using mobile cameras. However, the city-scale Lidar maps generally need to be compressed for storage efficiency, which increases the difficulty of direct visual place recognition in compressed Lidar maps.…

Cited by 7SourcecodeScholar
2023

ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding

ICRA 2023poster

Recently, a growing number of work design unsupervised paradigms for point cloud processing to alleviate the limitation of expensive manual annotation and poor transferability of supervised methods. Among them, CrossPoint follows the contrastive learning framework and exploits image and point cloud…

Cited by 11SourcecodeScholar