← Search

Qian Xie

13 accepted papers

2026

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

ICRA 2026poster

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while …

2025

EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models

AAAI 2025technical

The rapid development of diffusion models has significantly advanced AI-generated content (AIGC), particularly in Text-to-Image (T2I) and Text-to-Video (T2V) generation. Text-based video editing, leveraging these generative capabilities, has emerged as a promising field, enabling precise modificatio…

2025

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

RA-L 2025

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while

Cited by 18SourceScholar
2025

Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data

IROS 2025

Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive

Cited by 0SourceScholar
2024

Cost-aware Bayesian Optimization via the Pandora's Box Gittins Index

NeurIPS 2024poster

Bayesian optimization is a technique for efficiently optimizing unknown functions in a black-box manner. To handle practical settings where gathering data requires use of finite resources, it is desirable to explicitly incorporate function evaluation costs into Bayesian optimization policies. To und…

2024

Towards Learning Group-Equivariant Features for Domain Adaptive 3D Detection

NeurIPS 2024poster

The performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a…

Cited by 0SourcePDFScholar
2022

Meta-Sampler: Almost-Universal yet Task-Oriented Sampling for Point Clouds

ECCV 2022poster

"Sampling is a key operation in point-cloud task and acts to increase computational efficiency and tractability by discarding redundant points. Universal sampling algorithms (e.g., Farthest Point Sampling) work without modification across different tasks, models, and datasets, but by their very natu…

2021

MLVSNet: Multi-Level Voting Siamese Network for 3D Visual Tracking

ICCV 2021poster

Benefiting from the excellent performance of Siamese-based trackers, huge progress on 2D visual tracking has been achieved. However, 3D visual tracking is still under-explored. Inspired by the idea of Hough voting in 3D object detection, in this paper, we propose a Multi-level Voting Siamese Network…

Cited by 66PDFcodeScholar
2021

Robust and Accurate RGB-D Reconstruction With Line Feature Constraints

RA-L 2021

Scene reconstruction with consumer-level RGB-D cameras has developed considerable momentum in both robotics and vision communities. In the literature of robotics, high-quality camera tracking, the key to accurate reconstruction, is challenging in geometric featureless scenes or under large lighting

Cited by 5SourceScholar
2021

VENet: Voting Enhancement Network for 3D Object Detection

ICCV 2021poster

Hough voting, as has been demonstrated in VoteNet, is effective for 3D object detection, where voting is a key step. In this paper, we propose a novel VoteNet-based 3D detector with vote enhancement to improve the detection accuracy in cluttered indoor scenes. It addresses the limitations of current…

Cited by 61PDFScholar
2020

MLCVNet: Multi-Level Context VoteNet for 3D Object Detection

CVPR 2020poster

In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects individually, without giving any consideration on contextual informatio…

Cited by 229PDFcodeScholar