← Search

Jia-Wang Bian

11 accepted papers

2026

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

ICML 2026poster

Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces…

Cited by 0SourceScholar
2025

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

RA-L 2025

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while

Cited by 18SourceScholar
2025

RoboPearls: Editable Video Simulation for Robot Manipulation

ICCV 2025poster

The development of generalist robot manipulation policies has seen significant progress, driven by large-scale demonstration data across diverse environments. However, the high cost and inefficiency of collecting real-world demonstrations hinder the scalability of data acquisition. While existing si…

Cited by 0SourcePDFScholar
2025

SurfaceSplat: Connecting Surface Reconstruction and Gaussian Splatting

ICCV 2025poster

Surface reconstruction and novel view rendering from sparse-view images are challenging. Signed Distance Function (SDF)-based methods struggle with fine details, while 3D Gaussian Splatting (3DGS)-based approaches lack global geometry coherence. We propose a novel hybrid method that combines both st…

2024

GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing

ECCV 2024poster

"We propose , a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS and edits them by using a pre-trained 2D diffusion model (ControlNet) based on the input prompt, which is then used to optimise t…

2024

Neural Refinement for Absolute Pose Regression with Feature Synthesis

CVPR 2024poster

Absolute Pose Regression (APR) methods use deep neural networks to directly regress camera poses from RGB images. However the predominant APR architectures only rely on 2D operations during inference resulting in limited accuracy of pose estimation due to the lack of 3D geometry constraints or prior…

2024

PORF: POSE RESIDUAL FIELD FOR ACCURATE NEURAL SURFACE RECONSTRUCTION

ICLR 2024poster

Neural surface reconstruction is sensitive to the camera pose noise, even when state-of-the-art pose estimators like COLMAP or ARKit are used. Existing Pose-NeRF joint optimisation methods have struggled to improve pose accuracy in challenging real-world scenarios. To overcome the challenges, we int…

2023

MobileBrick: Building LEGO for 3D Reconstruction on Mobile Devices

CVPR 2023poster

High-quality 3D ground-truth shapes are critical for 3D object reconstruction evaluation. However, it is difficult to create a replica of an object in reality, and even 3D reconstructions generated by 3D scanners have artefacts that cause biases in evaluation. To address this issue, we introduce a n…

2023

NoPe-NeRF: Optimising Neural Radiance Field With No Pose Prior

CVPR 2023highlight

Training a Neural Radiance Field (NeRF) without pre-computed camera poses is challenging. Recent advances in this direction demonstrate the possibility of jointly optimising a NeRF and camera poses in forward-facing scenes. However, these methods still face difficulties during dramatic camera moveme…

2021

Diverse Knowledge Distillation for End-to-End Person Search

AAAI 2021technical

Person search aims to localize and identify a specific person from a gallery of images. Recent methods can be categorized into two groups, i.e., two-step and end-to-end approaches. The former views person search as two independent tasks and achieves dominant results using separately trained person d…

Cited by 47SourcePDFScholar
2020

Visual Odometry Revisited: What Should Be Learnt?

ICRA 2020poster

In this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance are based on geometry and have to be carefully designed for different application scenarios. Moreover, most monocular s…

Cited by 230SourcecodeScholar