← Search

Jia-Xing Zhong

10 accepted papers

2026

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

ICRA 2026poster

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while …

2025

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

RA-L 2025

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while

Cited by 18SourceScholar
2024

Learning to Catch Reactive Objects with a Behavior Predictor

ICRA 2024poster

Tracking and catching moving objects is an important ability for robots in a dynamic world. Whilst some objects have highly predictable state evolution e.g., the ballistic trajectory of a tennis ball, reactive targets alter their behavior in response to motion of the manipulator. Reactive applicatio…

Cited by 2SourcecodeScholar
2024

Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search

IROS 2024poster

In this paper, we address a recent trend in robotic home appliances to include vision systems on personal devices, capable of personalizing the appliances on the fly. In particular, we formulate and address an important technical task of personal object search, which involves localization and identi…

Cited by 2SourcecodeScholar
2023

DynPoint: Dynamic Neural Point For View Synthesis

NeurIPS 2023poster

The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require extensive training time specific to each new scenario. To tackle t…

Cited by 18SourcePDFScholar
2023

Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion Estimation

NeurIPS 2023poster

A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a traini…

2022

No Pain, Big Gain: Classify Dynamic Point Cloud Sequences With Static Models by Fitting Feature-Level Space-Time Surfaces

CVPR 2022poster

Scene flow is a powerful tool for capturing the motion field of 3D point clouds. However, it is difficult to directly apply flow-based models to dynamic point cloud classification since the unstructured points make it hard or even impossible to efficiently and effectively trace point-wise correspond…

Cited by 29PDFcodeScholar
2020

ROIMIX: Proposal-Fusion Among Multiple Images for Underwater Object Detection

ICASSP 2020accepted

Generic object detection algorithms have proven their excellent performance in recent years. However, object detection on underwater datasets is still less explored. In contrast to generic datasets, underwater images usually have color shift and low contrast; sediment would cause blurring in underwa…

Cited by 0SourceScholar
2019

Graph Convolutional Label Noise Cleaner: Train a Plug-And-Play Action Classifier for Anomaly Detection

CVPR 2019poster

Video anomaly detection under weak labels is formulated as a typical multiple-instance learning problem in previous works. In this paper, we provide a new perspective, i.e., a supervised learning task under noisy labels. In such a viewpoint, as long as cleaning away label noise, we can directly appl…

Cited by 590PDFcodeScholar