← Search

Runmin Zhang

10 accepted papers

2026

Large Depth Completion Model from Sparse Observations

ICLR 2026poster

This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without relying on complex architectural designs, LDCM generates metric-accurate dense depth maps in one large transformer. It outpe…

Cited by 0SourceScholar
2026

Rethinking Unsupervised Cross-modal Flow Estimation: Learning from Decoupled Optimization and Consistency Constraint

ICLR 2026poster

This work presents DCFlow, a novel self-supervised cross-modal flow estimation framework that integrates a decoupled optimization strategy and a cross-modal consistency constraint. Unlike previous unsupervised approaches that implicitly learn flow estimation solely from appearance similarity, we int…

Cited by 0SourceScholar
2025

Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume Construction

ICCV 2025poster

This work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate…

2025

EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration

ICCV 2025poster

Previous deep image registration methods that employ single homography, multi-grid homography, or thin-plate spline often struggle with real scenes containing depth disparities due to their inherent limitations. To address this, we propose an Exponential-Decay Free-Form Deformation Network (EDFFDNet…

Cited by 0SourcePDFScholar
2025

Language Driven Occupancy Prediction

ICCV 2025poster

We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model…

2025

Learned Image Transmission with Hierarchical Variational Autoencoder

AAAI 2025technical

In this paper, we introduce an innovative hierarchical joint source-channel coding (HJSCC) framework for image transmission, utilizing a hierarchical variational autoencoder (VAE). Our approach leverages a combination of bottom-up and top-down paths at the transmitter to autoregressively generate mu…

2025

SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split Optimization

CVPR 2025highlight

We propose a novel unsupervised cross-modal homography estimation learning framework, named Split Supervised Homography estimation Network (SSHNet). SSHNet reformulates the unsupervised cross-modal homography estimation into two supervised sub-problems, each addressed by its specialized network: a h…

2024

Context and Geometry Aware Voxel Transformer for Semantic Scene Completion

NeurIPS 2024spotlight

Vision-based Semantic Scene Completion (SSC) has gained much attention due to its widespread applications in various 3D perception tasks. Existing sparse-to-dense approaches typically employ shared context-independent queries across various input images, which fails to capture distinctions among the…

2023

Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus Transformer

CVPR 2023poster

We propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the a…