← Search

Runze Zhang

20 accepted papers

2026

CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation

CVPR 2026

Despite major advances brought by diffusion-based models, current 3D texture generation systems remain hindered by cross-view inconsistency -- textures that appear convincing from one viewpoint often fail to align across others. We find that this issue arises from attention ambiguity, where unstruct

Cited by 0SourceScholar
2026

CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling

CVPR 2026

Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional roles in a canonical space -- wings extend laterally, handles protrude to the side, and legs support from below. Psychoph

Cited by 0SourcecodeScholar
2026

LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context

ICLR 2026poster

Physically-based rendering (PBR) provides a principled standard for realistic material–lighting interactions in computer graphics. Despite recent advances in generating PBR textures, existing methods fail to address two fundamental challenges: 1) materials decomposition from image prompts under limi…

Cited by 0SourcecodeScholar
2026

Velocity Potential Field Modulation for Dense Coordination of Polytopic Swarms and Its Application to Assistive Robotic Furniture

ICRA 2026poster

We explore the use of mobile furniture swarms that are intended to assist users with limited mobility in their daily indoor activities. We focus on the multi-agent coordination problem for a mobile furniture swarm when a dense target pose configuration is required, such as in an apartment setting. I…

Cited by 0SourceScholar
2025

ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse Points

CVPR 2025highlight

We introduce ArcPro, a novel learning framework built on architectural programs to recover structured 3D abstractions from highly sparse and low-quality point clouds. Specifically, we design a domain-specific language (DSL) to hierarchically represent building structures as a program, which can be e…

Cited by 0SourcePDFScholar
2025

DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation

ICCV 2025poster

Spatio-temporal consistency is a critical topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily foc…

2025

Velocity Potential Field Modulation for Dense Coordination of Polytopic Swarms and Its Application to Assistive Robotic Furniture

RA-L 2025

We explore the use of a mobile furniture swarm that are intended to assist users with limited mobility in their daily indoor activities. We focus on the multi-robot coordination problem when a dense target pose configuration is required, such as in an apartment setting. In those cases, the convergen

Cited by 2SourceScholar
2024

Image Content Generation with Causal Reasoning

AAAI 2024technical

The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning…

2023

JR2Net: Joint Monocular 3D Face Reconstruction and Reenactment

AAAI 2023technical

Face reenactment and reconstruction benefit various applications in self-media, VR, etc. Recent face reenactment methods use 2D facial landmarks to implicitly retarget facial expressions and poses from driving videos to source images, while they suffer from pose and expression preservation issues fo…

Cited by 3SourcePDFScholar
2022

LiDAL: Inter-Frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation

ECCV 2022poster

"We propose LiDAL, a novel active learning method for 3D LiDAR semantic segmentation by exploiting inter-frame uncertainty among LiDAR frames. Our core idea is that a well-trained model should generate robust results irrespective of viewpoints for scene scanning and thus the inconsistencies in model…

2021

VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation

ICCV 2021poster

In recent years, sparse voxel-based methods have become the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and str…

Cited by 75PDFcodeScholar
2020

Dense Hybrid Recurrent Multi-view Stereo Net with Dynamic Consistency Checking

ECCV 2020poster

In this paper, we propose an efficient and effective dense hybrid recurrent multi-view stereo net with dynamic consistency checking, namely $D^{2}$HC-RMVSNet, for accurate dense point cloud reconstruction. Our novel hybrid recurrent multi-view stereo net consists of two core modules: 1) a light DREN…

2020

Pyramid Multi-view Stereo Net with Self-adaptive View Aggregation

ECCV 2020poster

In this paper, we propose an effective and efficient pyramid multi-view stereo (MVS) net with self-adaptive view aggregation for accurate and complete dense point cloud reconstruction. Different from using mean square variance to generate cost volume in previous deep-learning based MVS methods, our…

2019

Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation

ICRA 2019poster

Accurate relative pose is one of the key components in visual odometry (VO) and simultaneous localization and mapping (SLAM). Recently, the self-supervised learning framework that jointly optimizes the relative pose and target image depth has attracted the attention of the community. Previous works…

Cited by 113SourcecodeScholar
2018

GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints

ECCV 2018poster

Learned local descriptors based on Convolutional Neural Networks (CNNs) have achieved significant improvements on patch-based benchmarks, whereas not having demonstrated strong generalization ability on recent benchmarks of image-based 3D reconstruction. In this paper, we mitigate this limitation by…

Cited by 216SourcePDFScholar
2018

Learning and Matching Multi-View Descriptors for Registration of Point Clouds

ECCV 2018poster

Critical to the registration of point clouds is the establishment of a set of accurate correspondences between points in 3D space. The correspondence problem is generally addressed by the design of discriminative 3D local descriptors on the one hand, and the development of robust matching strategies…

Cited by 58SourcePDFScholar
2018

Very Large-Scale Global SfM by Distributed Motion Averaging

CVPR 2018poster

Global Structure-from-Motion (SfM) techniques have demonstrated superior efficiency and accuracy than the conventional incremental approach in many recent studies. This work proposes a divide-and-conquer framework to solve very large global SfM at the scale of millions of images. Specifically, we fi…

Cited by 183SourcePDFScholar
2015

Joint Camera Clustering and Surface Segmentation for Large-Scale Multi-View Stereo

ICCV 2015poster

In this paper, we propose an optimal decomposition approach to large-scale multi-view stereo from an initial sparse reconstruction. The success of the approach depends on the introduction of surface-segmentation-based camera clustering rather than sparse-point-based camera clustering, which suffers…

Cited by 29PDFScholar