← Search

Zhengzhe Liu

20 accepted papers

2026

EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

ICML 2026poster

This paper introduces EPS3D, a new end-to-end feed-forward framework for open-vocabulary 3D panoptic segmentation. Unlike existing methods relying on additional preprocessing, we design an end-to-end architecture, with a distillation-based training strategy on diverse 3D scenes to predict 3D-aware s…

Cited by 0SourceScholar
2026

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

ICML 2026poster

Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the firs…

Cited by 0SourceScholar
2025

COS3D: Collaborative Open-Vocabulary 3D Segmentation

NeurIPS 2025poster

Open-vocabulary 3D segmentation is a fundamental yet challenging task, requiring a mutual understanding of both segmentation and language. However, existing Gaussian-splatting-based methods rely either on a single 3D language field, leading to inferior segmentation, or on pre-computed class-agnostic…

Cited by 0SourceScholar
2025

How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach

ICCV 2025poster

Recent advancements in video diffusion models enable the generation of photorealistic videos with impressive 3D consistency and temporal coherence. However, the extent to which these AI-generated videos simulate the 3D visual world remains underexplored. In this paper, we introduce Learned 3D Evalua…

Cited by 0SourcePDFScholar
2025

Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting

CVPR 2025poster

Lifting multi-view 2D instance segmentation to a radiance field has proven effective to enhance 3D understanding. Existing works rely on direct matching for end-to-end lifting, yielding inferior results, or employ a two-stage solution constrained by complex pre- or post-processing. In this work, we…

2024

Let the Avatar Talk using Texts without Paired Training Data

ECCV 2024poster

"This paper introduces text-driven talking avatar generation, a task that uses text to instruct both the generation and animation of an avatar. One significant obstacle in this task is the absence of paired text and talking avatar data for model training, limiting data-driven methodologies. To this…

Cited by 0SourcePDFScholar
2024

Make-A-Shape: a Ten-Million-scale 3D Shape Model

ICML 2024poster

The progression in large-scale 3D generative models has been impeded by significant resource requirements for training and challenges like inefficient representations. This paper introduces Make-A-Shape, a novel 3D generative model trained on a vast scale, using 10 million publicly-available shapes.…

2023

Command-Driven Articulated Object Understanding and Manipulation

CVPR 2023poster

We present Cart, a new approach towards articulated-object manipulations by human commands. Beyond the existing work that focuses on inferring articulation structures, we further support manipulating articulated shapes to align them subject to simple command templates. The key of Cart is to utilize…

2023

Dense Depth Completion Based on Multi-Scale Confidence and Self-Attention Mechanism for Intestinal Endoscopy

ICRA 2023poster

Doctors perform limited one-way intestine endoscopy, in which advanced surgical robots with depth sensors, such as stereo and ToF endoscopes, can only provide sparse and incomplete depth information. However, dense, accurate and instant depth estimation during endoscopy is vital for doctors to judge…

Cited by 10SourceScholar
2023

ISS: Image as Stepping Stone for Text-Guided 3D Shape Generation

ICLR 2023top-25%

Text-guided 3D shape generation remains challenging due to the absence of large paired text-shape dataset, the substantial semantic gap between these two modalities, and the structural complexity of 3D shapes. This paper presents a new framework called Image as Stepping Stone (ISS) for the task by i…

2023

MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly Detection

AAAI 2023technical

Weakly supervised detection of anomalies in surveillance videos is a challenging task. Going beyond existing works that have deficient capabilities to localize anomalies in long videos, we propose a novel glance and focus network to effectively integrate spatial-temporal information for accurate ano…

2023

Texture Generation on 3D Meshes with Point-UV Diffusion

ICCV 2023oral

In this work, we focus on synthesizing high-quality textures on 3D meshes. We present Point-UV diffusion, a coarse-to-fine pipeline that marries the denoising diffusion model with UV mapping to generate 3D consistent and high-quality texture images in UV space. We start with introducing a point diff…

Cited by 45PDFcodeScholar
2022

Sparse2Dense: Learning to Densify 3D Features for 3D Object Detection

NeurIPS 2022accept

LiDAR-produced point clouds are the major source for most state-of-the-art 3D object detectors. Yet, small, distant, and incomplete objects with sparse or few points are often hard to detect. We present Sparse2Dense, a new framework to efficiently boost 3D detection performance by learning to densif…

2022

TWIST: Two-Way Inter-Label Self-Training for Semi-Supervised 3D Instance Segmentation

CVPR 2022poster

We explore the way to alleviate the label-hungry problem in a semi-supervised setting for 3D instance segmentation. To leverage the unlabeled data to boost model performance, we present a novel Two-Way Inter-label Self-Training framework named TWIST. It exploits inherent correlations between semanti…

Cited by 29PDFcodeScholar
2021

One Thing One Click: A Self-Training Approach for Weakly Supervised 3D Semantic Segmentation

CVPR 2021poster

Point cloud semantic segmentation often requires largescale annotated training data, but clearly, point-wise labels are too tedious to prepare. While some recent methods propose to train a 3D network with small percentages of point labels, we take the approach to an extreme and propose "One Thing On…

Cited by 171PDFcodeScholar
2018

GeoNet: Geometric Neural Network for Joint Depth and Surface Normal Estimation

CVPR 2018poster

In this paper, we propose Geometric Neural Network (GeoNet) to jointly predict depth and surface normal maps from a single image. Building on top of two-stream CNNs, our GeoNet incorporates geometric relation between depth and surface normal via the new depth-to-normal and normal- to-depth networks.…

Cited by 428SourcePDFScholar