← Search

Zhen Wei

8 accepted papers

2025

DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

ICML 2025poster

Several recent studies have attempted to autoregressively generate continuous speech representations without discrete speech tokens by combining diffusion and autoregressive models, yet they often face challenges with excessive computational loads or suboptimal outcomes. In this work, we propose Dif…

Cited by 1SourcePDFScholar
2023

Fast and Accurate Binary Neural Networks Based on Depth-Width Reshaping

AAAI 2023technical

Network binarization (i.e., binary neural networks, BNNs) can efficiently compress deep neural networks and accelerate model inference but cause severe accuracy degradation. Existing BNNs are mainly implemented based on the commonly used full-precision network backbones, and then the accuracy is imp…

2023

Robust Outlier Rejection for 3D Registration With Variational Bayes

CVPR 2023poster

Learning-based outlier (mismatched correspondence) rejection for robust 3D registration generally formulates the outlier removal as an inlier/outlier classification problem. The core for this to be successful is to learn the discriminative inlier/outlier feature representations. In this paper, we de…

2021

FBNetV3: Joint Architecture-Recipe Search Using Predictor Pretraining

CVPR 2021poster

Neural Architecture Search (NAS) yields state-of-the-art neural networks that outperform their best manually-designed counterparts. However, previous NAS methods search for architectures under one set of training hyper-parameters (i.e., a training recipe), overlooking superior architecture-recipe co…

Cited by 133PDFScholar
2019

Building Detail-Sensitive Semantic Segmentation Networks With Polynomial Pooling

CVPR 2019poster

Semantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classif…

Cited by 34PDFScholar
2018

Augmenting Crowd-Sourced 3D Reconstructions Using Semantic Detections

CVPR 2018poster

Image-based 3D reconstruction for Internet photo collections has become a robust technology to produce impressive virtual representations of real-world scenes. However, several fundamental challenges remain for Structure-from-Motion (SfM) pipelines, namely: the placement and reconstruction of transi…

Cited by 9SourcePDFScholar