← Search

Weiyue Wang

16 accepted papers

2026

DeepL Voice: Real-Time Speech-to-Speech Translation

IJCAI 2026

DeepL Voice is a real-time speech-to-speech translation system for global business communication, following a pragmatic incremental approach: developing a production-grade cascaded speech-to-speech-translation (S2ST) system, while exploring end-to-end solutions in parallel. The production system (la

Cited by 0Scholar
2024

WOMD-LiDAR: Raw Sensor Dataset Benchmark for Motion Forecasting

ICRA 2024poster

Widely adopted motion forecasting datasets sub-stitute the observed sensory inputs with higher-level abstractions such as 3D boxes and polylines. These sparse shapes are inferred through annotating the original scenes with perception systems’ predictions. Such intermediate representations tie the qu…

Cited by 28SourceScholar
2022

Multi-Class 3D Object Detection with Single-Class Supervision

ICRA 2022poster

While multi-class 3D detectors are needed in many robotics applications, training them with fully labeled datasets can be expensive in labeling cost. An alternative approach is to have targeted single-class labels on disjoint data samples. In this paper, we are interested in training a multi-class 3…

Cited by 2SourceScholar
2022

PseudoAugment: Learning to Use Unlabeled Data for Data Augmentation in Point Clouds

ECCV 2022poster

"Data augmentation is an important technique to improve data efficiency and to save labeling cost for 3D detection in point clouds. Yet, existing augmentation policies have so far been designed to only utilize labeled data, which limits the data diversity. In this paper, we recognize that pseudo lab…

Cited by 18SourcePDFScholar
2022

SWFormer: Sparse Window Transformer for 3D Object Detection in Point Clouds

ECCV 2022poster

"3D object detection in point clouds is a core component for modern robotics and autonomous driving systems. A key challenge in 3D object detection comes from the inherent sparse nature of point occupancy within the 3D scene. In this paper, we propose Sparse Window Transformer (SWFormer ), a scalabl…

Cited by 140SourcePDFScholar
2021

RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object Detection

CVPR 2021poster

The detection of 3D objects from LiDAR data is a critical component in most autonomous driving systems. Safe, high speed driving needs larger detection ranges, which are enabled by new LiDARs. These larger detection ranges require more efficient and accurate detection models. Towards this goal, we p…

Cited by 204PDFScholar
2021

SPG: Unsupervised Domain Adaptation for 3D Object Detection via Semantic Point Generation

ICCV 2021poster

In autonomous driving, a LiDAR-based object detector should perform reliably at different geographic locations and under various weather conditions. While recent 3D detection research focuses on improving performance within a single domain, our study reveals that the performance of modern detectors…

Cited by 198PDFcodeScholar
2021

To the Point: Efficient 3D Object Detection in the Range Image With Graph Convolution Kernels

CVPR 2021poster

3D object detection is vital for many robotics applications. For tasks where a 2D perspective range image exists, we propose to learn a 3D representation directly from this range image view. To this end, we designed a 2D convolutional network architecture that carries the 3D spherical coordinates of…

Cited by 87PDFScholar
2020

Neural Language Modeling for Named Entity Recognition

COLING 2020main

Named entity recognition is a key component in various natural language processing systems, and neural architectures provide significant improvements over conventional approaches. Regardless of different word embedding and hidden layer structures of the networks, a conditional random field layer is…

2019

DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction

NeurIPS 2019poster

Reconstructing 3D shapes from single-view images has been a long-standing research problem. In this paper, we present DISN, a Deep Implicit Surface Net- work which can generate a high-quality detail-rich 3D mesh from a 2D image by predicting the underlying signed distance fields. In addition to util…

2018

SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation

CVPR 2018poster

We introduce Similarity Group Proposal Network (SGPN), a simple and intuitive deep learning framework for 3D object instance segmentation on point clouds. SGPN uses a single network to predict point grouping proposals and a corresponding semantic class for each proposal, from which we can directly…

2017

Self-paced cross-modality transfer learning for efficient road segmentation

ICRA 2017poster

Accurate road segmentation is a prerequisite for autonomous driving. Current state-of-the-art methods are mostly based on convolutional neural networks (CNNs). Nevertheless, their good performance is at expense of abundant annotated data and high computational cost. In this work, we address these tw…

Cited by 20SourceScholar
2017

Shape Inpainting Using 3D Generative Adversarial Network and Recurrent Convolutional Networks

ICCV 2017poster

Recent advances in convolutional neural networks have shown promising results in 3D shape completion. But due to GPU memory limitations, these methods can only produce low-resolution outputs. To inpaint 3D models with semantic plausibility and contextual details, we introduce a hybrid framework that…

Cited by 214PDFScholar