← Search

Wenping Wang

90 accepted papers

2026

3DGS-HPC: Distractor-free 3D Gaussian Splatting with Hybrid Patch-wise Classification

ICML 2026poster

3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and 3D scene reconstruction, but its quality often degrades in real-world environments due to transient distractors, such as moving objects and varying shadows. Existing methods commonly rely on semantic cue…

Cited by 0SourceScholar
2026

DentalGS: Pose-Free 3D Gaussian Splatting from Five Intraoral Images for Novel View Synthesis

AAAI 2026technical

Orthodontic treatment needs regular tooth alignment checks, but current methods depend on clinic visits, limiting remote care. With the emergence of 3D Gaussian Splatting (3DGS), realistic novel views can be synthesized, making it possible for clinicians to remotely monitor orthodontic conditions. H

Cited by 0SourcePDFScholar
2026

Dynamic Gaussian Scene Reconstruction from Unsynchronized Videos

AAAI 2026technical

Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 3D Gaussian Splatting have demonstrated impressive capabilities in dynamic scene reconstruction, they typically rely on

Cited by 0SourcePDFScholar
2026

Learning What Matters: Adaptive Information Theoretic Objectives for Robot Exploration

RSS 2026poster

Designing learnable information-theoretic objectives for robot exploration remains challenging. Such objectives aim to guide exploration toward data that reduces uncertainty in model parameters, yet it is often unclear what information the collected data can actually reveal. Although reinforcement l…

Cited by 0SourceScholar
2026

MEGS^{2}: Memory-Efficient Gaussian Splatting via Spherical Gaussians and Unified Pruning

ICLR 2026poster

3D Gaussian Splatting (3DGS) has emerged as a dominant novel-view synthesis technique, but its high memory consumption severely limits its applicability on edge devices. A growing number of 3DGS compression methods have been proposed to make 3DGS more efficient, yet most only focus on storage compre…

Cited by 0SourcecodeScholar
2026

MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly

CVPR 2026

Scaling artist-designed meshes to high triangle numbers remains challenging for autoregressive generative models. Existing transformer-based methods suffer from long-sequence bottlenecks and limited quantization resolution, primarily due to the large number of tokens required and constrained quantiz

Cited by 0SourcecodeScholar
2026

Metric–-Phase Fields: Decoupling Distance and Sign for Thin-Structure Reconstruction from Unoriented Point Clouds

ICML 2026poster

Neural Signed Distance Functions (SDFs) excel at reconstructing watertight manifolds but fail on thin structures and open boundaries due to strict inside-outside constraints. Conversely, Unsigned Distance Fields (UDFs) accommodate general geometries but suffer from gradient singularities at the zero…

Cited by 0SourceScholar
2026

NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representations

AAAI 2026technical

3D Gaussian Splatting (3DGS) achieves impressive quality and rendering speed, but with millions of 3D Gaussians and significant storage and transmission costs. In this paper, we aim to develop a simple yet effective method called NeuralGS that compresses the original 3DGS into a compact representati

Cited by 0SourcePDFScholar
2026

UniSER: A Foundation Model for Unified Soft Effects Removal

CVPR 2026

Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models tha

Cited by 0SourceScholar
2025

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

CVPR 2025highlight

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the i…

Cited by 14SourcePDFScholar
2025

CADDreamer: CAD Object Generation from Single-view Images

CVPR 2025highlight

The field of diffusion-based 3D generation has experienced tremendous progress in recent times. However, existing 3D generative models often produce overly dense and unstructured meshes, which are in stark contrast to the compact, structured and clear-edged CAD models created by human modelers. We i…

Cited by 0SourcePDFScholar
2025

CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMs

ICLR 2025poster

In this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point…

Cited by 0SourcePDFScholar
2025

Controllable 3D Outdoor Scene Generation via Scene Graphs

ICCV 2025poster

Three-dimensional scene generation is crucial in computer vision, with applications spanning autonomous driving and gaming. However, current methods offer limited or non-intuitive user control. In this work, we propose a method that uses scene graph as a user-friendly control format to generate outd…

2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild

CVPR 2025poster

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task.Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view reconstruction faces significant challenges due to inherent ambiguities an…

2025

GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global Illumination

ICCV 2025poster

3D Gaussian Splatting (3DGS) is a prevailing technique to reconstruct large-scale 3D scenes from multiview images for novel view synthesis, like a room, a block, and even a city. Such large-scale scenes are not static with changes constantly happening in these scenes, like a new building being built…

Cited by 0SourcePDFScholar
2025

MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors

ICLR 2025poster

In this paper, we propose MoDGS, a new pipeline to render novel-view images in dynamic scenes using only casually captured monocular videos. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid movement of input cameras to construct multiview consistency but fail…

2025

Simplification Is All You Need against Out-of-Distribution Overconfidence

CVPR 2025poster

Deep neural networks (DNNs) often exhibit out-of-distribution (OOD) overconfidence, producing overly confident predictions on OOD samples. We attribute this issue to the inherent over-complexity of DNNs and investigate two key aspects: capacity and nonlinearity. First, we demonstrate that reducing m…

Cited by 3SourcePDFScholar
2025

Surprise3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes

NeurIPS 2025poster

The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between objects, remains underexplored in current 3D vision-languag…

Cited by 0SourceScholar
2025

VistaDream: Sampling multiview consistent images for single-view scene reconstruction

ICCV 2025poster

In this paper, we propose VistaDream, a novel framework to reconstruct a 3D scene from a single-view image. Recent diffusion models enable generating high-quality novel-view images from a single-view input image. Most existing methods only concentrate on building the consistency between the input im…

Cited by 0SourcePDFScholar
2025

🎧MOSPA: Human Motion Generation Driven by Spatial Audio

NeurIPS 2025spotlight

Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have…

Cited by 0SourcecodeScholar
2024

"EMDM: Efficient Motion Diffusion Model for Fast, High-Quality Human Motion Generation"

ECCV 2024poster

"We introduce Efficient Motion Diffusion Model (EMDM) for fast and high-quality human motion generation. Current state-of-the-art generative diffusion models have produced impressive results but struggle to achieve fast generation without sacrificing quality. On the one hand, previous works, like mo…

2024

CORES: Convolutional Response-based Score for Out-of-distribution Detection

CVPR 2024poster

Deep neural networks (DNNs) often display overconfidence when encountering out-of-distribution (OOD) samples posing significant challenges in real-world applications. Capitalizing on the observation that responses on convolutional kernels are generally more pronounced for in-distribution (ID) sample…

Cited by 6SourcePDFScholar
2024

Collaborative Tooth Motion Diffusion Model in Digital Orthodontics

AAAI 2024technical

Tooth motion generation is an essential task in digital orthodontic treatment for precise and quick dental healthcare, which aims to generate the whole intermediate tooth motion process given the initial pathological and target ideal tooth alignments. Most prior works for multi-agent motion plannin…

Cited by 2SourcePDFScholar
2024

DynoSurf: Neural Deformation-based Temporally Consistent Dynamic Surface Reconstruction

ECCV 2024poster

"This paper explores the problem of reconstructing temporally consistent surfaces from a 3D point cloud sequence without correspondence. To address this challenging task, we propose DynoSurf, an unsupervised learning framework integrating a template surface representation with a learnable deformatio…

2024

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

NeurIPS 2024poster

In this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resu…

Cited by 7SourcePDFScholar
2024

Flatten Anything: Unsupervised Neural Surface Parameterization

NeurIPS 2024poster

Surface parameterization plays an essential role in numerous computer graphics and geometry processing applications. Traditional parameterization approaches are designed for high-quality meshes laboriously created by specialized 3D modelers, thus unable to meet the processing demand for the current…

2024

FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators

ICLR 2024poster

Matching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning m…

2024

GaussianPro: 3D Gaussian Splatting with Progressive Propagation

ICML 2024poster

3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain tex…

2024

Generalizable Thermal-based Depth Estimation via Pre-trained Visual Foundation Model

ICRA 2024poster

Depth estimation is a crucial task in computer vision, applicable to various domains such as 3D reconstruction, robotics, and autonomous driving. In particular, thermal-based depth estimation has unique advantages, including night-time vision. However, the existing depth estimation method remains ch…

Cited by 0SourceScholar
2024

Language-Augmented Symbolic Planner for Open-World Task Planning

RSS 2024poster

Enabling robotic agents to perform complex long-horizon tasks has been a long-standing goal in robotics and artificial intelligence (AI). Despite the potential shown by large language models (LLMs), their planning capabilities remain limited to short-horizon tasks and they are unable to replace the…

2024

MMPI: a Flexible Radiance Field Representation by Multiple Multi-plane Images Blending

ICRA 2024poster

This paper presents a flexible representation of neural radiance fields based on multi-plane images (MPI), for high-quality view synthesis of complex scenes. MPI with Normalized Device Coordinate (NDC) parameterization is widely used in NeRF learning for its simple definition, easy calculation, and…

Cited by 4SourceScholar
2024

Manifold Constraints for Imperceptible Adversarial Attacks on Point Clouds

AAAI 2024technical

Adversarial attacks on 3D point clouds often exhibit unsatisfactory imperceptibility, which primarily stems from the disregard for manifold-aware distortion, i.e., distortion of the underlying 2-manifold surfaces. In this paper, we develop novel manifold constraints to reduce such distortion, aiming…

Cited by 11SourcePDFScholar
2024

PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction

ICLR 2024spotlight

We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in ~1.3 seconds on a single A100 GPU. PF-LRM is a highly scalable method utilizing self-atten…

2024

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

IJCAI 2024poster

In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and multimodal visual data. Utilizing a teleoperation system, we seamlessly synchronize human-robot hand poses in real time. Th…

2024

SMaRt: Improving GANs with Score Matching Regularity

ICML 2024poster

Generative adversarial networks (GANs) usually struggle in learning from highly diverse data, whose underlying manifold is complex. In this work, we revisit the mathematical foundations of GANs, and theoretically reveal that the native adversarial loss for GAN training is insufficient to fix the pro…

2024

Semantic Human Mesh Reconstruction with Textures

CVPR 2024poster

The field of 3D detailed human mesh reconstruction has made significant progress in recent years. However current methods still face challenges when used in industrial applications due to unstable results low-quality meshes and a lack of UV unwrapping and skinning weights. In this paper we present S…

2024

SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

ICLR 2024spotlight

In this paper, we present a novel diffusion model called SyncDreamer that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, recent work Zero123 demonstrates the ability to generate plausible novel views from a single-view image of an ob…

2024

TLControl: Trajectory and Language Control for Human Motion Synthesis

ECCV 2024poster

"Controllable human motion synthesis is essential for applications in AR/VR, gaming and embodied AI. Existing methods often focus solely on either language or full trajectory control, lacking precision in synthesizing motions aligned with user-specified trajectories, especially for multi-joint contr…

Cited by 49SourcePDFScholar
2024

Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner

CVPR 2024poster

A diffusion model which is formulated to produce an image using thousands of denoising steps usually suffers from a slow inference speed. Existing acceleration algorithms simplify the sampling by skipping most steps yet exhibit considerable performance degradation. By viewing the generation of diffu…

2024

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

CVPR 2024highlight

In this work we introduce Wonder3D a novel method for generating high-fidelity textured meshes from single-view images with remarkable efficiency. Recent methods based on the Score Distillation Sampling (SDS) loss methods have shown the potential to recover 3D geometry from 2D diffusion priors but t…

Cited by 414SourcePDFScholar
2023

Aligning Gradient and Hessian for Neural Signed Distance Function

NeurIPS 2023poster

The Signed Distance Function (SDF), as an implicit surface representation, provides a crucial method for reconstructing a watertight surface from unorganized point clouds. The SDF has a fundamental relationship with the principles of surface vector calculus. Given a smooth surface, there exists a th…

Cited by 4SourcePDFScholar
2023

Batch-based Model Registration for Fast 3D Sherd Reconstruction

ICCV 2023poster

3D reconstruction techniques have widely been used for digital documentation of archaeological fragments. However, efficient digital capture of fragments remains as a challenge. In this work, we aim to develop a portable, high-throughput, and accurate reconstruction system for efficient digitization…

Cited by 2PDFScholar
2023

CLIP2Scene: Towards Label-Efficient 3D Scene Understanding by CLIP

CVPR 2023poster

Contrastive Language-Image Pre-training (CLIP) achieves promising results in 2D zero-shot and few-shot learning. Despite the impressive performance in 2D, applying CLIP to help the learning in 3D scene understanding has yet to be explored. In this paper, we make the first attempt to investigate how…

2023

Deep Manifold Attack on Point Clouds via Parameter Plane Stretching

AAAI 2023technical

Adversarial attack on point clouds plays a vital role in evaluating and improving the adversarial robustness of 3D deep learning models. Current attack methods are mainly applied by point perturbation in a non-manifold manner. In this paper, we formulate a novel manifold attack, which deforms the un…

Cited by 17SourcePDFScholar
2023

F2-NeRF: Fast Neural Radiance Field Training With Free Camera Trajectories

CVPR 2023highlight

This paper presents a novel grid-based NeRF called F^2-NeRF (Fast-Free-NeRF) for novel view synthesis, which enables arbitrary input camera trajectories and only costs a few minutes for training. Existing fast grid-based NeRF training frameworks, like Instant-NGP, Plenoxels, DVGO, or TensoRF, are ma…

2023

GeoUDF: Surface Reconstruction from 3D Point Clouds via Geometry-guided Distance Representation

ICCV 2023poster

We present a learning-based method, namely GeoUDF, to tackle the long-standing and challenging problem of reconstructing a discrete surface from a sparse point cloud. To be specific, we propose a geometry-guided learning method for UDF and its gradient estimation that explicitly formulates the unsig…

Cited by 27PDFcodeScholar
2023

Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition From Egocentric RGB Videos

CVPR 2023poster

Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the differ…

2023

NeTO:Neural Reconstruction of Transparent Objects with Self-Occlusion Aware Refraction-Tracing

ICCV 2023poster

We present a novel method called NeTO, for capturing the 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very challenging task, which is ill-suited for general-purpose reconstruction techniques due to the specular light transport…

Cited by 12PDFScholar
2023

NeuralUDF: Learning Unsigned Distance Fields for Multi-View Reconstruction of Surfaces With Arbitrary Topologies

CVPR 2023poster

We present a novel method, called NeuralUDF, for reconstructing surfaces with arbitrary topologies from 2D images via volume rendering. Recent advances in neural rendering based reconstruction have achieved compelling results. However, these methods are limited to objects with closed surfaces since…

Cited by 67SourcePDFScholar
2023

NeuroGF: A Neural Representation for Fast Geodesic Distance and Path Queries

NeurIPS 2023poster

Geodesics play a critical role in many geometry processing applications. Traditional algorithms for computing geodesics on 3D mesh models are often inefficient and slow, which make them impractical for scenarios requiring extensive querying of arbitrary point-to-point geodesics. Recently, deep impli…

2023

Robust Multiview Point Cloud Registration With Reliable Pose Graph Initialization and History Reweighting

CVPR 2023poster

In this paper, we present a new method for the multiview registration of point cloud. Previous multiview registration methods rely on exhaustive pairwise registration to construct a densely-connected pose graph and apply Iteratively Reweighted Least Square (IRLS) on the pose graph to compute the sca…

2023

Surface Extraction from Neural Unsigned Distance Fields

ICCV 2023poster

We propose a method, named DualMesh-UDF, to extract a surface from unsigned distance functions (UDFs), encoded by neural networks, or neural UDFs. Neural UDFs are becoming increasingly popular for surface representation because of their versatility in presenting surfaces with arbitrary topologies, a…

Cited by 8PDFcodeScholar
2023

TORE: Token Reduction for Efficient Human Mesh Recovery with Transformer

ICCV 2023poster

In this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based structures. However, they suffer from high model complexity and computation cost caus…

Cited by 51PDFcodeScholar
2023

Towards Label-free Scene Understanding by Vision Foundation Models

NeurIPS 2023poster

Vision foundation models such as Contrastive Vision-Language Pre-training (CLIP) and Segment Anything (SAM) have demonstrated impressive zero-shot performance on image classification and segmentation tasks. However, the incorporation of CLIP and SAM for label-free scene understanding has yet to be e…

2023

mCLIP: Multilingual CLIP via Cross-lingual Transfer

ACL 2023long

Large-scale vision-language pretrained (VLP) models like CLIP have shown remarkable performance on various downstream cross-modal tasks. However, they are usually biased towards English due to the lack of sufficient non-English image-text pairs. Existing multilingual VLP methods often learn retrieva…

2022

DISP6D: Disentangled Implicit Shape and Pose Learning for Scalable 6D Pose Estimation

ECCV 2022poster

"Scalable 6D pose estimation for rigid objects from RGB images aims at handling multiple objects and generalizing to novel objects. Building on a well-known auto-encoding framework to cope with object symmetry and the lack of labeled training data, we achieve scalability by disentangling the latent…

2022

FaceFormer: Speech-Driven 3D Facial Animation With Transformers

CVPR 2022oral

Speech-driven 3D facial animation is challenging due to the complex geometry of human faces and the limited availability of 3D audio-visual data. Prior works typically focus on learning phoneme-level features of short audio windows with limited context, occasionally resulting in inaccurate lip movem…

Cited by 258PDFcodeScholar
2022

Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images

ECCV 2022poster

"In this paper, we present a generalizable model-free 6-DoF object pose estimator called Gen6D. Existing generalizable pose estimators either need the high-quality object models or require additional depth maps or object masks in test time, which significantly limits their application scope. In cont…

2022

Kinematic Analysis of Soft Continuum Manipulators Based on Sparse Workspace Mapping

RA-L 2022

Soft robots, with advantages of high adaptability to the environment, relatively easy and simple fabrication process as well as promising performances, have been thoroughly investigated and widely applied lately, the superiority of which has been proved in areas such as medicine, industry, daily lif

Cited by 9SourceScholar
2022

Learn to Predict How Humans Manipulate Large-Sized Objects From Interactive Motions

RA-L 2022

Understanding human intentions during interactions has been a long-lasting theme, that has applications in human-robot interaction, virtual reality and surveillance. In this study, we focus on full-body human interactions with large-sized daily objects and aim to predict the future states of objects

Cited by 35SourceScholar
2022

NeuRIS: Neural Reconstruction of Indoor Scenes Using Normal Priors

ECCV 2022poster

"Reconstructing 3D indoor scenes from 2D images is an important task in many computer vision and graphics applications. A main challenge in this task is that large texture-less areas in typical indoor scenes make existing methods struggle to produce satisfactory reconstruction results. We propose a…

Cited by 113SourcePDFScholar
2022

Neural Rays for Occlusion-Aware Image-Based Rendering

CVPR 2022poster

We present a new neural representation, called Neural Ray (NeuRay), for the novel view synthesis task. Recent works construct radiance fields from image features of input views to render novel view images, which enables the generalization to new scenes. However, due to occlusions, a 3D point may be…

Cited by 234PDFcodeScholar
2022

ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild

ECCV 2022poster

"Estimating the pose of a moving camera from monocular video is a challenging problem, especially due to the presence of moving objects in dynamic environments, where the performance of existing camera pose estimation methods are susceptible to pixels that are not geometrically consistent. To tackle…

2022

RestoreFormer: High-Quality Blind Face Restoration From Undegraded Key-Value Pairs

CVPR 2022poster

Blind face restoration is to recover a high-quality face image from unknown degradations. As face image contains abundant contextual information, we propose a method, RestoreFormer, which explores fully-spatial attentions to model contextual information and surpasses existing works that use local co…

Cited by 123PDFcodeScholar
2022

Self-Supervised Image Representation Learning With Geometric Set Consistency

CVPR 2022poster

We propose a method for self-supervised image representation learning under the guidance of 3D geometric consistency. Our intuition is that 3D geometric consistency priors such as smooth regions and surface discontinuities may imply consistent semantics or object boundaries, and can act as strong cu…

Cited by 8PDFScholar
2022

SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views

ECCV 2022poster

"We introduce SparseNeuS, a novel neural rendering based method for the task of surface reconstruction from multi-view images. This task becomes more difficult when only sparse images are provided as input, a scenario where existing neural reconstruction approaches usually produce incomplete or dist…

Cited by 195SourcePDFScholar
2022

Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine Translation

ACL 2022long

This paper demonstrates that multilingual pretraining and multilingual fine-tuning are both critical for facilitating cross-lingual transfer in zero-shot translation, where the neural machine translation (NMT) model is tested on source languages unseen during supervised training. Following this idea…

2022

Visual-tactile Sensing for Real-time Liquid Volume Estimation in Grasping

IROS 2022poster

We propose a deep visuo-tactile model for real-time estimation of the liquid inside a deformable container in a proprioceptive way. We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our specific tactile sensor without any extra sensor calibrati…

Cited by 16SourceScholar
2021

AdaFit: Rethinking Learning-Based Normal Estimation on Point Clouds

ICCV 2021poster

This paper presents a neural network for robust normal estimation on point clouds, named AdaFit, that can deal with point clouds with noise and density variations. Existing works use a network to learn point-wise weights for weighted least squares surface fitting to estimate the normals, which has d…

Cited by 56PDFcodeScholar
2021

Adaptive Surface Normal Constraint for Depth Estimation

ICCV 2021poster

We present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints, or are limited to the difficulty of reliably capturing geometric context, which leads to a bottleneck of depth estimat…

Cited by 72PDFcodeScholar
2021

CODEs: Chamfer Out-of-Distribution Examples Against Overconfidence Issue

ICCV 2021poster

Overconfident predictions on out-of-distribution (OOD) samples is a thorny issue for deep neural networks. The key to resolve the OOD overconfidence issue inherently is to build a subset of OOD samples and then suppress predictions on them. This paper proposes the Chamfer OOD examples (CODEs), whose…

Cited by 39PDFScholar
2021

Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks

CVPR 2021poster

We present a novel method for multi-view depth estimation from a single video, which is a critical task in various applications, such as perception, reconstruction and robot navigation. Although previous learning-based methods have demonstrated compelling results, most works estimate depth maps of i…

Cited by 88PDFcodeScholar
2021

NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction

NeurIPS 2021spotlight

We present a novel neural surface reconstruction method, called NeuS, for reconstructing objects and scenes with high fidelity from 2D image inputs. Existing neural surface reconstruction approaches, such as DVR [Niemeyer et al., 2020] and IDR [Yariv et al., 2020], require foreground mask as supervi…

2021

PR-Net: Preference Reasoning for Personalized Video Highlight Detection

ICCV 2021poster

Personalized video highlight detection aims to shorten a long video to interesting moments according to a user's preference, which has recently raised the community's attention. Current methods regard the user's history as holistic information to predict the user's preference but negating the inhere…

Cited by 14PDFScholar
2021

Point2Skeleton: Learning Skeletal Representations from Point Clouds

CVPR 2021poster

We introduce Point2Skeleton, an unsupervised method to learn skeletal representations from point clouds. Existing skeletonization methods are limited to tubular shapes and the stringent requirement of watertight input, while our method aims to produce more generalized skeletal representations for co…

Cited by 70PDFScholar
2021

Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained Encoders

EMNLP 2021main

Previous work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT. However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transfera…

2020

A Hybrid Underwater Manipulator System With Intuitive Muscle-Level sEMG Mapping Control

RA-L 2020

Soft-robotic manipulators, with their closed-chamber elastomeric actuators, natural water-sealing and inherent compliance, are ideal for underwater applications for compact, lightweight, and dexterous manipulation tasks. However, their low structure rigidity makes soft robots highly prone to underwa

Cited by 11SourceScholar
2020

Edge Enhanced Implicit Orientation Learning With Geometric Prior for 6D Pose Estimation

RA-L 2020

Estimating 6D poses of rigid objects from RGB images is an important but challenging task. This is especially true for textureless objects with strong symmetry, since they have only sparse visual features to be leveraged for the task and their symmetry leads to pose ambiguity. The implicit encoding

Cited by 34SourcecodeScholar
2020

Mapping in a Cycle: Sinkhorn Regularized Unsupervised Learning for Point Cloud Shapes

ECCV 2020poster

We propose an unsupervised learning framework with the pretext task of finding dense correspondences between point cloud shapes from the same category based on the cycle-consistency formulation. In order to learn discriminative pointwise features from point cloud data, we incorporate in the formulat…

2020

Occlusion-Aware Depth Estimation with Adaptive Normal Constraints

ECCV 2020poster

We present a new learning-based method for multi-frame depth estimation from a color video, which is a fundamental problem in scene understanding, robot navigation or handheld 3D reconstruction. While recent learning-based methods estimate depth at high accuracy, 3D point clouds exported from their…

2020

TANet: Towards Fully Automatic Tooth Arrangement

ECCV 2020poster

Determining optimal target tooth arrangements is a key step of treatment planning in digital orthodontics. Existing practice for specifying the target tooth arrangement involves tedious manual operations with the outcome quality depending heavily on the experience of individual specialists, leading…

Cited by 34SourcePDFScholar
2020

Unsupervised Learning of Intrinsic Structural Representation Points

CVPR 2020poster

Learning structures of 3D shapes is a fundamental problem in the field of computer graphics and geometry processing. We present a simple yet interpretable unsupervised method for learning a new structural representation in the form of 3D structure points. The 3D structure points produced by our meth…

Cited by 70PDFcodeScholar
2019

ToothNet: Automatic Tooth Instance Segmentation and Identification From Cone Beam CT Images

CVPR 2019poster

This paper proposes a method that uses deep convolutional neural networks to achieve automatic and accurate tooth instance segmentation and identification from CBCT (cone beam CT) images for digital dentistry. The core of our method is a two-stage network. In the first stage, an edge map is extracte…

Cited by 284PDFScholar
2015

Harvesting Discriminative Meta Objects With Deep CNN Features for Scene Classification

ICCV 2015poster

Recent work on scene classification still makes use of generic CNN features in a rudimentary manner. In this paper, we present a novel pipeline built upon deep CNN features to harvest discriminative visual objects and parts for scene classification. We first use a region proposal technique to genera…

Cited by 153PDFScholar