← Search

Yuchao Dai

73 accepted papers

2026

EC-MVSNet: Enhanced Cascaded Multi-View Stereo with Cross-Scale Relevance Integration

AAAI 2026technical

Cascade-based multi-scale architectures are currently the mainstream in Multi-view Stereo (MVS), achieving a balance between computational efficiency and reconstruction accuracy. However, existing cascade MVS methods suffer from significant limitations in cross-scale information utilization, where d

Cited by 0SourcePDFScholar
2026

Learning Spatial Decay for Vision Transformers

AAAI 2026technical

Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, a

Cited by 0SourcePDFScholar
2026

RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

CVPR 2026

Humans perceive the 3D world from limited 2D observations. While recent feed-forward generalizable 3D reconstruction models can recover structures from sparse images, they typically represent only observed regions, leaving unseen geometry unmodeled. This raises a fundamental question: Can we infer c

Cited by 0SourceScholar
2026

SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

CVPR 2026

Adversarial robustness of BEV 3D object detectors is critical for autonomous driving (AD). Existing invasive attacks require altering the target vehicle itself (e.g. attaching patches), making them unrealistic and impractical for real-world evaluation. While non-invasive attacks that place adversari

Cited by 0SourceScholar
2026

SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priors

CVPR 2026

Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost Learning-based Multi-View Stereo (MVS) methods have become the mainstream in the field, relying on the construction of cost volumes through multi-view feature similarity comput

Cited by 0SourcecodeScholar
2025

Deep Non-Rigid Structure-from-Motion Revisited: Canonicalization and Sequence Modeling

AAAI 2025technical

Non-Rigid Structure-from-Motion (NRSfM) is a classic 3D vision problem, where a 2D sequence is taken as input to estimate the corresponding 3D sequence. Recently, the deep neural networks have greatly advanced the task of NRSfM. However, existing deep NRSfM methods still have limitations in handling…

Cited by 0SourcePDFScholar
2025

Event-aided Dense and Continuous Point Tracking: Everywhere and Anytime

ICCV 2025poster

Recent point tracking methods have made great strides in recovering the trajectories of any point (especially key points) in long video sequences associated with large motions. However, the spatial and temporal granularities of point trajectories remain constrained by limited motion estimation accur…

Cited by 0SourcePDFScholar
2025

LoopRefine: Deep Camera Pose Estimation With Loop Consistency

RA-L 2025

Recently, pose estimation under sparse views (<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\leq 10$</tex-math></inline-formula>) has witnessed significant advances with the development of deep learning. Most exi

Cited by 2SourceScholar
2025

MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation

ICCV 2025poster

We present MixRI, a lightweight network that solves the CAD-based novel object pose estimation problem in RGB images. It can be instantly applied to a novel object at test time without finetuning. We design our network to meet the demands of real-world applications, emphasizing reduced memory requir…

Cited by 0SourcePDFScholar
2025

PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction

ICCV 2025poster

Wide-baseline panorama reconstruction has emerged as a highly effective and pivotal approach for not only achieving geometric reconstruction of the surrounding 3D environment, but also generating highly realistic and immersive novel views. Although existing methods have shown remarkable performance…

Cited by 0SourcePDFScholar
2024

3D Focusing-and-Matching Network for Multi-Instance Point Cloud Registration

NeurIPS 2024poster

Multi-instance point cloud registration aims to estimate the pose of all instances of a model point cloud in the whole scene. Existing methods all adopt the strategy of first obtaining the global correspondence and then clustering to obtain the pose of each instance. However, due to the cluttered an…

2024

Adaptive Feature Enhanced Multi-View Stereo With Epipolar Line Information Aggregation

RA-L 2024

Despite the promising performance achieved by the learning-based multi-view stereo (MVS) methods, the commonly used feature extractors still struggle with the perspective transformation across different viewpoints. Furthermore, existing methods generally employ a “one-to-many” strategy, computing th

Cited by 1SourceScholar
2024

Improving Audio-Visual Segmentation with Bidirectional Generation

AAAI 2024technical

The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the contribution of each modality is implicitly or explicitly mod…

2024

Improving Depth Completion via Depth Feature Upsampling

CVPR 2024poster

The encoder-decoder network (ED-Net) is a commonly employed choice for existing depth completion methods but its working mechanism is ambiguous. In this paper we visualize the internal feature maps to analyze how the network densifies the input sparse depth. We find that the encoder feature of ED-Ne…

2024

Non-Rigid Structure-from-Motion: Temporally-Smooth Procrustean Alignment and Spatially-Variant Deformation Modeling

CVPR 2024poster

Even though Non-rigid Structure-from-Motion (NRSfM) has been extensively studied and great progress has been made there are still key challenges that hinder their broad real-world applications: 1) the inherent motion/rotation ambiguity requires either explicit camera motion recovery with extra const…

Cited by 1SourcePDFScholar
2024

PaReNeRF: Toward Fast Large-scale Dynamic NeRF with Patch-based Reference

CVPR 2024poster

With photo-realistic image generation Neural Radiance Field (NeRF) is widely used for large-scale dynamic scene reconstruction as autonomous driving simulator. However large-scale scene reconstruction still suffers from extremely long training time and rendering time. Low-resolution (LR) rendering c…

Cited by 1SourcePDFScholar
2024

Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking Cameras

NeurIPS 2024poster

The spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network ar…

Cited by 1SourcePDFScholar
2023

Fine-Grained Audible Video Description

CVPR 2023poster

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of each object, the actions of moving objects, and the sounds i…

2023

Forward Flow for Novel View Synthesis of Dynamic Scenes

ICCV 2023oral

This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canoni…

Cited by 48PDFcodeScholar
2023

Joint Appearance and Motion Learning for Efficient Rolling Shutter Correction

CVPR 2023poster

Rolling shutter correction (RSC) is becoming increasingly popular for RS cameras that are widely used in commercial and industrial applications. Despite the promising performance, existing RSC methods typically employ a two-stage network structure that ignores intrinsic information interactions and…

2023

LRRU: Long-short Range Recurrent Updating Networks for Depth Completion

ICCV 2023poster

Existing deep learning-based depth completion methods generally employ massive stacked layers to predict the dense depth map from sparse input data. Although such approaches greatly advance this task, their accompanied huge computational complexity hinders their practical applications. To accomplish…

Cited by 55PDFcodeScholar
2023

Masked Representation Learning for Domain Generalized Stereo Matching

CVPR 2023poster

Recently, many deep stereo matching methods have begun to focus on cross-domain performance, achieving impressive achievements. However, these methods did not deal with the significant volatility of generalization performance among different training epochs. Inspired by masked representation learnin…

Cited by 29SourcePDFScholar
2023

Modeling the Distributional Uncertainty for Salient Object Detection Models

CVPR 2023poster

Most of the existing salient object detection (SOD) models focus on improving the overall model performance, without explicitly explaining the discrepancy between the training and testing distributions. In this paper, we investigate a particular type of epistemic uncertainty, namely distributional u…

Cited by 23SourcePDFScholar
2023

Multimodal Variational Auto-encoder based Audio-Visual Segmentation

ICCV 2023poster

We propose an Explicit Conditional Multimodal Variational Auto-Encoder (ECMVAE) for audio-visual segmentation (AVS), aiming to segment sound sources in the video sequence. Existing AVS methods focus on implicit feature fusion strategies, where models are trained to fit the discrete samples in the da…

Cited by 42PDFcodeScholar
2023

RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow Estimation

ICCV 2023poster

Recently, the RGB images and point clouds fusion methods have been proposed to jointly estimate 2D optical flow and 3D scene flow. However, as both conventional RGB cameras and LiDAR sensors adopt a frame-based data acquisition mechanism, their performance is limited by the fixed low sampling rates,…

Cited by 24PDFcodeScholar
2023

Toeplitz Neural Network for Sequence Modeling

ICLR 2023top-25%

Sequence modeling has important applications in natural language processing and computer vision. Recently, the transformer-based models have shown strong performance on various sequence modeling tasks, which rely on attention to capture pairwise token relations, and position embedding to inject posi…

2022

Context-Aware Video Reconstruction for Rolling Shutter Cameras

CVPR 2022poster

With the ubiquity of rolling shutter (RS) cameras, it is becoming increasingly attractive to recover the latent global shutter (GS) video from two consecutive RS frames, which also places a higher demand on realism. Existing solutions, using deep neural networks or optimization, achieve promising pe…

Cited by 30PDFcodeScholar
2022

Efficient Spatial-Temporal Information Fusion for LiDAR-Based 3D Moving Object Segmentation

IROS 2022poster

Accurate moving object segmentation is an es-sential task for autonomous driving. It can provide effective information for many downstream tasks, such as collision avoidance, path planning, and static map construction. How to effectively exploit the spatial-temporal information is a critical questio…

Cited by 86SourcecodeScholar
2022

End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud Registration

AAAI 2022technical

Even though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-ba…

Cited by 33SourcePDFScholar
2022

PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching

ECCV 2022poster

"Existing deep learning based stereo matching methods either focus on achieving optimal performances on the target dataset while with poor generalization for other datasets or focus on handling the cross-domain generalization by suppressing the domain sensitive features which results in a significan…

Cited by 98SourcePDFScholar
2021

Complementary Patch for Weakly Supervised Semantic Segmentation

ICCV 2021poster

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has been greatly advanced by exploiting the outputs of Class Activation Map (CAM) to generate the pseudo labels for semantic segmentation. However, CAM merely discovers seeds from a small number of regions, which may be insuf…

Cited by 172PDFcodeScholar
2021

Deep Two-View Structure-From-Motion Revisited

CVPR 2021poster

Two-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem in ways that are fundamentally ill-posed, relying on training data to overcome the inherent difficulties. In contrast, we propose a return to th…

Cited by 63PDFcodeScholar
2021

Inverting a Rolling Shutter Camera: Bring Rolling Shutter Images to High Framerate Global Shutter Video

ICCV 2021poster

Rolling shutter (RS) images can be viewed as the result of the row-wise combination of global shutter (GS) images captured by a virtual moving GS camera over the period of camera readout time. The RS effect brings tremendous difficulties for the downstream applications. In this paper, we propose to…

Cited by 43PDFScholar
2021

Neural Image Compression via Attentional Multi-Scale Back Projection and Frequency Decomposition

ICCV 2021poster

In recent years, neural image compression emerges as a rapidly developing topic in computer vision, where the state-of-the-art approaches now exhibit superior compression performance than their conventional counterparts. Despite the great progress, current methods still have limitations in preservin…

Cited by 89PDFScholar
2021

PR-RRN: Pairwise-Regularized Residual-Recursive Networks for Non-Rigid Structure-From-Motion

ICCV 2021poster

We propose PR-RRN, a novel neural-network based method for Non-rigid Structure-from-Motion (NRSfM). PR-RRN consists of Residual-Recursive Networks (RRN) and two extra regularization losses. RRN is designed to effectively recover 3D shape and camera from 2D keypoints with novel residual-recursive str…

Cited by 13PDFScholar
2021

RGB-D Saliency Detection via Cascaded Mutual Information Minimization

ICCV 2021poster

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to explicitly model the multi-modal information between RGB im…

Cited by 139PDFcodeScholar
2021

Simultaneously Localize, Segment and Rank the Camouflaged Objects

CVPR 2021poster

Camouflage is a key defence mechanism across species that is critical to survival. Common camouflage include background matching, imitating the color and pattern of the environment, and disruptive coloration, disguising body outlines. Camouflaged object detection (COD) aims to segment camouflaged ob…

Cited by 472PDFcodeScholar
2021

UASNet: Uncertainty Adaptive Sampling Network for Deep Stereo Matching

ICCV 2021poster

Recent studies have shown that cascade cost volume can play a vital role in deep stereo matching to achieve high resolution depth map with efficient hardware usage. However, how to construct good cascade volume as well as effective sampling for them are still under in-depth study. Previous cascade-b…

Cited by 31PDFScholar
2021

Uncertainty-Aware Joint Salient Object and Camouflaged Object Detection

CVPR 2021poster

Visual salient object detection (SOD) aims at finding the salient object(s) that attract human attention, while camouflaged object detection (COD) on the contrary intends to discover the camouflaged object(s) that hidden in the surrounding. In this paper, we propose a paradigm of leveraging the cont…

Cited by 307PDFcodeScholar
2020

Channel Attention Based Iterative Residual Learning for Depth Map Super-Resolution

CVPR 2020poster

Despite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what…

Cited by 103PDFScholar
2020

Displacement-Invariant Matching Cost Learning for Accurate Optical Flow Estimation

NeurIPS 2020poster

Learning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly…

2020

Hierarchical Neural Architecture Search for Deep Stereo Matching

NeurIPS 2020poster

To reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the netw…

2020

Joint 3D Instance Segmentation and Object Detection for Autonomous Driving

CVPR 2020poster

Currently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle…

Cited by 132PDFScholar
2020

UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders

CVPR 2020oral

In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following…

Cited by 420PDFScholar
2020

Weakly-Supervised Salient Object Detection via Scribble Annotations

CVPR 2020poster

Compared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 1 2 seconds to label one image. However, using scribble labels to learn salient object detection has not been explored. In this paper, we propose a weakly-supervised salient object detec…

Cited by 335PDFcodeScholar
2019

ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving

CVPR 2019poster

Autonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision communit…

Cited by 224PDFcodeScholar
2019

Bringing a Blurry Frame Alive at High Frame-Rate With an Event Camera

CVPR 2019oral

Event-based cameras can measure intensity changes (called 'events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the active pixel sensor (APS), the event camera allows simultaneous output of the intensity frames. However, the output images are captured…

Cited by 312PDFScholar
2019

Phase-Only Image Based Kernel Estimation for Single Image Blind Deblurring

CVPR 2019poster

The image motion blurring process is generally modelled as the convolution of a blur kernel with a latent image. Therefore, the estimation of the blur kernel is essentially important for blind image deblurring. Unlike existing approaches which focus on approaching the problem by enforcing various pr…

Cited by 80PDFScholar
2019

Unsupervised Deep Epipolar Flow for Stationary or Dynamic Scenes

CVPR 2019poster

Unsupervised deep learning for optical flow computation has achieved promising results. Most existing deep-net based methods rely on image brightness consistency and local smoothness constraint to train the networks. Their performance degrades at regions where repetitive textures or occlusions occ…

Cited by 83PDFScholar
2018

Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling Perspective

CVPR 2018poster

The success of current deep saliency detection methods heavily depends on the availability of large-scale supervision in the form of per-pixel labeling. Such supervision, while labor-intensive and not always possible, tends to hinder the generalization ability of the learned models. By contrast, tra…

Cited by 219SourcePDFScholar
2018

Scalable Dense Non-Rigid Structure-From-Motion: A Grassmannian Perspective

CVPR 2018poster

This paper addresses the task of dense non-rigid structure-from-motion (NRSfM) using multiple images. State-of-the-art methods to this problem are often hurdled by scalability, expensive computations, and noisy measurements. Further, recent methods to NRSfM usually either assume a small number of sp…

Cited by 58SourcePDFScholar
2017

"Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction From Multiple Perspective Views

ICCV 2017poster

Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of "…

Cited by 20PDFScholar
2017

Monocular Dense 3D Reconstruction of a Complex Dynamic Scene From Two Perspective Frames

ICCV 2017poster

This paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel oversegmentation to the image, we model a generically dynamic (hence non-rigid) scene with a piecewise planar and rigid approximation. In this way,…

Cited by 74PDFScholar
2016

Robust Optical Flow Estimation of Double-Layer Images Under Transparency or Reflection

CVPR 2016poster

This paper deals with a challenging, frequently encountered, yet not properly investigated problem in two-frame optical flow estimation. That is, the input frames are compounds of two imaging layers -- one desired background layer of the scene, and one distracting, possibly moving layer due to trans…

Cited by 62PDFScholar
2015

Depth and Surface Normal Estimation From Monocular Images Using Regression on Deep Features and Hierarchical CRFs

CVPR 2015poster

Predicting the depth (or surface normal) of a scene from single monocular color images is a challenging task. This paper tackles this challenging and essentially under-determined problem by regression on deep convolutional neural network (DCNN) features, combined with a post-processing refining step…

Cited by 753SourcePDFScholar