← Search

Prune Truong

11 accepted papers

2026

AnyUp: Universal Feature Upsampling

ICLR 2026oral

We introduce AnyUp, a method for feature upsampling that can be applied to any vision feature at any resolution, without encoder-specific training. Existing learning-based upsamplers for features like DINO or CLIP need to be re-trained for every feature extractor and thus do not generalize to differ…

Cited by 0SourcecodeScholar
2026

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

CVPR 2026

The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present a controllable, direct end-to-end method for upgrading a conventional video to a binocular one. Our approach, based on (conditional) latent diffusion, avoids artifacts due to explicit dep

Cited by 0SourcecodeScholar
2026

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

ICLR 2026oral

The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could obtain a formidable 3D scene generator if one were able to combine the power of a modern latent text-to-video model as…

Cited by 0SourcecodeScholar
2025

One2Any: One-Reference 6D Pose Estimation for Any Object

CVPR 2025poster

6D object pose estimation remains challenging for many applications due to dependencies on complete 3D models, multi-view images, or training limited to specific object categories. These requirements make generalization to novel objects difficult for which neither 3D models nor multi-view images may…

2023

SPARF: Neural Radiance Fields From Sparse and Noisy Poses

CVPR 2023highlight

Neural Radiance Field (NeRF) has recently emerged as a powerful representation to synthesize photorealistic novel views. While showing impressive performance, it relies on the availability of dense input views with highly accurate camera poses, thus limiting its application in real-world scenarios.…

2022

Probabilistic Warp Consistency for Weakly-Supervised Semantic Correspondences

CVPR 2022poster

We propose Probabilistic Warp Consistency, a weakly-supervised learning objective for semantic matching. Our approach directly supervises the dense matching scores predicted by the network, encoded as a conditional probability distribution. We first construct an image triplet by applying a known war…

Cited by 38PDFcodeScholar
2021

Learning Accurate Dense Correspondences and When To Trust Them

CVPR 2021poster

Establishing dense correspondences between a pair of images is an important and general problem. However, dense flow estimation is often inaccurate in the case of large displacements or homogeneous regions. For most applications and down-stream tasks, such as pose estimation, image manipulation, or…

Cited by 149PDFcodeScholar
2021

Warp Consistency for Unsupervised Learning of Dense Correspondences

ICCV 2021poster

The key challenge in learning dense correspondences lies in the lack of ground-truth matches for real image pairs. While photometric consistency losses provide unsupervised alternatives, they struggle with large appearance changes, which are ubiquitous in geometric and semantic matching tasks. Moreo…

Cited by 51PDFcodeScholar
2020

GLU-Net: Global-Local Universal Network for Dense Flow and Correspondences

CVPR 2020oral

Establishing dense correspondences between a pair of images is an important and general problem, covering geometric matching, optical flow and semantic correspondences. While these applications share fundamental challenges, such as large displacements, pixel-accuracy, and appearance changes, they ar…

Cited by 228PDFcodeScholar
2020

GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural Network

NeurIPS 2020poster

The feature correlation layer serves as a key neural network module in numerous computer vision problems that involve dense correspondences between image pairs. It predicts a correspondence volume by evaluating dense scalar products between feature vectors extracted from pairs of locations in two im…

2019

GLAMpoints: Greedily Learned Accurate Match Points

ICCV 2019accepted

We introduce a novel CNN-based feature point detector - Greedily Learned Accurate Match Points (GLAMpoints) - learned in a semi-supervised manner. Our detector extracts repeatable, stable interest points with a dense coverage, specifically designed to maximize the correct matching in a specific doma…

Cited by 86SourcePDFScholar