← Search

Daniyar Turmukhambetov

15 accepted papers

2026

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

CVPR 2026

We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level and by satellite. Faithful reconstructions require good camera coverage, but ground imagery is time-consuming and hard to capture at scale for large outdoor s

Cited by 0SourcecodeScholar
2025

MVSAnywhere: Zero-Shot Multi-View Stereo

CVPR 2025poster

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision.However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs outdoor). Training a general-purpose multi-view stereo model is challenging an…

2024

Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer

ECCV 2024oral

"We address the task of estimating camera parameters from a set of images depicting a scene. Popular feature-based structure-from-motion (SfM) tools solve this task by incremental reconstruction: they repeat triangulation of sparse 3D points and registration of more camera views to the sparse point…

2023

DiffusioNeRF: Regularizing Neural Radiance Fields With Denoising Diffusion Models

CVPR 2023poster

Under good conditions, Neural Radiance Fields (NeRFs) have shown impressive results on novel view synthesis tasks. NeRFs learn a scene's color and density fields by minimizing the photometric discrepancy between training views and differentiable renderings of the scene. Once trained from a sufficien…

2023

Two-View Geometry Scoring Without Correspondences

CVPR 2023poster

Camera pose estimation for two-view geometry traditionally relies on RANSAC. Normally, a multitude of image correspondences leads to a pool of proposed hypotheses, which are then scored to find a winning model. The inlier count is generally regarded as a reliable indicator of "consensus". We examine…

2022

Map-Free Visual Relocalization: Metric Pose Relative to a Single Image

ECCV 2022poster

"Can we relocalize in a scene represented by a single reference image? Standard visual relocalization requires hundreds of images and scale calibration to build a scene-specific 3D map. In contrast, we propose Map-free Relocalization, i.e., using only one photo of a scene to enable instant, metric s…

2021

Learning to Predict Repeatability of Interest Points

ICRA 2021poster

Many robotics applications require interest points that are highly repeatable under varying viewpoints and lighting conditions. However, this requirement is very challenging as the environment changes continuously and indefinitely, leading to appearance changes of interest points with respect to tim…

Cited by 2SourceScholar
2021

Single Image Depth Prediction With Wavelet Decomposition

CVPR 2021poster

We present a novel method for predicting accurate depths from monocular images with high efficiency. This optimal efficiency is achieved by exploiting wavelet decomposition, which is integrated in a fully differentiable encoder-decoder architecture. We demonstrate that we can reconstruct high-fideli…

Cited by 82PDFcodeScholar
2020

Learning Stereo from Single Images

ECCV 2020poster

Supervised deep networks are among the best methods for finding correspondences in stereo image pairs. Like all supervised approaches, these networks require ground truth data during training. However, collecting large quantities of accurate dense correspondence data is very challenging. We propose…

2020

Predicting Visual Overlap of Images Through Interpretable Non-Metric Box Embeddings

ECCV 2020poster

To what extent are two images picturing the same 3D surfaces? Even when this is a known scene, the answer typically requires an expensive search across scale space, with matching and geometric verification of large sets of local features. This expense is further multiplied when a query image is eval…

2020

Single-Image Depth Prediction Makes Feature Matching Easier

ECCV 2020poster

Good local features improve the robustness of many 3D re-localization and multi-view reconstruction pipelines. The problem is that viewing angle and distance severely impact the recognizability of a local feature. Attempts to improve appearance invariance by choosing better local feature points or b…

2017

Harmonic Networks: Deep Translation and Rotation Equivariance

CVPR 2017poster

Translating or rotating an input image should not affect the results of many computer vision tasks. Convolutional neural networks (CNNs) are already translation equivariant: input image translations produce proportionate feature map translations. This is not the case for rotations. Global rotation e…

Cited by 867PDFScholar
2017

Interpretable Transformations With Encoder-Decoder Networks

ICCV 2017poster

Deep feature spaces have the capacity to encode complex transformations of their input data. However, understanding the relative feature-space relationship between two transformed encoded images is difficult. For instance, what is the relative feature space relationship between two rotated images? W…

Cited by 113PDFScholar
2015

Modeling Object Appearance Using Context-Conditioned Component Analysis

CVPR 2015poster

Subspace models have been very successful at modeling the appearance of structured image datasets when the visual objects have been aligned in the images (e.g., faces). Even with extensions that allow for global transformations or dense warps of the image, the set of visual objects whose appearance…

Cited by 8SourcePDFScholar