← Search

Nikita Araslanov

12 accepted papers

2026

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

CVPR 2026

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the pixel level. Existing frameworks either train on image-base

Cited by 0SourceScholar
2026

Scene-Centric Unsupervised Video Panoptic Segmentation

CVPR 2026

Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focuse

Cited by 0SourceScholar
2025

Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction

ICCV 2025poster

Traditional SLAM systems, which rely on bundle adjustment, struggle with the highly dynamic scenes commonly found in casual videos. Such videos entangle the motion of dynamic elements, undermining the assumption of static environments required by traditional systems. Existing techniques either filte…

Cited by 0SourcePDFScholar
2025

It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data

CVPR 2025poster

The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar. This suggests that as foundation models mature, it may become possible to match…

2025

Scene-Centric Unsupervised Panoptic Segmentation

CVPR 2025highlight

Unsupervised panoptic segmentation aims to partition an image into semantically meaningful regions and distinct object instances without training on manually annotated data. In contrast to prior work on unsupervised panoptic scene understanding, we eliminate the need for object-centric training data…

2024

An Analytical Solution to Gauss-Newton Loss for Direct Image Alignment

ICLR 2024oral

Direct image alignment is a widely used technique for relative 6DoF pose estimation between two images, but its accuracy strongly depends on pose initialization. Therefore, recent end-to-end frameworks increase the convergence basin of the learned feature descriptors with special training objectives…

Cited by 0SourcePDFScholar
2024

Flattening the Parent Bias: Hierarchical Semantic Segmentation in the Poincare Ball

CVPR 2024poster

Hierarchy is a natural representation of semantic taxonomies including the ones routinely used in image segmentation. Indeed recent work on semantic segmentation reports improved accuracy from supervised training leveraging hierarchical label structures. Encouraged by these results we revisit the fu…