← Search

Maks Ovsjanikov

52 accepted papers

2026

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

CVPR 2026

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different modalities. For instance, the feature embedding for an RGB image and its corresponding depth map of the same scene exhibit

Cited by 0SourcecodeScholar
2026

Beyond ReLU: Bifurcation, Oversmoothing, and Topological Priors

ICML 2026spotlight

Graph Neural Networks (GNNs) learn node representations through iterative network-based message-passing. While powerful, deep GNNs suffer from oversmoothing, where node features converge to a homogeneous, non-informative state. We re-frame this problem of representational collapse from a \emph{bifur…

Cited by 0SourceScholar
2026

Dynamic Reflections: Probing Video Representations with Text Alignment

ICLR 2026poster

The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal natur…

Cited by 0SourceScholar
2026

Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication

ICML 2026poster

Graph alignment, the problem of identifying corresponding nodes across multiple graphs, is fundamental to numerous applications. Most existing unsupervised methods embed node features into latent representations to enable cross-graph comparison without ground-truth correspondences. However, these me…

Cited by 0SourceScholar
2026

PaNDaS: Learnable Shape Interpolation Modeling with Localized Control

CVPR 2026

We present PaNDaS, a novel deep learning framework for Partial Non-Rigid Deformations and interpolations of Surfaces (PaNDaS). PaNDaS learns a per-face feature field on the source mesh and fuses it with a global encoding of the target. A deformation generator predicts a Jacobian field and recovers a

Cited by 0SourceScholar
2026

PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding

CVPR 2026

Current foundation models for 3D shapes excel at global tasks (retrieval, classification) but transfer poorly to local part-level reasoning. Recent approaches leverage vision and language foundation models to directly solve dense tasks through multi-view renderings and text queries. While promising,

Cited by 0SourcecodeScholar
2026

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

ICML 2026poster

Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled data from external vision sources or synthetic…

Cited by 0SourceScholar
2026

Unfolding Generative Flows with Koopman Operators: Trajectory-Preserving Linearization

ICML 2026poster

Continuous Normalizing Flows (CNFs) enable elegant generative modeling but remain bottlenecked by their iterative nature requiring costly sampling and lacking interpretability of the intermediate states. Recent approaches accelerate sampling by straightening trajectories or distilling endpoints, yet…

Cited by 0SourceScholar
2026

Unified Brain Surface and Volume Registration

ICLR 2026poster

Accurate registration of brain MRI scans is fundamental for cross-subject analysis in neuroscientific studies. This involves aligning both the cortical surface of the brain and the interior volume. Traditional methods treat volumetric and surface-based registration separately, which often leads to i…

Cited by 0SourceScholar
2026

Unique Lives, Shared World: Learning from Single-Life Videos

CVPR 2026

We introduce the "single-life" learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual. We leverage the multiple viewpoints naturally captured within a single life to learn a visual encoder in a self-supervised manner. Our experiments dem

Cited by 0SourceScholar
2025

AtomSurf: Surface Representation for Learning on Protein Structures

ICLR 2025poster

While there has been significant progress in evaluating and comparing different representations for learning on protein data, the role of surface-based learning approaches remains not well-understood. In particular, there is a lack of direct and fair benchmark comparison between the best available s…

2025

DiffuMatch: Category-Agnostic Spectral Diffusion Priors for Robust Non-rigid Shape Matching

ICCV 2025poster

Deep functional maps have recently emerged as a powerful tool for solving non-rigid shape correspondence tasks. Methods that use this approach combine the power and flexibility of the functional map framework, with data-driven learning for improved accuracy and generality. However, most existing met…

2025

Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces

CVPR 2025poster

Recent works have shown that, when trained at scale, uni-modal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising from different representations. However, the role of 3D encoders with respect to other modalities remains unexplored. F…

Cited by 0SourcePDFScholar
2025

LACONIC: A 3D Layout Adapter for Controllable Image Creation

ICCV 2025poster

Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric structure, underlying the scene. In this paper, we propose a n…

Cited by 0SourcePDFScholar
2025

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications

ICCV 2025poster

In recent years, there has been a proliferation of spatiotemporal foundation models in different scientific disciplines. While promising, these models are often domain-specific and are only assessed within the particular applications for which they are designed. Given that many tasks can be represen…

2025

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

ICCV 2025poster

We propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP. Traditional methods for 3D keypoint detection rely heavily on ann…

Cited by 0SourcePDFScholar
2024

Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D Features

CVPR 2024poster

With the immense growth of dataset sizes and computing resources in recent years so-called foundation models have become popular in NLP and vision tasks. In this work we propose to explore foundation models for the task of keypoint detection on 3D shapes. A unique characteristic of keypoint detectio…

2024

DeBaRA: Denoising-Based 3D Room Arrangement Generation

NeurIPS 2024poster

Generating realistic and diverse layouts of furnished indoor 3D scenes unlocks multiple interactive applications impacting a wide range of industries. The inherent complexity of object interactions, the limited amount of available data and the requirement to fulfill spatial constraints all make gene…

Cited by 2SourcePDFScholar
2024

To Supervise or Not to Supervise: Understanding and Addressing the Key Challenges of Point Cloud Transfer Learning

ECCV 2024poster

"Transfer learning has long been a key factor in the advancement of many fields including 2D image analysis. Unfortunately, its applicability in 3D data processing has been relatively limited. While several approaches for point cloud transfer learning have been proposed in recent literature, with co…

Cited by 0SourcePDFScholar
2023

Affection: Learning Affective Explanations for Real-World Visual Data

CVPR 2023poster

In this work, we explore the space of emotional reactions induced by real-world images. For this, we first introduce a large-scale dataset that contains both categorical emotional reactions and free-form textual explanations for 85,007 publicly available images, analyzed by 6,283 annotators who were…

Cited by 18SourcePDFScholar
2023

Generalizable Local Feature Pre-Training for Deformable Shape Analysis

CVPR 2023highlight

Transfer learning is fundamental for addressing problems in settings with little training data. While several transfer learning approaches have been proposed in 3D, unfortunately, these solutions typically operate on an entire 3D object or even scene-level and thus, as we show, fail to generalize to…

2023

SATR: Zero-Shot Semantic Segmentation of 3D Shapes

ICCV 2023poster

We explore the task of zero-shot semantic segmentation of 3D shapes by using large-scale off-the-shelf 2D im- age recognition models. Surprisingly, we find that modern zero-shot 2D object detectors are better suited for this task than contemporary text/image similarity predictors or even zero-shot 2…

Cited by 40PDFcodeScholar
2023

Shape Non-rigid Kinematics (SNK): A Zero-Shot Method for Non-Rigid Shape Matching via Unsupervised Functional Map Regularized Reconstruction

NeurIPS 2023poster

We present Shape Non-rigid Kinematics (SNK), a novel zero-shot method for non-rigid shape matching that eliminates the need for extensive training or ground truth data.SNK operates on a single pair of shapes, and employs a reconstruction-based strategy using an encoder-decoder architecture, which de…

2023

Spatially and Spectrally Consistent Deep Functional Maps

ICCV 2023poster

Cycle consistency has long been exploited as a powerful prior for jointly optimizing maps within a collection of shapes. In this paper, we investigate its utility in the approaches of Deep Functional Maps, which are considered state-of-the-art in non-rigid shape matching. We first justify that under…

Cited by 27PDFcodeScholar
2023

TIDE: Time Derivative Diffusion for Deep Learning on Graphs

ICML 2023poster

A prominent paradigm for graph neural networks is based on the message-passing framework. In this framework, information communication is realized only between neighboring nodes. The challenge of approaches that use this paradigm is to ensure efficient and accurate long-distance communication betwee…

2023

Understanding and Improving Features Learned in Deep Functional Maps

CVPR 2023highlight

Deep functional maps have recently emerged as a successful paradigm for non-rigid 3D shape correspondence tasks. An essential step in this pipeline consists in learning feature functions that are used as constraints to solve for a functional map inside the network. However, the precise nature of the…

2023

VoroMesh: Learning Watertight Surface Meshes with Voronoi Diagrams

ICCV 2023poster

In stark contrast to the case of images, finding a concise, learnable discrete representation of 3D surfaces remains a challenge. In particular, while polygon meshes are arguably the most common surface representation used in geometry processing, their irregular and combinatorial structure often mak…

Cited by 15PDFScholar
2022

Deep Orientation-Aware Functional Maps: Tackling Symmetry Issues in Shape Matching

CVPR 2022poster

State-of-the-art fully intrinsic network for non-rigid shape matching are unable to disambiguate between shape inner symmetries. Meanwhile, recent advances in the functional map framework allow to enforce orientation preservation using a functional representation for tangent vector field transfer, t…

Cited by 52PDFcodeScholar
2022

Implicit Field Supervision for Robust Non-rigid Shape Matching

ECCV 2022poster

"Establishing a correspondence between two non-rigidly deforming shapes is one of the most fundamental problems in visual computing. Existing methods often show weak resilience when presented with challenges innate to real-world data such as noise, outliers, self-occlusion etc. On the other hand, au…

2022

Learning Multi-resolution Functional Maps with Spectral Attention for Robust Shape Matching

NeurIPS 2022accept

In this work, we present a novel non-rigid shape matching framework based on multi-resolution functional maps with spectral attention. Existing functional map learning methods all rely on the critical choice of the spectral resolution hyperparameter, which can severely affect the overall accuracy or…

2022

NCP: Neural Correspondence Prior for Effective Unsupervised Shape Matching

NeurIPS 2022accept

We present Neural Correspondence Prior (NCP), a new paradigm for computing correspondences between 3D shapes. Our approach is fully unsupervised and can lead to high quality correspondences even in challenging cases such as sparse point clouds or non-isometric meshes, where current methods fail. Our…

2022

Reduced Representation of Deformation Fields for Effective Non-rigid Shape Matching

NeurIPS 2022accept

In this work we present a novel approach for computing correspondences between non-rigid objects, by exploiting a reduced representation of deformation fields. Different from existing works that represent deformation fields by training a general-purpose neural network, we advocate for an approximati…

2021

ArtEmis: Affective Language for Visual Art

CVPR 2021poster

We present a novel large-scale dataset and accompanying machine learning models aimed at providing a detailed understanding of the interplay between visual content, its emotional effect, and explanations for the latter in language. In contrast to most existing annotation datasets in computer vision,…

Cited by 201PDFcodeScholar
2021

Fast Sinkhorn Filters: Using Matrix Scaling for Non-Rigid Shape Correspondence With Functional Maps

CVPR 2021poster

In this paper, we provide a theoretical foundation for pointwise map recovery from functional maps and highlight its relation to a range of shape correspondence methods based on spectral alignment. With this analysis in hand, we develop a novel spectral registration technique: Fast Sinkhorn Filters,…

Cited by 64PDFcodeScholar
2021

Learning Delaunay Surface Elements for Mesh Reconstruction

CVPR 2021poster

We present a method for reconstructing triangle meshes from point clouds. Existing learning-based methods for mesh reconstruction mostly generate triangles individually, making it hard to create manifold meshes. We leverage the properties of 2D Delaunay triangulations to construct a mesh from manifo…

Cited by 57PDFcodeScholar
2021

Physically-Aware Generative Network for 3D Shape Modeling

CVPR 2021poster

Shapes are often designed to satisfy structural properties and serve a particular functionality in the physical world. Unfortunately, most existing generative models focus primarily on the geometric or visual plausibility, ignoring the physical or structural constraints. To remedy this, we present a…

Cited by 30PDFScholar
2020

Correspondence learning via linearly-invariant embedding

NeurIPS 2020poster

In this paper, we propose a fully differentiable pipeline for estimating accurate dense correspondences between 3D point clouds. The proposed pipeline is an extension and a generalization of the functional maps framework. However, instead of using the Laplace-Beltrami eigenfunctions as done in virtu…

2020

Deep Geometric Functional Maps: Robust Feature Learning for Shape Correspondence

CVPR 2020oral

We present a novel learning-based approach for computing correspondences between non-rigid 3D shapes. Unlike previous methods that either require extensive training data or operate on handcrafted input descriptors and thus generalize poorly across diverse datasets, our approach is both accurate and…

Cited by 203PDFcodeScholar
2020

Intrinsic Point Cloud Interpolation via Dual Latent Space Navigation

ECCV 2020poster

We present a learning-based method for interpolating and manipulating 3D shapes represented as point clouds, that is explicitly designed to preserve intrinsic shape properties. Our approach is based on constructing a dual encoding space that enables shape synthesis and, at the same time, provides li…

2019

Isospectralization, or How to Hear Shape, Style, and Correspondence

CVPR 2019poster

The question whether one can recover the shape of a geometric object from its Laplacian spectrum ('hear the shape of the drum') is a classical problem in spectral geometry with a broad range of implications and applications. While theoretically the answer to this question is negative (there exist ex…

Cited by 65PDFScholar
2019

OperatorNet: Recovering 3D Shapes From Difference Operators

ICCV 2019poster

This paper proposes a learning-based framework for reconstructing 3D shapes from functional operators, compactly encoded as small-sized matrices. To this end we introduce a novel neural architecture, called OperatorNet, which takes as input a set of linear operators representing a shape and produces…

Cited by 18PDFcodeScholar
2017

Region-Based Correspondence Between 3D Shapes via Spatially Smooth Biclustering

ICCV 2017poster

Region-based correspondence (RBC) is a highly relevant and non-trivial computer vision problem. Given two 3D shapes, RBC seeks segments/regions on these shapes that can be reliably put in correspondence. The problem thus consists both in finding the regions and determining the correspondences betwee…

Cited by 10PDFScholar