← Search

Krystian Mikolajczyk

23 accepted papers

2026

From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis

CVPR 2026

In this paper, we introduce NAS3R, a self-supervised feed-forward framework that jointly learns explicit 3D geometry and camera parameters with no ground-truth annotations and no pretrained priors. During training, NAS3R reconstructs 3D Gaussians from uncalibrated and unposed context views and rende

Cited by 0SourcecodeScholar
2025

Closed Loop Interactive Embodied Reasoning for Robot Manipulation

ICRA 2025

Embodied reasoning systems integrate robotic hardware and cognitive processes to perform complex tasks, typically in response to a natural language query about a specific physical environment. This usually involves changing the belief about the scene or physically interacting and changing the scene

Cited by 5SourceScholar
2025

Hypo3D: Exploring Hypothetical Reasoning in 3D

ICML 2025poster

The rise of vision-language foundation models marks an advancement in bridging the gap between human and machine capabilities in 3D scene reasoning. Existing 3D reasoning benchmarks assume real-time scene accessibility, which is impractical due to the high cost of frequent scene updates. To this end…

2025

MIMO Channel as a Neural Function: Implicit Neural Representations for Extreme CSI Compression

ICASSP 2025accepted

Acquiring and utilizing accurate channel state information (CSI) is crucial for realizing the benefits of massive multiple-input multiple-output (MIMO) technology. Current CSI feedback approaches improve precision by employing advanced deep-learning methods to learn representative CSI features for a…

Cited by 0SourceScholar
2025

No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views

ICCV 2025poster

We introduce SPFSplat, an efficient framework for 3D Gaussian splatting from sparse multi-view images, requiring no ground-truth poses during training or inference. It employs a shared feature extraction backbone, enabling simultaneous prediction of 3D Gaussian primitives and camera poses in a canon…

2024

Interactive Learning of Physical Object Properties Through Robot Manipulation and Database of Object Measurements

IROS 2024poster

This work presents a framework for automatically extracting physical object properties, such as material composition, mass, volume, and stiffness, through robot manipulation and a database of object measurements. The framework involves exploratory action selection to maximize learning about objects…

Cited by 2SourcecodeScholar
2024

OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images

NeurIPS 2024poster

Recent open-world 3D representation learning methods using Vision-Language Models (VLMs) to align 3D point clouds with image-text information have shown superior 3D zero-shot performance. However, CAD-rendered images for this alignment often lack realism and texture variation, compromising alignment…

2022

NinjaDesc: Content-Concealing Visual Descriptors via Adversarial Learning

CVPR 2022poster

In the light of recent analyses on privacy-concerning scene revelation from visual descriptors, we develop descriptors that conceal the input image content. In particular, we propose an adversarial learning framework for training visual descriptors that prevent image reconstruction, while maintainin…

Cited by 27PDFScholar
2021

Embodied Reasoning for Discovering Object Properties via Manipulation

ICRA 2021poster

In this paper, we present an integrated system that includes reasoning from visual and natural language inputs, action and motion planning, executing tasks by a robotic arm, manipulating objects, and discovering their properties. A vision to action module recognises the scene with objects and their…

Cited by 5SourceScholar
2020

Deep Joint Source-Channel Coding for Wireless Image Retrieval

ICASSP 2020accepted

Motivated by surveillance applications with wireless cameras or drones, we consider the problem of image retrieval over a wireless channel. Conventional systems apply lossy compression on query images to reduce the data that must be transmitted over a bandwidth and power limited wireless link. We fi…

Cited by 0SourceScholar
2020

HyNet: Learning Local Descriptor with Hybrid Similarity Measure and Triplet Loss

NeurIPS 2020poster

In this paper, we investigate how L2 normalisation affects the back-propagated descriptor gradients during training. Based on our observations, we propose HyNet, a new local descriptor that leads to state-of-the-art results in matching. HyNet introduces a hybrid similarity measure for triplet margin…

2020

SOLAR: Second-Order Loss and Attention for Image Retrieval

ECCV 2020poster

Recent works in deep-learning have shown that second-order information is beneficial in many computer-vision tasks. Second-order information can be enforced both in the spatial context and the abstract feature dimensions. In this work, we explore two second-order components. One is focused on second…

2019

Key.Net: Keypoint Detection by Handcrafted and Learned CNN Filters

ICCV 2019poster

We introduce a novel approach for keypoint detection task that combines handcrafted and learned CNN filters within a shallow multi-scale architecture. Handcrafted filters provide anchor structures for learned filters, which localize, score and rank repeatable features. Scale-space representation is…

Cited by 364PDFcodeScholar
2017

HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors

CVPR 2017poster

In this paper, we propose a novel benchmark for evaluating local image descriptors. We demonstrate that the existing datasets and evaluation protocols do not specify unambiguously all aspects of evaluation, leading to ambiguities and inconsistencies in results reported in the literature. Furthermore…

Cited by 980PDFScholar
2016

Hierarchical online domain adaptation of deformable part-based models

ICRA 2016

We propose an online domain adaptation method for the deformable part-based model (DPM). The online domain adaptation is based on a two-level hierarchical adaptation tree, which consists of instance models in the leaf nodes and a category model at the root node. Moreover, combined with a multiple ob

Cited by 15SourceScholar
2015

BOLD - Binary Online Learned Descriptor For Efficient Image Matching

CVPR 2015poster

In this paper we propose a novel approach to generate a binary descriptor optimized for each image patch independently. The approach is inspired by the linear discriminant embedding that simultaneously increases inter and decreases intra class distances. A set of discriminative and uncorrelated b…

Cited by 100SourcePDFScholar