← Search

Anton Konushin

21 accepted papers

2026

TUN3D: Towards Real-World Scene Understanding from Unposed Images

ICRA 2026poster

Layout estimation and 3D object detection are two fundamental tasks in indoor scene understanding. When combined, they enable the creation of a compact yet semantically rich spatial representation of a scene. Existing approaches typically rely on point cloud input, which poses a major limitation sin…

2026

Visual Implicit Geometry Transformer for Autonomous Driving

IJCAI 2026

We introduce the Visual Implicit Geometry Transformer (ViGT), an autonomous driving geometric model that estimates continuous 3D occupancy fields from surround-view camera rigs. ViGT represents a step towards foundational geometric models for autonomous driving, prioritizing scalability, architectur

Cited by 0Scholar
2026

Zoo3D: Zero-Shot 3D Object Detection at Scene Level

CVPR 2026

3D object detection is fundamental for spatial understanding. Real-world environments demand models capable of recognizing diverse, previously unseen objects, which remains a major limitation of closed-set methods. Existing open-vocabulary 3D detectors relax annotation requirements but still depend

Cited by 0SourcecodeScholar
2026

cadrille: Multi-modal CAD Reconstruction with Reinforcement Learning

ICLR 2026oral

Computer-Aided Design (CAD) plays a central role in engineering and manufacturing, making it possible to create precise and editable 3D models. Using a variety of sensor or user-provided data as inputs for CAD reconstruction can democratize access to design applications. However, most existing metho…

Cited by 0SourcecodeScholar
2025

A3D: Does Diffusion Dream about 3D Alignment?

ICLR 2025poster

We tackle the problem of text-driven 3D generation from a geometry alignment perspective. Given a set of text prompts, we aim to generate a collection of objects with semantically corresponding parts aligned across them. Recent methods based on Score Distillation have succeeded in distilling the kno…

Cited by 0SourcePDFScholar
2025

DepthART: Monocular Depth Estimation as Autoregressive Refinement Task

IJCAI 2025

Monocular depth estimation has seen significant advances through discriminative approaches, yet their performance remains constrained by the limitations of training datasets. While generative approaches have addressed this challenge by leveraging priors from internet-scale datasets, with recent stud

Cited by 0SourcePDFScholar
2025

UniDet3D: Multi-dataset Indoor 3D Object Detection

AAAI 2025technical

Growing customer demand for smart solutions in robotics and augmented reality has attracted considerable attention to 3D object detection from point clouds. Yet, existing indoor datasets taken individually are too small and insufficiently diverse to train a powerful and general 3D object detection m…

2024

OneFormer3D: One Transformer for Unified Point Cloud Segmentation

CVPR 2024poster

Semantic instance and panoptic segmentation of 3D point clouds have been addressed using task-specific models of distinct design. Thereby the similarity of all segmentation tasks and the implicit relationship between them have not been utilized effectively. This paper presents a unified simple and e…

Cited by 76SourcePDFScholar
2024

RClicks: Realistic Click Simulation for Benchmarking Interactive Segmentation

NeurIPS 2024poster

The emergence of Segment Anything (SAM) sparked research interest in the field of interactive segmentation, especially in the context of image editing tasks and speeding up data annotation. Unlike common semantic segmentation, interactive segmentation methods allow users to directly influence their…

2024

TETRIS: Towards Exploring the Robustness of Interactive Segmentation

AAAI 2024technical

Interactive segmentation methods rely on user inputs to iteratively update the selection mask. A click specifying the object of interest is arguably the most simple and intuitive interaction type, and thereby the most common choice for interactive segmentation. However, user clicking patterns in the…

Cited by 2SourcePDFScholar
2023

Independent Component Alignment for Multi-Task Learning

CVPR 2023poster

In a multi-task learning (MTL) setting, a single model is trained to tackle a diverse set of tasks jointly. Despite rapid progress in the field, MTL remains challenging due to optimization issues such as conflicting and dominating gradients. In this work, we propose using a condition number of a lin…

2022

FCAF3D: Fully Convolutional Anchor-Free 3D Object Detection

ECCV 2022poster

"Recently, promising applications in robotics and augmented reality have attracted considerable attention to 3D object detection from point clouds. In this paper, we present FCAF3D -- a first-in-class fully convolutional anchor-free indoor 3D object detection method. It is a simple yet effective met…

2022

Single-Stage 3D Geometry-Preserving Depth Estimation Model Training on Dataset Mixtures With Uncalibrated Stereo Data

CVPR 2022poster

Nowadays, robotics, AR, and 3D modeling applications attract considerable attention to single-view depth estimation (SVDE) as it allows estimating scene geometry from a single RGB image. Recent works have demonstrated that the accuracy of an SVDE method hugely depends on the diversity and volume of…

Cited by 7PDFScholar
2021

Decoder Modulation for Indoor Depth Completion

IROS 2021poster

Depth completion recovers a dense depth map from sensor measurements. Current methods are mostly tailored for very sparse depth measurements from LiDARs in outdoor settings, while for indoor scenes Time-of-Flight (ToF) or structured light sensors are mostly used. These sensors provide semi-dense map…

Cited by 52SourceScholar
2020

F-BRS: Rethinking Backpropagating Refinement for Interactive Segmentation

CVPR 2020oral

Deep neural networks have become a mainstream approach to interactive segmentation. As we show in our experiments, while for some images a trained network provides accurate segmentation result with just a few clicks, for some unknown objects it cannot achieve satisfactory result even with a large am…

Cited by 267PDFcodeScholar
2019

DISCOMAN: Dataset of Indoor SCenes for Odometry, Mapping And Navigation

IROS 2019poster

We present a novel dataset for training and benchmarking semantic SLAM methods. The dataset consists of 200 long sequences, each one containing 3000-5000 data frames. We generate the sequences using realistic home layouts. For that we sample trajectories that simulate motions of a simple home robot,…

Cited by 25SourceScholar
2019

Double Refinement Network for Efficient Monocular Depth Estimation

IROS 2019poster

Monocular depth estimation is the task of obtaining a measure of distance for each pixel using a single image. It is an important problem in computer vision and is usually solved using neural networks. Though recent works in this area have shown significant improvement in accuracy, the state-of-the-…

Cited by 15SourceScholar