← Search

Kevis-Kokitsi Maninis

13 accepted papers

2026

Benchmarking Open-ended Segmentation

ICLR 2026poster

Open-ended segmentation requires models capable of generating free-form descriptions of previously unseen concepts and regions. Despite advancements in model development, current evaluation protocols for open-ended segmentation tasks fail to capture the true semantic accuracy of the generated descri…

Cited by 0SourceScholar
2026

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

CVPR 2026

Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch r

Cited by 0SourcecodeScholar
2025

TIPS: Text-Image Pretraining with Spatial awareness

ICLR 2025poster

While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense visio…

2024

OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects

ECCV 2024oral

"We propose OmniNOCS, a large-scale monocular dataset with 3D Normalized Object Coordinate Space (NOCS) maps, object masks, and 3D bounding box annotations for indoor and outdoor scenes. OmniNOCS has 20 times more object classes and 200 times more instances than existing NOCS datasets (NOCS-Real275,…

2024

Probing the 3D Awareness of Visual Foundation Models

CVPR 2024poster

Recent advances in large-scale pretraining have yielded visual foundation models with strong capabilities. Not only can recent models generalize to arbitrary images for their training task their intermediate representations are useful for other visual tasks such as detection and segmentation. Given…

2023

CAD-Estate: Large-scale CAD Model Annotation in RGB Videos

ICCV 2023poster

We propose a method for annotating videos of complex multi-object scenes with a globally-consistent 3D representation of the objects. We annotate each object with a CAD model from a database, and place it in the 3D coordinate frame of the scene with a 9-DoF pose transformation. Our method is semi-au…

Cited by 7PDFcodeScholar
2023

Estimating Generic 3D Room Structures from 2D Annotations

NeurIPS 2023poster

Indoor rooms are among the most common use cases in 3D scene understanding. Current state-of-the-art methods for this task are driven by large annotated datasets. Room layouts are especially important, consisting of structural elements in 3D, such as wall, floor, and ceiling. However, they are diffi…

2023

NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations

NeurIPS 2023poster

Recent advances in neural reconstruction enable high-quality 3D object reconstruction from casually captured image collections. Current techniques mostly analyze their progress on relatively simple image collections where SfM techniques can provide ground-truth (GT) camera poses. We note that SfM te…

2022

RayTran: 3D Pose Estimation and Shape Reconstruction of Multiple Objects from Videos with Ray-Traced Transformers

ECCV 2022poster

"We propose a transformer-based neural network architecture for multi-object 3D reconstruction from RGB videos. It relies on two alternative ways to represent its knowledge: as a global 3D grid of features and an array of view-specific 2D grids. We progressively exchange information between the two…

2018

Automatic Tool Landmark Detection for Stereo Vision in Robot-Assisted Retinal Surgery

RA-L 2018

Computer vision and robotics are being increasingly applied in medical interventions. Especially in interventions where extreme precision is required, they could make a difference. One such application is robot-assisted retinal microsurgery. In recent works, such interventions are conducted under a

Cited by 49SourceScholar
2018

Deep Extreme Cut: From Extreme Points to Object Segmentation

CVPR 2018poster

This paper explores the use of extreme points in an object (left-most, right-most, top, bottom pixels) as input to obtain precise object segmentation for images and videos. We do so by adding an extra channel to the image in the input of a convolutional neural network (CNN), which contains a Gaussia…

Cited by 545SourcePDFScholar
2017

One-Shot Video Object Segmentation

CVPR 2017poster

This paper tackles the task of semi-supervised video object segmentation, i.e., the separation of an object from the background in a video, given the mask of the first frame. We present One-Shot Video Object Segmentation (OSVOS), based on a fully-convolutional neural network architecture that is abl…

Cited by 1167PDFScholar