← Search

Nikolaos Gkanatsios

9 accepted papers

2024

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

CoRL 2024poster

Diffusion policies are conditional diffusion models that learn robot action distributions conditioned on the robot and environment state. They have recently shown to outperform both deterministic and alternative action distribution learning formulations. 3D robot policies use 3D scene feature repres…

Cited by 117SourceScholar
2024

Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous and Instruction-guided Driving

CVPR 2024poster

Diffusion models excel at modeling complex and multimodal trajectory distributions for decision-making and control. Reward-gradient guided denoising has been recently proposed to generate trajectories that maximize both a differentiable reward function and the likelihood under the data distribution…

Cited by 6SourcePDFScholar
2024

ODIN: A Single Model for 2D and 3D Segmentation

CVPR 2024highlight

State-of-the-art models on contemporary 3D segmentation benchmarks like ScanNet consume and label dataset-provided 3D point clouds obtained through post processing of sensed multiview RGB-D images. They are typically trained in-domain forego large-scale 2D pre-training and outperform alternatives th…

2023

Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

CoRL 2023poster

3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically demands high-resolution 3D feature grids that are computationa…

Cited by 71SourcecodeScholar
2023

Analogy-Forming Transformers for Few-Shot 3D Parsing

ICLR 2023poster

We present Analogical Networks, a model that segments 3D object scenes with analogical reasoning: instead of mapping a scene to part segments directly, our model first retrieves related scenes from memory and their corresponding part structures, and then predicts analogous part structures in the inp…

Cited by 5SourcePDFScholar
2023

ChainedDiffuser: Unifying Trajectory Diffusion and Keypose Prediction for Robotic Manipulation

CoRL 2023poster

We present ChainedDiffuser, a policy architecture that unifies action keypose prediction and trajectory diffusion generation for learning robot manipulation from demonstrations. Our main innovation is to use a global transformer-based action predictor to predict actions at keyframes, a task that req…

Cited by 88SourceScholar
2023

Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement

RSS 2023poster

Language is compositional; an instruction can express multiple relation constraints to hold among objects in a scene that a robot is tasked to rearrange. Our focus in this work is an instructable scene-rearranging framework that generalizes to longer instructions and to spatial concept compositions…

2022

Bottom Up Top down Detection Transformers for Language Grounding in Images and Point Clouds

ECCV 2022poster

"Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector. This is limiting because an utterance may refer to visual entities at various levels of granularity, such as the chair, the…

2021

Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection

ICCV 2021poster

Scene Graph Generators (SGGs) are models that, given an image, build a directed graph where each edge represents a predicted subject predicate object triplet. Most SGGs silently exploit datasets' bias on relationships' context, i.e. its subject and object, to improve recall and neglect spatial and v…

Cited by 13PDFcodeScholar