← Search

Kostas Daniilidis

110 accepted papers

2026

Active Next-Best-View Optimization for Risk-Averse Path Planning

ICRA 2026poster

Safe navigation in uncertain environments requires planning methods that integrate risk aversion with active perception. In this work, we present a unified frame- work that refines a coarse reference path by construct- ing tail-sensitive risk maps from Average Value-at-Risk statistics on an online-u…

2026

Recurrent Equivariant Constraint Modulation: Learning Per-Layer Symmetry Relaxation from Data

ICML 2026spotlight

Equivariant neural networks exploit underlying task symmetries to improve generalization, but strict equivariance constraints can induce more complex optimization dynamics that can hinder learning. Prior work addresses these limitations by relaxing strict equivariance during training, but typically …

Cited by 0SourceScholar
2025

A Scalable, Causal, and Energy Efficient Framework for Neural Decoding with Spiking Neural Networks

NeurIPS 2025poster

Brain-computer interfaces (BCIs) promise to enable vital functions, such as speech and prosthetic control, for individuals with neuromotor impairments. Central to their success are neural decoders, models that map neural activity to intended behavior. Current learning-based decoding approaches fall…

Cited by 0SourceScholar
2025

Continuous-Time Human Motion Field from Event Cameras

ICCV 2025poster

This paper addresses the challenges of estimating a continuous-time field from a stream of events. Existing Human Mesh Recovery (HMR) methods rely predominantly on frame-based approaches, which are prone to aliasing and inaccuracies due to limited temporal resolution and motion blur. In this work, w…

Cited by 0SourcePDFScholar
2025

DIMO: Diverse 3D Motion Generation for Arbitrary Objects

ICCV 2025poster

We present DIMO, a generative approach capable of generating diverse 3D motions for arbitrary objects from a single image. The core idea of our work is to leverage the rich priors in well-trained video models to extract the common motion patterns and then embed them into a shared low-dimensional lat…

Cited by 0SourcePDFScholar
2025

ETAP: Event-based Tracking of Any Point

CVPR 2025highlight

Tracking any point (TAP) recently shifted the motion estimation paradigm from focusing on individual salient points with local templates to tracking arbitrary points with global image contexts. However, while research has mostly focused on driving the accuracy of models in nominal settings, addressi…

2025

EqNIO: Subequivariant Neural Inertial Odometry

ICLR 2025poster

Neural network-based odometry using accelerometer and gyroscope readings from a single IMU can achieve robust, and low-drift localization capabilities, through the use of _neural displacement priors (NDPs)_. These priors learn to produce denoised displacement measurements but need to ignore data var…

2025

MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

CVPR 2025highlight

We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundati…

2025

Multimodal LLM Guided Exploration and Active Mapping using Fisher Information

ICCV 2025poster

We present an active mapping system that could plan for long-horizon exploration goals and short-term actions with a 3D Gaussian Splatting (3DGS) representation. Existing methods either did not take advantage of recent developments in multimodal Large Language Models (LLM) or did not consider challe…

Cited by 0SourcePDFScholar
2025

Neural Inertial Odometry from Lie Events

RSS 2025poster

Neural displacement priors (NDPs) can reduce the drift in inertial odometry and provide uncertainty estimates that can be readily fused with off-the-shelf filters. However, they fail to generalize to different IMU sampling rates and trajectory profiles, which limits their robustness in diverse setti…

Cited by 0PDFcodeScholar
2025

Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting

ICRA 2025

We propose a framework for active next best view and touch selection for robotic manipulators using 3D Gaussian Splatting (3DGS). 3DGS is emerging as a useful explicit 3D scene representation for robotics, as it has the ability to represent scenes in a both photorealistic and geometrically accurate

Cited by 11SourcecodeScholar
2025

PromptHMR: Promptable Human Mesh Recovery

CVPR 2025poster

Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challe…

Cited by 0SourcePDFScholar
2025

RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies

CoRL 2025oral

Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standardization, either by specifying fixed evaluation tasks and environments, or by hosting centralized "robot challenges", an…

Cited by 0SourceScholar
2024

$SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation

NeurIPS 2024poster

Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior,…

Cited by 1SourcePDFScholar
2024

DynMF: Neural Motion Factorization for Real-time Dynamic View Synthesis with 3D Gaussian Splatting

ECCV 2024poster

"Accurately and efficiently modeling dynamic scenes and motions is considered so challenging a task due to temporal dynamics and motion complexity. To address these challenges, we propose , a compact and efficient representation that decomposes a dynamic scene into a few neural trajectories. We argu…

Cited by 95SourcePDFScholar
2024

GART: Gaussian Articulated Template Models

CVPR 2024highlight

We introduce Gaussian Articulated Template Model (GART) an explicit efficient and expressive representation for non-rigid articulated subject capturing and rendering from monocular videos. GART utilizes a mixture of moving 3D Gaussians to explicitly approximate a deformable subject's geometry and ap…

Cited by 94SourcePDFScholar
2024

Improving Equivariant Model Training via Constraint Relaxation

NeurIPS 2024poster

Equivariant neural networks have been widely used in a variety of applications due to their ability to generalize well in tasks where the underlying data symmetries are known. Despite their successes, such networks can be difficult to optimize and require careful hyperparameter tuning to train succe…

2024

Motion-prior Contrast Maximization for Dense Continuous-Time Motion Estimation

ECCV 2024poster

"Current optical flow and point-tracking methods rely heavily on synthetic datasets. Event cameras are novel vision sensors with advantages in challenging visual conditions, but state-of-the-art frame-based methods cannot be easily adapted to event data due to the limitations of current event simula…

2024

Neural decoding from stereotactic EEG: accounting for electrode variability across subjects

NeurIPS 2024poster

Deep learning based neural decoding from stereotactic electroencephalography (sEEG) would likely benefit from scaling up both dataset and model size. To achieve this, combining data across multiple subjects is crucial. However, in sEEG cohorts, each subject has a variable number of electrodes placed…

Cited by 0SourcePDFScholar
2024

Scalable Networked Feature Selection with Randomized Algorithm for Robot Navigation

IROS 2024poster

We address the problem of sparse selection of visual features for localizing a team of robots navigating in an unknown environment, where robots can exchange relative position measurements with neighbors. We select a set of the most informative features by anticipating their importance in robots loc…

Cited by 1SourceScholar
2024

TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos

ECCV 2024poster

"We propose TRAM, a two-stage method to reconstruct a human’s global trajectory and motion from in-the-wild videos. TRAM robustifies SLAM to recover the camera motion in the presence of dynamic humans and uses the scene background to derive the motion scale. Using the recovered camera as a metric-sc…

2024

Track Everything Everywhere Fast and Robustly

ECCV 2024poster

"We propose a novel test-time optimization approach for efficiently and robustly tracking any pixel at any time in a video. The latest state-of-the-art optimization-based tracking technique, OmniMotion, requires a prohibitively long optimization time, rendering it impractical for downstream applicat…

Cited by 6SourcePDFScholar
2024

Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies

IROS 2024poster

Large-scale robotic policies trained on data from diverse tasks and robotic platforms hold great promise for enabling general-purpose robots; however, reliable generalization to new environment conditions remains a major challenge. Toward addressing this challenge, we propose a novel approach for un…

Cited by 1SourcecodeScholar
2023

$\mathrm{SE}(3)$-Equivariant Attention Networks for Shape Reconstruction in Function Space

ICLR 2023poster

We propose a method for 3D shape reconstruction from unoriented point clouds. Our method consists of a novel SE(3)-equivariant coordinate-based network (TF-ONet), that parametrizes the occupancy field of the shape and respects the inherent symmetries of the problem. In contrast to previous shape rec…

Cited by 33SourcePDFScholar
2023

Banana: Banach Fixed-Point Network for Pointcloud Segmentation with Inter-Part Equivariance

NeurIPS 2023spotlight

Equivariance has gained strong interest as a desirable network property that inherently ensures robust generalization. However, when dealing with complex systems such as articulated objects or multi-object scenes, effectively capturing inter-part transformations poses a challenge, as it becomes enta…

Cited by 15SourcePDFScholar
2023

EFEM: Equivariant Neural Field Expectation Maximization for 3D Object Segmentation Without Scene Supervision

CVPR 2023poster

We introduce Equivariant Neural Field Expectation Maximization (EFEM), a simple, effective, and robust geometric algorithm that can segment objects in 3D scenes without annotations or training on scenes. We achieve such unsupervised segmentation by exploiting single object shape priors. We make two…

Cited by 22SourcePDFScholar
2023

Learning to Navigate in Turbulent Flows With Aerial Robot Swarms: A Cooperative Deep Reinforcement Learning Approach

RA-L 2023

Aerial operation in turbulent environments is a challenging problem due to the chaotic behavior of the flow. This problem is made even more complex when a team of aerial robots is trying to achieve coordinated motion in turbulent wind conditions. In this letter, we present a novel multi-robot contro

Cited by 8SourceScholar
2023

NAP: Neural 3D Articulated Object Prior

NeurIPS 2023poster

We propose Neural 3D Articulated object Prior (NAP), the first 3D deep generative model to synthesize 3D articulated object models. Despite the extensive research on generating 3D static objects, compositions, or scenes, there are hardly any approaches on capturing the distribution of articulated ob…

Cited by 16SourcePDFScholar
2023

NeuS2: Fast Learning of Neural Implicit Surfaces for Multi-view Reconstruction

ICCV 2023poster

Recent methods for neural surface representation and rendering, for example NeuS, have demonstrated the remarkably high-quality reconstruction of static scenes. However, the training of NeuS takes an extremely long time (8 hours), which makes it almost impossible to apply them to dynamic scenes with…

Cited by 276PDFcodeScholar
2022

CaDeX: Learning Canonical Deformation Coordinate Space for Dynamic Surface Representation via Neural Homeomorphism

CVPR 2022poster

While neural representations for static 3D shapes are widely studied, representations for deformable surfaces are limited to be template-dependent or to lack efficiency. We introduce Canonical Deformation Coordinate Space (CaDeX), a unified representation of both shape and nonrigid motion. Our key i…

Cited by 59PDFScholar
2022

Cross-Modal Map Learning for Vision and Language Navigation

CVPR 2022poster

We consider the problem of Vision-and-Language Navigation (VLN). The majority of current methods for VLN are trained end-to-end using either unstructured memory such as LSTM, or using cross-modal attention over the egocentric observations of the agent. In contrast to other works, our key insight is…

Cited by 83PDFcodeScholar
2022

EV-Catcher: High-Speed Object Catching Using Low-Latency Event-Based Neural Networks

RA-L 2022

Event-based sensors have recently drawn increasing interest in robotic perception due to their lower latency, higher dynamic range, and lower bandwidth requirements compared to standard CMOS-based imagers. These properties make them ideal tools for real-time perception tasks in highly dynamic enviro

Cited by 26SourceScholar
2022

EvAC3D: From Event-Based Apparent Contours to 3D Models via Continuous Visual Hulls

ECCV 2022poster

"3D reconstruction from multiple views is a successful computer vision field with multiple deployments in applications. State of the art is based on traditional RGB frames that enable optimization of photo-consistency cross views. In this paper, we study the problem of 3D reconstruction from event-c…

2022

Learning to Map for Active Semantic Goal Navigation

ICLR 2022poster

We consider the problem of object goal navigation in unseen environments. Solving this problem requires learning of contextual semantic priors, a challenging endeavour given the spatial and semantic variability of indoor environments. Current methods learn to implicitly encode these priors through g…

Cited by 94SourcePDFScholar
2022

Uncertainty-driven Planner for Exploration and Navigation

ICRA 2022poster

We consider the problems of exploration and pointgoal navigation in previously unseen environments, where the spatial complexity of indoor scenes and partial observability constitute these tasks challenging. We argue that learning occupancy priors over indoor maps provides significant advantages tow…

Cited by 70SourcecodeScholar
2022

Unified Fourier-based Kernel and Nonlinearity Design for Equivariant Networks on Homogeneous Spaces

ICML 2022spotlight

We introduce a unified framework for group equivariant networks on homogeneous spaces derived from a Fourier perspective. We consider tensor-valued feature fields, before and after a convolutional layer. We present a unified derivation of kernels via the Fourier domain by leveraging the sparsity of…

Cited by 21SourcePDFScholar
2021

An Adversarial Objective for Scalable Exploration

IROS 2021poster

Collecting new experience is costly in many robotic tasks, so determining how to efficiently explore in a new environment to learn as much as possible in as few trials as possible is an important problem for robotics. In this paper, we propose a method for exploring for the purpose of learning a dyn…

Cited by 9SourcecodeScholar
2021

Birds of a Feather: Capturing Avian Shape Models From Images

CVPR 2021poster

Animals are diverse in shape, but building a deformable shape model for a new species is not always possible due to the lack of 3D data. We present a method to capture new species using an articulated template and images of that species. In this work, we focus mainly on birds. Although birds represe…

Cited by 32PDFcodeScholar
2021

Deformable Linear Object Prediction Using Locally Linear Latent Dynamics

ICRA 2021poster

We propose a framework for deformable linear object prediction. Prediction of deformable objects (e.g., rope) is challenging due to their non-linear dynamics and infinite-dimensional configuration spaces. By mapping the dynamics from a non-linear space to a linear space, we can use the good properti…

Cited by 27SourcecodeScholar
2021

Discovering and Achieving Goals via World Models

NeurIPS 2021poster

How can artificial agents learn to solve many diverse tasks in complex visual environments without any supervision? We decompose this question into two challenges: discovering new goals and learning to reliably achieve them. Our proposed agent, Latent Explorer Achiever (LEXA), addresses both challen…

2021

Model-Based Reinforcement Learning via Latent-Space Collocation

ICML 2021spotlight

The ability to plan into the future while utilizing only raw high-dimensional observations, such as images, can provide autonomous agents with broad and general capabilities. However, realistic tasks require performing temporally extended reasoning, and cannot be solved with only myopic, short-sight…

2021

Probabilistic Modeling for Human Mesh Recovery

ICCV 2021poster

This paper focuses on the problem of 3D human reconstruction from 2D evidence. Although this is an inherently ambiguous problem, the majority of recent works avoid the uncertainty modeling and typically regress a single estimate for a given input. In contrast to that, in this work, we propose to emb…

Cited by 212PDFcodeScholar
2021

Self-Supervised Optical Flow with Spiking Neural Networks and Event Based Cameras

IROS 2021poster

Optical flow can be leveraged in robotic systems for obstacle detection where low latency solutions are critical in highly dynamic settings. While event-based cameras have changed the dominant paradigm of sending by encoding stimuli into spike trails, offering low bandwidth and latency, events are s…

Cited by 19SourceScholar
2021

Simple and Effective VAE Training with Calibrated Decoders

ICML 2021spotlight

Variational autoencoders (VAEs) provide an effective and simple method for modeling complex distributions. However, training VAEs often requires considerable hyperparameter tuning to determine the optimal amount of information retained by the latent variable. We study the impact of calibrated decode…

2020

3D Bird Reconstruction: a Dataset, Model, and Shape Recovery from a Single View

ECCV 2020poster

Model, and Shape Recovery from a Single View","Automated capture of animal pose is transforming how we study neuroscience and social behavior. Movements carry important social cues, but current methods are not able to robustly estimate pose and shape of animals, particularly for social animals such…

2020

A Low-Rank Matrix Approximation Approach to Multiway Matching with Applications in Multi-Sensory Data Association

ICRA 2020poster

Consider the case of multiple visual sensors perceiving the same scene from different viewpoints. In order to achieve consistent visual perception, the problem of data association, in this case establishing correspondences between observed features, must be first solved. In this work, we consider mu…

Cited by 5SourceScholar
2020

Coherent Reconstruction of Multiple Humans From a Single Image

CVPR 2020poster

In this work, we address the problem of multi-person 3D pose estimation from a single image. A typical regression approach in the top-down setting of this problem would first detect all humans and then reconstruct each one of them independently. However, this type of prediction suffers from incohere…

Cited by 208PDFcodeScholar
2020

Learning Predictive Models from Observation and Interaction

ECCV 2020poster

Learning predictive models from interaction with the world allows an agent, such as a robot, to learn about how the world works, and then use this learned model to plan coordinated sequences of actions to bring about desired outcomes. However, learning a model that captures the dynamics of complex s…

Cited by 65SourcePDFScholar
2020

Planning to Explore via Self-Supervised World Models

ICML 2020poster

Reinforcement learning allows solving complex tasks, however, the learning tends to be task-specific and the sample efficiency remains a challenge. We present Plan2Explore, a self-supervised reinforcement learning agent that tackles both these challenges through a new approach to self-supervised exp…

2020

Reactive Semantic Planning in Unexplored Semantic Environments Using Deep Perceptual Feedback

RA-L 2020

This letter presents a reactive planning system that enriches the topological representation of an environment with a tightly integrated semantic representation, achieved by incorporating and exploiting advances in deep perceptual learning and probabilistic semantic reasoning. Our architecture combi

Cited by 34SourceScholar
2020

Reinforcement Learning with Videos: Combining Offline Observations with Interaction

CoRL 2020

Reinforcement learning is a powerful framework for robots to acquire skills from experience, but often requires a substantial amount of online data collection. As a result, it is difficult to collect sufficiently diverse experiences that are needed for robots to generalize broadly. Videos of humans,

2020

Spike-FlowNet: Event-based Optical Flow Estimation with Energy-Efficient Hybrid Neural Networks

ECCV 2020poster

Event-based cameras display great potential for a variety of tasks such as high-speed motion detection and navigation in low-light environments where conventional frame-based cameras suffer critically. This is attributed to their high temporal resolution, high dynamic range, and low-power consumptio…

2020

TLIO: Tight Learned Inertial Odometry

RA-L 2020

In this letter we propose a tightly-coupled Extended Kalman Filter framework for IMU-only state estimation. Strap-down IMU measurements provide relative state estimates based on IMU kinematic motion model. However the integration of measurements is sensitive to sensor bias and noise, causing signifi

Cited by 241SourcecodeScholar
2020

The Tiercel: A novel autonomous micro aerial vehicle that can map the environment by flying into obstacles

ICRA 2020poster

Autonomous flight through unknown environments in the presence of obstacles is a challenging problem for micro aerial vehicles (MAVs). A majority of the current state-of-art research assumes obstacles as opaque objects that can be easily sensed by optical sensors such as cameras or LiDARs. However i…

Cited by 32SourceScholar
2019

Autonomous Precision Pouring From Unknown Containers

RA-L 2019

We autonomously pour from unknown symmetric containers found in a typical wet laboratory for the development of a robot-assisted, rapid experiment preparation system. The robot estimates the pouring container symmetric geometry, then leverages simulated pours as priors for a given fluid to pour prec

Cited by 46SourceScholar
2019

Convolutional Mesh Regression for Single-Image Human Shape Reconstruction

CVPR 2019oral

This paper addresses the problem of 3D human pose and shape estimation from a single image. Previous approaches consider a parametric model of the human body, SMPL, and attempt to regress the model parameters that give rise to a mesh consistent with image evidence. This parameter regression has been…

Cited by 670PDFScholar
2019

Cross-Domain 3D Equivariant Image Embeddings

ICML 2019oral

Spherical convolutional networks have been introduced recently as tools to learn powerful feature representations of 3D shapes. Spherical CNNs are equivariant to 3D rotations making them ideally suited to applications where 3D data may be observed in arbitrary orientations. In this paper we learn 2D…

Cited by 28SourcePDFScholar
2019

Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the Loop

ICCV 2019poster

Model-based human pose estimation is currently approached through two different paradigms. Optimization-based methods fit a parametric body model to 2D observations in an iterative manner, leading to accurate image-model alignments, but are often slow and sensitive to the initialization. In contrast…

Cited by 1232PDFScholar
2019

Learning what you can do before doing anything

ICLR 2019poster

Intelligent agents can learn to represent the action spaces of other agents simply by observing them act. Such representations help agents quickly learn to predict the effects of their own actions on the environment and to plan complex action sequences. In this work, we address the problem of learni…

2019

TexturePose: Supervising Human Mesh Estimation With Texture Consistency

ICCV 2019poster

This work addresses the problem of model-based human pose estimation. Recent approaches have made significant progress towards regressing the parameters of parametric human body models directly from images. Because of the absence of images with 3D shape ground truth, relevant approaches rely on 2D a…

Cited by 131PDFScholar
2019

Unsupervised Event-Based Learning of Optical Flow, Depth, and Egomotion

CVPR 2019poster

In this work, we propose a novel framework for unsupervised learning for event cameras that learns motion information from only the event stream. In particular, we propose an input representation of the events in the form of a discretized volume that maintains the temporal distribution of the events…

Cited by 648PDFScholar
2018

EV-FlowNet: Self-Supervised Optical Flow Estimation for Event-based Cameras

RSS 2018poster

Event-based cameras have shown great promise in a variety of situations where frame based cameras suffer, such as high speed motions and high dynamic range scenes. However, developing algorithms for event measurements requires a new class of hand crafted algorithms. Deep learning has shown great suc…

2018

Learning SO(3) Equivariant Representations with Spherical CNNs

ECCV 2018poster

We address the problem of 3D rotation equivariance in convolutional neural networks. 3D rotations have been a challenging nuisance in 3D classification tasks requiring higher capacity and extended data augmentation in order to tackle it. We model 3D data with multi-valued spherical functions and we…

2018

Learning to Estimate 3D Human Pose and Shape From a Single Color Image

CVPR 2018poster

This work addresses the problem of estimating the full body 3D human pose and shape from a single color image. This is a task where iterative optimization-based solutions have typically prevailed, while Convolutional Networks (ConvNets) have suffered because of the lack of training data and their lo…

Cited by 785SourcePDFScholar
2018

Semi-Dense Visual-Inertial Odometry and Mapping for Quadrotors with SWAP Constraints

ICRA 2018poster

Micro Aerial Vehicles have the potential to assist humans in real life tasks involving applications such as smart homes, search and rescue, and architecture construction. To enhance autonomous navigation capabilities these vehicles need to be able to create dense 3D maps of the environment, while co…

Cited by 13SourceScholar
2018

The Multivehicle Stereo Event Camera Dataset: An Event Camera Dataset for 3D Perception

RA-L 2018

Event-based cameras are a new passive sensing modality with a number of benefits over traditional cameras, including extremely low latency, asynchronous data acquisition, high dynamic range, and very low power consumption. There has been a lot of recent interest and development in applying algorithm

Cited by 583SourceScholar
2018

Understanding image motion with group representations

ICLR 2018poster

Motion is an important signal for agents in dynamic environments, but learning to represent motion from unlabeled video is a difficult and underconstrained problem. We propose a model of motion based on elementary group properties of transformations and use it to train a representation of image moti…

Cited by 5SourcePDFScholar
2017

6-DoF object pose from semantic keypoints

ICRA 2017poster

This paper presents a novel approach to estimating the continuous six degree of freedom (6-DoF) pose (3D translation and rotation) of an object from a single RGB image. The approach combines semantic keypoints predicted by a convolutional network (convnet) with a deformable shape model. Unlike prior…

Cited by 540SourceScholar
2017

Active end-effector pose selection for tactile object recognition through Monte Carlo tree search

IROS 2017poster

This paper considers the problem of active object recognition using touch only. The focus is on adaptively selecting a sequence of wrist poses that achieves accurate recognition by enclosure grasps. It seeks to minimize the number of touches and maximize recognition confidence. The actions are formu…

Cited by 29SourceScholar
2017

Autonomous Flight for Detection, Localization, and Tracking of Moving Targets With a Small Quadrotor

RA-L 2017

In this letter, we address the autonomous flight of a small quadrotor, enabling tracking of a moving object. The 15-cm diameter, 250-g robot relies only on onboard sensors (a single camera and an inertial measurement unit) and computers, and can detect, localize, and track moving objects. Our key co

Cited by 118SourceScholar
2017

Coarse-To-Fine Volumetric Prediction for Single-Image 3D Human Pose

CVPR 2017spotlight

This paper addresses the challenge of 3D human pose estimation from a single color image. Despite the general success of the end-to-end learning paradigm, top performing approaches employ a two-step solution consisting of a Convolutional Network (ConvNet) for 2D joint localization and a subsequent o…

Cited by 1177PDFScholar
2017

Distributed consistent data association via permutation synchronization

ICRA 2017poster

Data association is one of the fundamental problems in multi-sensor systems. Most current techniques rely on pairwise data associations which can be spurious even after the employment of outlier rejection schemes. Considering multiple pairwise associations at once significantly increases accuracy an…

Cited by 52SourceScholar
2017

Harvesting Multiple Views for Marker-Less 3D Human Pose Annotations

CVPR 2017spotlight

Recent advances with Convolutional Networks (ConvNets) have shifted the bottleneck for many computer vision tasks to annotated data collection. In this paper, we present a geometry-driven approach to automatically collect annotations for human pose prediction tasks. Starting from a generic ConvNet f…

Cited by 248PDFScholar
2017

PennCOSYVIO: A challenging Visual Inertial Odometry benchmark

ICRA 2017poster

We present PennCOSYVIO, a new challenging Visual Inertial Odometry (VIO) benchmark with synchronized data from a VI-sensor (stereo camera and IMU), two Project Tango hand-held devices, and three GoPro Hero 4 cameras. Recorded at UPenn's Singh center, the 150m long path of the hand-held rig crosses f…

Cited by 115SourcecodeScholar
2017

Precise dispensing of liquids using visual feedback

IROS 2017poster

Robotic pouring is an important step in improving the safety, productivity and repeatability in the biotechnology industry and generally increasing the effectiveness of robotics in human based environments. In this work we present a method to autonomously dispense a precise amount of fluid using onl…

Cited by 33SourceScholar
2017

Probabilistic data association for semantic SLAM

ICRA 2017poster

Traditional approaches to simultaneous localization and mapping (SLAM) rely on low-level geometric features such as points, lines, and planes. They are unable to assign semantic labels to landmarks observed in the environment. Furthermore, loop closure recognition based on low-level features is ofte…

Cited by 605SourceScholar
2017

Shape-based object classification and recognition through continuum manipulation

IROS 2017poster

We introduce a novel approach to shape-based object classification and recognition through the use of a continuum manipulator. Noticing the fact that when a continuum manipulator wraps around an object in a whole-arm grasping, its own shape is indicative of the shape of the object, our approach enab…

Cited by 10SourceScholar
2016

A triangle histogram for object classification by tactile sensing

IROS 2016poster

We present a new descriptor for tactile 3D object classification. It is invariant to object movement and simple to construct, using only the relative geometry of points on the object surface. We demonstrate successful classification of 185 objects in 10 categories, at sparse to dense surface samplin…

Cited by 35SourceScholar
2016

Articulated motion estimation from a monocular image sequence using spherical tangent bundles

ICRA 2016

We propose a second order stochastic dynamical model for generic articulated objects whose state space is a Riemannian manifold naturally suggested by the articulation constraints. We derive the equations of a Riemannian Extended Kalman Filter to perform the structure estimation from an image sequen

Cited by 23SourceScholar
2016

Seeing Glassware: from Edge Detection to Pose Estimation and Shape Recovery

RSS 2016poster

Perception of transparent objects has been an open challenge in robotics despite advances in sensors and data- driven learning approaches. In this paper, we introduce a new approach that combines recent advances in learnt object detectors with perceptual grouping in 2D, and projective geometry of ap…

Cited by 67SourcePDFScholar
2016

Sparseness Meets Deepness: 3D Human Pose Estimation From Monocular Video

CVPR 2016spotlight

This paper addresses the challenge of 3D full-body human pose estimation from a monocular image sequence. Here, two cases are considered: (i) the image locations of the human joints are provided and (ii) the image locations of joints are unknown. In the former case, a novel approach is introduced th…

Cited by 545PDFScholar
2016

Visual Servoing of Quadrotors for Perching by Hanging From Cylindrical Objects

RA-L 2016

This letter addresses vision-based localization and servoing for quadrotors to enable autonomous perching by hanging from cylindrical structures using only a monocular camera. We focus on the problems of relative pose estimation, control, and trajectory planning for maneuvering a robot relative to c

Cited by 101SourceScholar
2015

3D Shape Estimation From 2D Landmarks: A Convex Relaxation Approach

CVPR 2015poster

We investigate the problem of estimating the 3D shape of an object, given a set of 2D landmarks in a single image. To alleviate the reconstruction ambiguity, a widely-used approach is to confine the unknown 3D shape within a shape space built upon existing shapes. While this approach has proven to b…

Cited by 141SourcePDFScholar
2015

A Metric Parametrization for Trifocal Tensors With Non-Colinear Pinholes

CVPR 2015poster

The trifocal tensor, which describes the relation between projections of points and lines in three views, is a fundamental entity of geometric computer vision. In this work, we investigate a new parametrization of the trifocal tensor for calibrated cameras with non-colinear pinholes obtained from a…

Cited by 21SourcePDFScholar
2015

Decentralized active information acquisition: Theory and application to multi-robot SLAM

ICRA 2015poster

This paper addresses the problem of controlling mobile sensing systems to improve the accuracy and efficiency of gathering information autonomously. It applies to scenarios such as environmental monitoring, search and rescue, surveillance and reconnaissance, and simultaneous localization and mapping…

Cited by 238SourceScholar
2015

Grasping surfaces of revolution: Simultaneous pose and shape recovery from two views

ICRA 2015poster

In many scenarios, robots encounter rotationally symmetric objects for which no known 3D model exists. To be able to grasp such objects using existing grasp point computation schemes, an estimate of their 3D-pose and shape is necessary. In this paper, we address the problem of recovering 3D-pose and…

Cited by 17SourceScholar
2015

Initialization techniques for 3D SLAM: A survey on rotation estimation and its use in pose graph optimization

ICRA 2015poster

Pose graph optimization is the non-convex optimization problem underlying pose-based Simultaneous Localization and Mapping (SLAM). If robot orientations were known, pose graph optimization would be a linear least-squares problem, whose solution can be computed efficiently and reliably. Since rotatio…

Cited by 302SourceScholar