← Search

Edward Johns

43 accepted papers

2026

Observer–Actor: Active Vision Imitation Learning with Sparse-View Gaussian Splatting

ICRA 2026poster

We propose Observer-Actor (ObAct), a novel framework for active vision imitation learning in which the observer moves to optimal visual observations for the actor. We study ObAct on a dual-arm robotic system equipped with wrist-mounted cameras. At test time, ObAct dynamically assigns observer and ac…

2026

ZeroBot: Learning From Scratch in Minutes With Generative Real2Sim

RA-L 2026

We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses

Cited by 0SourcecodeScholar
2025

Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE Solutions

NeurIPS 2025poster

Stochastic differential equations (SDEs) are well suited to modelling noisy and/or irregularly-sampled time series, which are omnipresent in finance, physics, and machine learning applications. Traditional approaches require costly simulation of numerical solvers when sampling between arbitrary time…

Cited by 0SourceScholar
2024

Adapting Skills to Novel Grasps: A Self-Supervised Approach

IROS 2024poster

In this paper, we study the problem of adapting manipulation trajectories involving grasped objects (e.g. tools) defined for a single grasp pose to novel grasp poses. A common approach to address this is to define a new trajectory for each possible grasp explicitly, but this is highly inefficient. I…

Cited by 1SourceScholar
2024

Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models

ICRA 2024poster

We introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline. This is achieved by the robot autonomously constructing a 3D representation of the scene, where objects can be rearranged virtually and an image of…

Cited by 19SourceScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2023

Learning Tethered Perching for Aerial Robots

ICRA 2023poster

Aerial robots have a wide range of applications, such as collecting data in hard-to-reach areas. This requires the longest possible operation time. However, because currently available commercial batteries have limited specific energy of roughly 300 W h kg-1, a drone's flight time is a bottleneck fo…

Cited by 16SourceScholar
2022

Bootstrapping Semantic Segmentation with Regional Contrast

ICLR 2022poster

We present ReCo, a contrastive learning framework designed at a regional level to assist learning in semantic segmentation. ReCo performs pixel-level contrastive learning on a sparse set of hard negative pixels, with minimal additional memory footprint. ReCo is easy to implement, being built on top…

2022

Demonstrate Once, Imitate Immediately (DOME): Learning Visual Servoing for One-Shot Imitation Learning

IROS 2022poster

We present DOME, a novel method for one-shot imitation learning, where a task can be learned from just a single demonstration and then be deployed immediately, without any further data collection or training. DOME does not require prior task or object knowledge, and can perform the task in novel obj…

Cited by 41SourceScholar
2022

Real-time Mapping of Physical Scene Properties with an Autonomous Robot Experimenter

CoRL 2022poster

Neural fields can be trained from scratch to represent the shape and appearance of 3D scenes efficiently. It has also been shown that they can densely map correlated properties such as semantics, via sparse interactions from a human labeller. In this work, we show that a robot can densely annotate a…

Cited by 5SourceScholar
2021

Coarse-to-Fine for Sim-to-Real: Sub-Millimetre Precision Across Wide Task Spaces

IROS 2021poster

In this paper, we study the problem of zero-shot sim-to-real when the task requires both highly precise control with sub-millimetre error tolerance, and wide task space generalisation. Our framework involves a coarse-to-fine controller, where trajectories begin with classical motion planning using I…

Cited by 21SourceScholar
2021

DROID: Minimizing the Reality Gap Using Single-Shot Human Demonstration

RA-L 2021

Reinforcement learning (RL) has demonstrated great success in the past several years. However, most of the scenarios focus on simulated environments. One of the main challenges of transferring the policy learned in a simulated environment to real world, is the discrepancy between the dynamics of the

Cited by 36SourceScholar
2021

Hybrid ICP

IROS 2021poster

ICP algorithms typically involve a fixed choice of data association method and a fixed choice of error metric. In this paper, we propose Hybrid ICP, a novel and flexible ICP variant which dynamically optimises both the data association method and error metric based on the live image of an object and…

Cited by 3SourceScholar
2020

Constrained-Space Optimization and Reinforcement Learning for Complex Tasks

RA-L 2020

Learning from demonstration is increasingly used for transferring operator manipulation skills to robots. In practice, it is important to cater for limited data and imperfect human demonstrations, as well as underlying safety constraints. This article presents a constrained-space optimization and re

Cited by 16SourceScholar
2020

Crossing the Gap: A Deep Dive into Zero-Shot Sim-to-Real Transfer for Dynamics

IROS 2020poster

Zero-shot sim-to-real transfer of tasks with complex dynamics is a highly challenging and unsolved problem. A number of solutions have been proposed in recent years, but we have found that many works do not present a thorough evaluation in the real world, or underplay the significant engineering eff…

Cited by 49SourceScholar
2020

Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning

IROS 2020poster

Dexterous manipulation of objects in virtual environments with our bare hands, by using only a depth sensor and a state-of-the-art 3D hand pose estimator (HPE), is challenging. While virtual environments are ruled by physics, e.g. object weights and surface frictions, the absence of force feedback m…

Cited by 62SourceScholar
2020

Shape Adaptor: A Learnable Resizing Module

ECCV 2020poster

We present a novel resizing module for neural networks: shape adaptor, a drop-in enhancement built on top of traditional resizing layers, such as pooling, bilinear sampling, and strided convolution. Whilst traditional resizing layers have fixed and deterministic reshaping factors, our module allows…

2019

Self-Supervised Generalisation with Meta Auxiliary Learning

NeurIPS 2019poster

Learning with auxiliary tasks can improve the ability of a primary task to generalise. However, this comes at the cost of manually labelling auxiliary data. We propose a new method which automatically learns appropriate labels for an auxiliary task, such that any supervised learning task can be impr…

2017

Application-oriented design space exploration for SLAM algorithms

ICRA 2017poster

In visual SLAM, there are many software and hardware parameters, such as algorithmic thresholds and GPU frequency, that need to be tuned; however, this tuning should also take into account the structure and motion of the camera. In this paper, we determine the complexity of the structure and motion…

Cited by 39SourceScholar
2017

Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task

CoRL 2017

End-to-end control for robot manipulation and grasping is emerging as an attractive alternative to traditional pipelined approaches. However, end-to-end methods tend to either be slow to train, exhibit little or no generalisability, or lack the ability to accomplish long-horizon or multi-stage tasks

Cited by 0SourcePDFScholar
2016

Deep learning a grasp function for grasping under gripper pose uncertainty

IROS 2016poster

This paper presents a new method for parallel-jaw grasping of isolated objects from depth images, under large gripper pose uncertainty. Whilst most approaches aim to predict the single best grasp pose from an image, our method first predicts a score for every possible grasp pose, which we denote the…

Cited by 307SourceScholar
2016

Pairwise Decomposition of Image Sequences for Active Multi-View Recognition

CVPR 2016oral

A multi-view image sequence provides a much richer capacity for object recognition than from a single image. However, most existing solutions to multi-view recognition typically adopt hand-crafted, model-based geometric methods, which do not readily embrace recent trends in deep learning. We propose…

Cited by 308PDFScholar