← Search

Walterio Mayol-Cuevas

18 accepted papers

2025

Focal Plane Visual Feature Generation and Matching on a Pixel Processor Array

ICCV 2025poster

Pixel Processor Arrays (PPAs) are vision sensors that embed data and processing into every pixel element. PPAs can execute visual processing directly at the point of light capture, and output only sparse, high-level information. This is in sharp contrast with the conventional visual pipeline, where…

Cited by 0SourcePDFScholar
2021

Weighted Node Mapping and Localisation on a Pixel Processor Array

ICRA 2021poster

This paper implements and demonstrates visual route mapping and localisation upon a Pixel Processor Array (PPA). The PPA sensor comprises of an array of Processing Elements (PEs), each of which can capture and process visual information directly. This provides significant parallel processing power a…

Cited by 6SourceScholar
2020

Action Modifiers: Learning From Adverbs in Instructional Videos

CVPR 2020poster

We present a method to learn a representation for adverbs from instructional videos using weak supervision from the accompanying narrations. Key to our method is the fact that the visual representation of the adverb is highly dependent on the action to which it applies, although the same adverb will…

Cited by 38PDFcodeScholar
2020

Centroids Triplet Network and Temporally-Consistent Embeddings for In-Situ Object Recognition

IROS 2020poster

This work proposes learning to recognize objects from a small number of training examples collected and deployed in-situ. That is, from data collected where the objects are commonly placed or being used, perhaps after first encountering them, the learning algorithm immediately is able to recognize t…

Cited by 5SourceScholar
2019

A Camera That CNNs: Towards Embedded Neural Networks on Pixel Processor Arrays

ICCV 2019oral

We present a convolutional neural network implementation for pixel processor array (PPA) sensors. PPA hardware consists of a fine-grained array of general-purpose processing elements, each capable of light capture, data storage, program execution, and communication with neighboring elements. This al…

Cited by 47PDFScholar
2019

Learning Discriminative Embeddings for Object Recognition on-the-fly

ICRA 2019poster

We address the problem of learning to recognize new objects on-the-fly efficiently. When using CNNs, a typical approach for learning new objects is by fine-tuning the model. However, this approach relies on the assumption that the original training set is available and requires high-end computationa…

Cited by 17SourceScholar
2019

The Pros and Cons: Rank-Aware Temporal Attention for Skill Determination in Long Videos

CVPR 2019poster

We present a new model to determine relative skill from long videos, through learnable temporal attention modules. Skill determination is formulated as a ranking problem, making it suitable for common and generic tasks. However, for long videos, parts of the video are irrelevant for assessing skill,…

Cited by 137PDFcodeScholar
2018

Perspective Correcting Visual Odometry for Agile MAVs using a Pixel Processor Array

IROS 2018poster

This paper presents a visual odometry approach using a Pixel Processor Array (PPA) camera, specifically, the SCAMP-5 vision chip. In this device, each pixel is capable of storing data and performing computation, enabling a variety of computer vision tasks to be carried out directly upon the sensor i…

Cited by 13SourceScholar
2018

Where can i do this? Geometric Affordances from a Single Example with the Interaction Tensor

ICRA 2018poster

This paper introduces and evaluates a new tensor field representation to express the geometric affordance of one object relative to another, a key competence for Cognitive and Autonomous robots. We expand the bisector surface representation to one that is weight-driven and that retains the provenanc…

Cited by 10SourceScholar
2018

Who's Better? Who's Best? Pairwise Deep Ranking for Skill Determination

CVPR 2018poster

This paper presents a method for assessing skill from video, applicable to a variety of tasks, ranging from surgery to drawing and rolling pizza dough. We formulate the problem as pairwise (who’s better?) and overall (who’s best?) ranking of video collections, using supervised deep ranking. We propo…

2017

Tracking control of a UAV with a parallel visual processor

IROS 2017poster

This paper presents a vision-based control strategy for tracking a ground target using a novel vision sensor featuring a processor for each pixel element. This enables computer vision tasks to be carried out directly on the focal plane in a highly efficient manner rather than using a separate genera…

Cited by 28SourceScholar
2017

Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric Video

ICCV 2017poster

Manual annotations of temporal bounds for object interactions (i.e. start and end times) are typical training input to recognition, localization and detection algorithms. For three publicly available egocentric datasets, we uncover inconsistencies in ground truth temporal bounds within and across an…

Cited by 36PDFScholar
2017

Visual Odometry for Pixel Processor Arrays

ICCV 2017spotlight

We present an approach of estimating constrained motion of a novel Cellular Processor Array (CPA) camera, on which each pixel is capable of limited processing and data storage allowing for fast low power parallel computation to be carried out directly on the focal-plane of the device. Rather than th…

Cited by 36PDFScholar
2015

Improving MAV control by predicting aerodynamic effects of obstacles

IROS 2015poster

Building on our previous work [1], in this paper we demonstrate how it is possible to improve flight control of a MAV that experiences aerodynamic disturbances caused by objects on its path. Predictions based on low resolution depth images taken at a distance are incorporated into the flight control…

Cited by 9SourceScholar
2015

Inverse depth for accurate photometric and geometric error minimisation in RGB-D dense visual odometry

ICRA 2015poster

In this paper we present a dense visual odometry system for RGB-D cameras performing both photometric and geometric error minimisation to estimate the camera motion between frames. Contrary to most works in the literature, we parametrise the geometric error by the inverse depth instead of the depth,…

Cited by 47SourceScholar
2015

What should I landmark? Entropy of normals in depth juts for place recognition in changing environments using RGB-D data

ICRA 2015poster

One open problem in the fields of place recognition and mapping is to be able to recognise a revisited place when its appearance and layout have changed between visits. In this paper, we investigate this problem in the context of RGB-D mapping in indoor environments. We propose to segment the scene…

Cited by 6SourceScholar