← Search

Simon Stent

16 accepted papers

2024

COCO-Periph: Bridging the Gap Between Human and Machine Perception in the Periphery

ICLR 2024poster

Evaluating deep neural networks (DNNs) as models of human perception has given rich insights into both human visual processing and representational properties of DNNs. We extend this work by analyzing how well DNNs perform compared to humans when constrained by peripheral vision -- which limits huma…

Cited by 3SourcePDFScholar
2024

Seeing Faces in Things: A Model and Dataset for Pareidolia

ECCV 2024poster

"The human visual system is well-tuned to detect faces of all shapes and sizes. While this brings obvious survival advantages, such as a better chance of spotting unknown predators in the bush, it also leads to spurious face detections. “Face pareidolia” describes the perception of face-like structu…

2023

Exploring perceptual straightness in learned visual representations

ICLR 2023poster

Humans have been shown to use a ''straightened'' encoding to represent the natural visual world as it evolves in time (Henaff et al. 2019). In the context of discrete video sequences, ''straightened'' means that changes between frames follow a more linear path in representation space at progressivel…

Cited by 5SourcePDFScholar
2023

Tracking Through Containers and Occluders in the Wild

CVPR 2023poster

Tracking objects with persistence in cluttered and dynamic environments remains a difficult challenge for computer vision systems. In this paper, we introduce TCOW, a new benchmark and model for visual tracking through heavy occlusion and containment. We set up a task where the goal is to, given a v…

2023

What You Can Reconstruct From a Shadow

CVPR 2023poster

3D reconstruction is a fundamental problem in computer vision, and the task is especially challenging when the object to reconstruct is partially or fully occluded. We introduce a method that uses the shadows cast by an unobserved object in order to infer the possible 3D volumes under occlusion. We…

Cited by 3SourcePDFScholar
2022

"Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications"

ECCV 2022poster

"Egocentric videos offer fine-grain information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding viewer’s behaviors and intentions. We provide a labeled dataset consisting of 11,235 egocentric images with per-pixel segmentation labe…

2022

Look Both Ways: Self-Supervising Driver Gaze Estimation and Road Scene Saliency

ECCV 2022poster

"We present a new on-road driving dataset, called “Look Both Ways”, which contains synchronized video of both driver faces and the forward road scene, along with ground truth gaze data registered from eye tracking glasses worn by the drivers. Our dataset supports the study of methods for non-intrusi…

2022

Revealing Occlusions With 4D Neural Fields

CVPR 2022oral

For computer vision systems to operate in dynamic situations, they need to be able to represent and reason about object permanence. We introduce a framework for learning to estimate 4D visual representations from monocular RGB-D video, which is able to persist objects, even once they become obstruct…

Cited by 15PDFScholar
2021

LocTex: Learning Data-Efficient Visual Representations From Localized Textual Supervision

ICCV 2021poster

Computer vision tasks such as object detection and semantic/instance segmentation rely on the painstaking annotation of large training datasets. In this paper, we propose LocTex that takes advantage of the low-cost localized textual annotations (i.e., captions and synchronized mouse-over gestures) t…

Cited by 14PDFScholar
2021

The Way to My Heart Is Through Contrastive Learning: Remote Photoplethysmography From Unlabelled Video

ICCV 2021poster

The ability to reliably estimate physiological signals from video is a powerful tool in low-cost, pre-clinical health monitoring. In this work we propose a new approach to remote photoplethysmography (rPPG) -- the measurement of blood volume changes from observations of a person's face or skin. Simi…

Cited by 133PDFcodeScholar
2019

Gaze360: Physically Unconstrained Gaze Estimation in the Wild

ICCV 2019poster

Understanding where people are looking is an informative social cue. In this work, we present Gaze360, a large-scale remote gaze-tracking dataset and method for robust 3D gaze estimation in unconstrained images. Our dataset consists of 238 subjects in indoor and outdoor environments with labelled 3D…

Cited by 475PDFScholar
2018

A Multi-Camera Deep Neural Network for Detecting Elevated Alertness in Drivers

ICASSP 2018accepted

We present a system for the detection of elevated levels of driver alertness in driver-facing video captured from multiple viewpoints. This problem is important in automotive safety as a helpful feedback signal to determine driver engagement and as a means of automatically flagging anomalous driving…

Cited by 0SourceScholar
2018

Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks

ECCV 2018poster

We introduce a saliency-based distortion layer for convolutional neural networks that helps to improve the spatial sampling of input data for a given task. Our differentiable layer can be added as a preprocessing block to existing task networks and trained altogether in an end-to-end fashion. The ef…

2016

SceneNet: An annotated model generator for indoor scene understanding

ICRA 2016

We introduce SceneNet, a framework for generating high-quality annotated 3D scenes to aid indoor scene understanding. SceneNet leverages manually-annotated datasets of real world scenes such as NYUv2 to learn statistics about object co-occurrences and their spatial relationships. Using a hierarchica

Cited by 110SourceScholar
2016

Street-View Change Detection with Deconvolutional Networks

RSS 2016poster

We propose a system for performing structural change detection in street-view videos captured by a vehicle- mounted monocular camera over time. Our approach is moti- vated by the need for more frequent and efficient updates in the large-scale maps used in autonomous vehicle navigation. Our method ch…

Cited by 401SourcePDFScholar
2016

Understanding Real World Indoor Scenes With Synthetic Data

CVPR 2016poster

Scene understanding is a prerequisite to many high level tasks for any automated intelligent machine operating in real world environments. Recent attempts with supervised learning have shown promise in this direction but also highlighted the need for enormous quantity of supervised data --- performa…

Cited by 448PDFScholar