← Search

Hazel Doughty

15 accepted papers

2025

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

CVPR 2025poster

We present a validation dataset of newly-collected kitchen based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, al…

Cited by 3SourcePDFScholar
2024

SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery

ECCV 2024poster

"In this paper, we address Generalized Category Discovery, aiming to simultaneously uncover novel categories and accurately classify known ones. Traditional methods, which lean heavily on self-supervision and contrastive learning, often fall short when distinguishing between fine-grained categories.…

2023

Learn to Categorize or Categorize to Learn? Self-Coding for Generalized Category Discovery

NeurIPS 2023poster

In the quest for unveiling novel categories at test time, we confront the inherent limitations of traditional supervised recognition models that are restricted by a predefined category set. While strides have been made in the realms of self-supervised and open-world learning towards test-time catego…

2023

Tubelet-Contrastive Self-Supervision for Video-Efficient Generalization

ICCV 2023poster

We propose a self-supervised method for learning motion-focused video representations. Existing approaches minimize distances between temporally augmented videos, which maintain high spatial similarity. We instead propose to learn similarities between videos with identical local motion dynamics but…

Cited by 13PDFcodeScholar
2022

How Severe Is Benchmark-Sensitivity in Video Self-Supervised Learning?

ECCV 2022poster

"Despite the recent success of video self-supervised learning models, there is much still to be understood about their generalization capability. In this paper, we investigate how sensitive video self-supervised learning is to the current conventional benchmark and whether methods generalize beyond…

2020

Action Modifiers: Learning From Adverbs in Instructional Videos

CVPR 2020poster

We present a method to learn a representation for adverbs from instructional videos using weak supervision from the accompanying narrations. Key to our method is the fact that the visual representation of the adverb is highly dependent on the action to which it applies, although the same adverb will…

Cited by 38PDFcodeScholar
2019

The Pros and Cons: Rank-Aware Temporal Attention for Skill Determination in Long Videos

CVPR 2019poster

We present a new model to determine relative skill from long videos, through learnable temporal attention modules. Skill determination is formulated as a ranking problem, making it suitable for common and generic tasks. However, for long videos, parts of the video are irrelevant for assessing skill,…

Cited by 137PDFcodeScholar
2018

Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

ECCV 2018poster

First-person vision is gaining interest as it offers a unique viewpoint on people’s interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of sufficiently large datasets. In this paper, we introduce EPIC-KI…

Cited by 1329SourcePDFScholar
2018

Who's Better? Who's Best? Pairwise Deep Ranking for Skill Determination

CVPR 2018poster

This paper presents a method for assessing skill from video, applicable to a variety of tasks, ranging from surgery to drawing and rolling pizza dough. We formulate the problem as pairwise (who’s better?) and overall (who’s best?) ranking of video collections, using supervised deep ranking. We propo…