← Search

Evangelos Kazakos

6 accepted papers

2025

Large-scale Pre-training for Grounded Video Caption Generation

ICCV 2025poster

We propose a novel approach for captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally dense bounding boxes. We introduce the following contributions. First, we present a large-scale automatic annotation method that aggregates frame-level…

2024

TIM: A Time Interval Machine for Audio-Visual Action Recognition

CVPR 2024poster

Diverse actions give rise to rich audio-visual signals in long videos. Recent works showcase that the two modalities of audio and video exhibit different temporal extents of events and distinct labels. We address the interplay between the two modalities in long videos by explicitly modelling the tem…

2023

Epic-Sounds: A Large-Scale Dataset of Actions that Sound

ICASSP 2023accepted

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100. We propose an annotation pipeline where annotators temporally label distinguishable audio segments and describe th…

Cited by 0SourceScholar
2019

EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition

ICCV 2019poster

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three modalities -- RGB, Flow and Audio -- and combine them with…

Cited by 438PDFcodeScholar
2018

Scaling Egocentric Vision: The EPIC-KITCHENS Dataset

ECCV 2018poster

First-person vision is gaining interest as it offers a unique viewpoint on people’s interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of sufficiently large datasets. In this paper, we introduce EPIC-KI…

Cited by 1329SourcePDFScholar