← Search

Gül Varol

18 accepted papers

2025

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

CVPR 2025poster

We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, automatic chaptering has the potential to enable efficient navigation and content retrieval in long-form videos. In this…

Cited by 0SourcePDFScholar
2025

Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues

CVPR 2025poster

Our objective is to translate continuous sign language into spoken language text. Inspired by the way human interpreters rely on context for accurate translation, we incorporate additional contextual cues together with the signing video, into a new translation framework. Specifically, besides visual…

Cited by 3SourcePDFScholar
2025

Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

ICCV 2025poster

Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video understanding. This includes extending temporal context to neighbo…

Cited by 0SourcePDFScholar
2024

AutoAD III: The Prequel - Back to the Pixels

CVPR 2024poster

Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently visual language models for AD generation are limited by a lack of suitable training data and also their evaluation is hampe…

Cited by 19SourcePDFScholar
2024

CoVR: Learning Composed Video Retrieval from Web Video Captions

AAAI 2024technical

Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database. Most CoIR approaches require manually annotated datasets, comprising image-text-image triplets, where the text describes a modifi…

Cited by 46SourcePDFScholar
2023

AutoAD: Movie Description in Context

CVPR 2023highlight

The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the dependency of the descriptions on context, and the limited amount of training data available. In this work, we leverage…

2023

SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation

ICCV 2023poster

Our goal is to synthesize 3D human motions given textual inputs describing simultaneous actions, for example `waving hand' while `walking' at the same time. We refer to generating such simultaneous movements as performing `spatial compositions'. In contrast to `temporal compositions' that seek to tr…

Cited by 48PDFScholar
2022

Automatic Dense Annotation of Large-Vocabulary Sign Language Videos

ECCV 2022poster

"Recently, sign language researchers have turned to sign language interpreted TV broadcasts, comprising (i) a video of continuous signing and (ii) subtitles corresponding to the audio content, as a readily available and large-scale source of training data. One key challenge in the usability of such…

Cited by 24SourcePDFScholar
2022

Sign Language Video Retrieval With Free-Form Textual Queries

CVPR 2022poster

Systems that can efficiently search collections of sign language videos have been highlighted as a useful application of sign language technology. However, the problem of searching videos beyond individual keywords has received limited attention in the literature. To address this gap, in this work w…

Cited by 41PDFScholar
2022

TEMOS: Generating Diverse Human Motions from Textual Descriptions

ECCV 2022poster

"We address the problem of generating diverse 3D human motions from textual descriptions. This challenging task requires joint modeling of both modalities: understanding and extracting useful human-centric information from the text, and then generating plausible and realistic sequences of human pose…

2021

Aligning Subtitles in Sign Language Videos

ICCV 2021poster

The goal of this work is to temporally align asynchronous subtitles in sign language videos. In particular, we focus on sign-language interpreted TV broadcast data comprising (i) a video of continuous signing, and (ii) subtitles corresponding to the audio content. Previous work exploiting such weakl…

Cited by 39PDFScholar
2021

Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

ICCV 2021poster

Our objective in this work is video-text retrieval - in particular a joint embedding that enables efficient text-to-video retrieval. The challenges in this area include the design of the visual architecture and the nature of the training data, in that the available large scale video-text training da…

Cited by 1303PDFcodeScholar
2021

SeeHear: Signer Diarisation and a New Dataset

ICASSP 2021accepted

In this work, we propose a framework to collect a large-scale, diverse sign language dataset that can be used to train automatic sign language recognition models.The first contribution of this work is SDTrack, a generic method for signer tracking and diarisation in the wild. Our second contribution…

Cited by 0SourceScholar
2021

Sign Language Segmentation with Temporal Convolutional Networks

ICASSP 2021accepted

The objective of this work is to determine the location of temporal boundaries between signs in continuous sign language videos. Our approach employs 3D convolutional neural network representations with iterative temporal segment refinement to resolve ambiguities between sign boundary cues. We demon…

Cited by 0SourceScholar
2020

BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues

ECCV 2020poster

Recent progress in fine-grained gesture and action classification, and machine translation, point to the possibility of automated sign language recognition becoming a reality. A key stumbling block in making progress towards this goal is a lack of appropriate training data, stemming from the high co…

Cited by 221SourcePDFScholar