← Search

Necati Cihan Camgoz

16 accepted papers

2026

OSMO: Open-vocabulary Self-eMOtion Tracking

CVPR 2026

We introduce the novel task of egocentric self-emotion tracking, which aims to infer an individual's evolving emotions from egocentric multimodal streams such as voice, visual surroundings, semantic subtext, and eye-tracking signals. To establish this research direction, we present: (1) OSMO dataset

Cited by 0SourcecodeScholar
2025

2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset Download PDF

ACL 2025finding

We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 91 spoken languages at the intersection of BELEBELE and FLEURS, and one sign language (ASL). As a by-product we also extend the Automatic Speech Recognit…

2025

Streaming VideoLLMs for Real-Time Procedural Video Understanding

ICCV 2025poster

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens -- verbalized text tokens, which provide compressed textual summaries of long-term observations, and visual tokens, encoded…

Cited by 9SourcePDFScholar
2024

POET: Prompt Offset Tuning for Continual Human Action Adaptation

ECCV 2024oral

"As extended reality (XR) is redefining how users interact with computing devices, research in human action recognition is gaining prominence. Typically, models deployed on immersive computing devices are static and limited to their default set of classes. The goal of our research is to provide user…

2024

Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language Translation

ICLR 2024poster

Automatic Sign Language Translation requires the integration of both computer vision and natural language processing to effectively bridge the communication gap between sign and spoken languages. However, the deficiency in large-scale training data to support sign language translation means we need…

Cited by 36SourcePDFScholar
2024

Towards Privacy-Aware Sign Language Translation at Scale

ACL 2024long

A major impediment to the advancement of sign language translation (SLT) is data scarcity. Much of the sign language data currently available on the web cannot be used for training supervised models due to the lack of aligned captions. Furthermore, scaling SLT using large-scale web-scraped datasets…

2023

Data-Free Class-Incremental Hand Gesture Recognition

ICCV 2023poster

This paper investigates data-free class-incremental learning (DFCIL) for hand gesture recognition from 3D skeleton sequences. In this class-incremental learning (CIL) setting, while incrementally registering the new classes, we do not have access to the training samples (i.e. data-free) of t…

Cited by 9PDFcodeScholar
2022

Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language Production

CVPR 2022poster

Sign languages are visual languages, with vocabularies as rich as their spoken language counterparts. However, current deep-learning based Sign Language Production (SLP) models produce under-articulated skeleton pose sequences from constrained vocabularies and this limits applicability. To be unders…

Cited by 76PDFcodeScholar
2021

Mixed SIGNals: Sign Language Production via a Mixture of Motion Primitives

ICCV 2021poster

It is common practice to represent spoken languages at their phonetic level. However, for sign languages, this implies breaking motion into its constituent motion primitives. Avatar based Sign Language Production (SLP) has traditionally done just this, building up animation from sequences of hand mo…

Cited by 69PDFScholar
2021

VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of Experts

CVPR 2021poster

Disentangled representations support a range of downstream tasks including causal reasoning, generative modeling, and fair machine learning. Unfortunately, disentanglement has been shown to be impossible without the incorporation of supervision or inductive bias. Given that supervision is often expe…

Cited by 10PDFcodeScholar
2020

Progressive Transformers for End-to-End Sign Language Production

ECCV 2020poster

The goal of automatic Sign Language Production (SLP) is to translate spoken language to a continuous stream of sign language video at a level comparable to a human translator. If this was achievable, then it would revolutionise Deaf hearing communications. Previous work on predominantly isolated SLP…

2020

Sign Language Transformers: Joint End-to-End Sign Language Recognition and Translation

CVPR 2020oral

Prior work on Sign Language Translation has shown that having a mid-level sign gloss representation (effectively recognizing the individual signs) improves the translation performance drastically. In fact, the current state-of-the-art in translation requires gloss level tokenization in order to work…

Cited by 726PDFcodeScholar
2018

Neural Sign Language Translation

CVPR 2018poster

Sign Language Recognition (SLR) has been an active research field for the last two decades. However, most research to date has considered SLR as a naive gesture recognition problem. SLR seeks to recognize a sequence of continuous signs but neglects the underlying rich grammatical and linguistic stru…

2017

SubUNets: End-To-End Hand Shape and Continuous Sign Language Recognition

ICCV 2017spotlight

We propose a novel deep learning approach to solve simultaneous alignment and recognition problems (referred to as "Sequence-to-sequence" learning). We decompose the problem into a series of specialised expert systems referred to as SubUNets. The spatio-temporal relationships between these SubUNets…

Cited by 420PDFcodeScholar