← Search

Antoine Yang

10 accepted papers

2026

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

CVPR 2026

Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language.Due to the complexity of the task and the high cost associated with manual

Cited by 1SourceScholar
2026

Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos

CVPR 2026

Multimodal AI agents are increasingly automating complex real-world workflows that involve online web execution. However, current web-agent benchmarks suffer from a critical limitation: they focus entirely on web-based interaction and perception, lacking grounding in the user's real-world physical s

Cited by 0SourcecodeScholar
2025

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

CVPR 2025poster

We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, automatic chaptering has the potential to enable efficient navigation and content retrieval in long-form videos. In this…

Cited by 0SourcePDFScholar
2024

CoVR: Learning Composed Video Retrieval from Web Video Captions

AAAI 2024technical

Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database. Most CoIR approaches require manually annotated datasets, comprising image-text-image triplets, where the text describes a modifi…

Cited by 46SourcePDFScholar
2023

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

CVPR 2023poster

In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architecture augments a language model with special time tokens, allowing it to seamlessly predict event boundaries and textual…

2023

VidChapters-7M: Video Chapters at Scale

NeurIPS 2023poster

Segmenting untrimmed videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly released datasets. To address this issue, we present VidChapters-7M, a dataset of 817K user-chaptered videos includ…

Cited by 36SourcePDFScholar
2022

TubeDETR: Spatio-Temporal Video Grounding With Transformers

CVPR 2022oral

We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal interactions. To address this task, we propose TubeDETR, a transformer-based arc…

Cited by 120PDFcodeScholar
2022

Zero-Shot Video Question Answering via Frozen Bidirectional Language Models

NeurIPS 2022accept

Video question answering (VideoQA) is a complex task that requires diverse multi-modal data for training. Manual annotation of question and answers for videos, however, is tedious and prohibits scalability. To tackle this problem, recent methods consider zero-shot settings with no manual annotation…

2021

Just Ask: Learning To Answer Questions From Millions of Narrated Videos

ICCV 2021poster

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for vid…

Cited by 352PDFcodeScholar