← Search

Zijia Lu

7 accepted papers

2025

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

CVPR 2025poster

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challen…

2024

Error Detection in Egocentric Procedural Task Videos

CVPR 2024poster

We present a new egocentric procedural error dataset containing videos with various types of errors as well as normal videos and propose a new framework for procedural error detection using error-free training videos only. Our framework consists of an action segmentation model and a contrastive step…

Cited by 15SourcePDFScholar
2024

FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action Segmentation

CVPR 2024poster

We study supervised action segmentation whose goal is to predict framewise action labels of a video. To capture temporal dependencies over long horizons prior works either improve framewise features with transformer or refine framewise predictions with learned action features. However they are compu…

2024

Self-Supervised Multi-Object Tracking with Path Consistency

CVPR 2024highlight

In this paper we propose a novel concept of path consistency to learn robust object matching without using manual object identity supervision. Our key idea is that to track a object through frames we can obtain multiple different association results from a model by varying the frames it can observe…

2021

Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces Learning

ICCV 2021poster

We address the problem of learning to segment actions from weakly-annotated videos, i.e., videos accompanied by transcripts (ordered list of actions). We propose a framework in which we model actions with a union of low-dimensional subspaces, learn the subspaces using transcripts and refine video fe…

Cited by 28PDFcodeScholar