← Search

Subarna Tripathi

10 accepted papers

2024

Action Scene Graphs for Long-Form Understanding of Egocentric Videos

CVPR 2024poster

We present Egocentric Action Scene Graphs (EASGs) a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos such as verb-noun action labels by providing a temporally evolving graph-based description of the act…

2023

SViTT: Temporal Learning of Sparse Video-Text Transformers

CVPR 2023poster

Do video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has revealed the strong tendency of video-text models towards frame-based spatial representations, while temporal reasoning remain…

2022

Joint Hand Motion and Interaction Hotspots Prediction From Egocentric Videos

CVPR 2022poster

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e., interaction hotspots). This relatively low-dimensional repres…

Cited by 105PDFcodeScholar
2022

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

ECCV 2022poster

"Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a novel spatial-temporal graph learning framework that can solv…

2022

Single-Stage Visual Relationship Learning using Conditional Queries

NeurIPS 2022accept

Research in scene graph generation (SGG) usually considers two-stage models, that is, detecting a set of entities, followed by combining them and labeling all possible relationships. While showing promising results, the pipeline structure induces large parameter and computation overhead, and typical…

Cited by 9SourcePDFScholar
2021

In Defense of Scene Graphs for Image Captioning

ICCV 2021poster

The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their structural semantics such as object entities, relationships and…

Cited by 53PDFcodeScholar
2019

PartNet: A Large-Scale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding

CVPR 2019poster

We present PartNet: a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information. Our dataset consists of 573,585 part instances over 26,671 3D models covering 24 object categories. This dataset enables and serves as a catalyst for…

Cited by 849PDFScholar