← Search

Du Tran

18 accepted papers

2025

SEAL: Semantic Attention Learning for Long Video Representation

CVPR 2025poster

Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces **S…

Cited by 0SourcePDFScholar
2023

Relational Space-Time Query in Long-Form Videos

CVPR 2023highlight

Egocentric videos are often available in the form of uninterrupted, uncurated long videos capturing the camera wearers' daily life activities.Understanding these videos requires models to be able to reason about activities, objects, and their interactions. However, current video benchmarks study the…

Cited by 14SourcePDFScholar
2022

Open-World Instance Segmentation: Exploiting Pseudo Ground Truth From Learned Pairwise Affinity

CVPR 2022poster

Open-world instance segmentation is the task of grouping pixels into object instances without any pre-determined taxonomy. This is challenging, as state-of-the-art methods rely on explicit class semantics obtained from large labeled datasets, and out-of-domain evaluation performance drops significan…

Cited by 54PDFcodeScholar
2020

Self-Supervised Learning by Cross-Modal Audio-Video Clustering

NeurIPS 2020spotlight

Visual and audio modalities are highly correlated, yet they contain different information. Their strong correlation makes it possible to predict the semantics of one from the other with good accuracy. Their intrinsic differences make cross-modal prediction a potentially more rewarding pretext task f…

2019

DistInit: Learning Video Representations Without a Single Labeled Video

ICCV 2019poster

Video recognition models have progressed significantly over the past few years, evolving from shallow classifiers trained on hand-crafted features to deep spatiotemporal networks. However, labeled video data required to train such models has not been able to keep up with the ever increasing depth an…

Cited by 75PDFScholar
2019

Large-Scale Weakly-Supervised Pre-Training for Video Action Recognition

CVPR 2019poster

Current fully-supervised video datasets consist of only a few hundred thousand videos and fewer than a thousand domain-specific labels. This hinders the progress towards advanced video architectures. This paper presents an in-depth study of using large volumes of web videos for pre-training video mo…

Cited by 391PDFcodeScholar
2019

Learning Temporal Pose Estimation from Sparsely-Labeled Videos

NeurIPS 2019poster

Modern approaches for multi-person pose estimation in video require large amounts of dense annotations. However, labeling every frame in a video is costly and labor intensive. To reduce the need for dense annotations, we propose a PoseWarper network that leverages training videos with sparse annotat…

2019

Video Classification With Channel-Separated Convolutional Networks

ICCV 2019poster

Group convolution has been shown to offer great computational savings in various 2D convolutional architectures for image classification. It is natural to ask: 1) if group convolution can help to alleviate the high computational cost of video classification networks; 2) what factors matter the most…

Cited by 784PDFcodeScholar
2018

A Closer Look at Spatiotemporal Convolutions for Action Recognition

CVPR 2018poster

In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition. Our motivation stems from the observation that 2D CNNs applied to individual frames of the video have remained solid performers in action recognition. In this work…

2018

Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization

NeurIPS 2018poster

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal synchronization. We demonstrate that a calibrated curriculum learning sch…

Cited by 570SourcePDFScholar
2018

Detect-and-Track: Efficient Pose Estimation in Videos

CVPR 2018poster

This paper addresses the problem of estimating and tracking human body keypoints in complex, multi-person video. We propose an extremely lightweight yet highly effective approach that builds upon the latest advancements in human detection and video understanding. Our method operates in two-stages: k…

Cited by 315SourcePDFScholar
2018

Scenes-Objects-Actions: A Multi-Task, Multi-Label Video Dataset

ECCV 2018poster

This paper introduces a large-scale, multi-label and multitask video dataset named Scenes-Objects-Actions (SOA). Most prior video datasets are based on a predened taxonomy, which is used to de- ne the keyword queries issued to search engines. The videos retrieved by the search engines are then verie…

Cited by 38SourcePDFScholar
2015

Learning Spatiotemporal Features With 3D Convolutional Networks

ICCV 2015poster

We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset. Our findings are three-fold: 1) 3D ConvNets are more suitable for spatiotemporal feature learning compared…

Cited by 11362PDFcodeScholar