← Search

Mihir Jain

8 accepted papers

2025

UVGS: Reimagining Unstructured 3D Gaussian Splatting using UV Mapping

CVPR 2025poster

3D Gaussian Splatting (3DGS) has demonstrated superior quality in modeling 3D objects and scenes. However, generating 3DGS remains challenging due to their discrete, unstructured, and permutation-invariant nature. In this work, we present a simple yet effective method to overcome these challenges. W…

Cited by 2SourcePDFScholar
2023

Few-Shot Common Action Localization via Cross-Attentional Fusion of Context and Temporal Dynamics

ICCV 2023poster

The goal of this paper is to localize action instances in a long untrimmed query video using just meager trimmed support videos representing a common action whose class information is not given. In this task, it is crucial to mine reliable temporal cues representing a common action from handful supp…

Cited by 7PDFScholar
2021

Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization

ICLR 2021poster

Temporally localizing actions in videos is one of the key components for video understanding. Learning from weakly-labeled data is seen as a potential solution towards avoiding expensive frame-level annotations. Different from other works which only depend on visual-modality, we propose to learn ric…

Cited by 77SourcePDFScholar
2021

Efficient Action Recognition via Dynamic Knowledge Propagation

ICCV 2021poster

Efficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To t…

Cited by 30PDFScholar
2021

Motion-Augmented Self-Training for Video Recognition at Smaller Scale

ICCV 2021poster

The goal of this paper is to self-train a 3D convolutional neural network on an unlabeled video collection for deployment on small-scale video collections. As smaller video datasets benefit more from motion than appearance, we strive to train our network using optical flow, but avoid its computation…

Cited by 24PDFScholar
2015

Objects2action: Classifying and Localizing Actions Without Any Video Example

ICCV 2015poster

The goal of this paper is to recognize actions in video without the need for examples. Different from traditional zero-shot approaches we do not demand the design and specification of attribute classifiers and class-to-attribute mappings to allow for transfer from seen classes to unseen classes. Our…

Cited by 189PDFScholar
2015

What do 15,000 Object Categories Tell Us About Classifying and Localizing Actions?

CVPR 2015poster

This paper contributes to automatic classification and localization of human actions in video. Whereas motion is the key ingredient in modern approaches, we assess the benefits of having objects in the video representation. Rather than considering a handful of carefully selected and localized object…

Cited by 225SourcePDFScholar