← Search

Akash Kumar

10 accepted papers

2025

Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding

ICLR 2025poster

In this work, we focus on Weakly Supervised Spatio-Temporal Video Grounding (WSTVG). It is a multimodal task aimed at localizing specific subjects spatio-temporally based on textual queries without bounding box supervision. Motivated by recent advancements in multi-modal foundation models for groun…

Cited by 2SourcePDFScholar
2025

MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective

ICCV 2025poster

In this paper, we propose MixA-Q, a mixed-precision activation quantization framework that leverages intra-layer activation sparsity (a concept widely explored in activation pruning methods) for efficient inference of quantized window-based vision transformers. For a given uniform-bit quantization c…

Cited by 0SourcePDFScholar
2025

STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding

CVPR 2025poster

In this work, we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Inspired by recent advances in vision-language foundation models, we investigate their u…

Cited by 1SourcePDFScholar
2025

Stable Mean Teacher for Semi-supervised Video Action Detection

AAAI 2025technical

In this work, we focus on semi-supervised learning for video action detection. Video action detection requires spatio-temporal localization in addition to classification, and a limited amount of labels makes the model prone to unreliable predictions. We present Stable Mean Teacher, a simple end-to-e…

2024

Semi-supervised Active Learning for Video Action Detection

AAAI 2024technical

In this work, we focus on label efficient learning for video action detection. We develop a novel semi-supervised active learning approach which utilizes both labeled as well as un- labeled data along with informative sample selection for ac- tion detection. Video action detection requires spatio-te…

2021

Teaching via Best-Case Counterexamples in the Learning-with-Equivalence-Queries Paradigm

NeurIPS 2021poster

We study the sample complexity of teaching, termed as "teaching dimension" (TD) in the literature, for the learning-with-equivalence-queries (LwEQ) paradigm. More concretely, we consider a learner who asks equivalence queries (i.e., "is the queried hypothesis the target hypothesis?"), and a teacher…

Cited by 1SourcePDFScholar