← Search

Srijan Das

17 accepted papers

2026

MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos

CVPR 2026

Temporal Action Detection (TAD) in untrimmed videos poses significant challenges, particularly for Activities of Daily Living (ADL) requiring models to (1) process long-duration videos, (2) capture temporal variations in actions, and (3) simultaneously detect dense overlapping actions. Existing CNN

Cited by 0SourcecodeScholar
2025

GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology

ICCV 2025poster

Pretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. While recent multimodal MIL pretraining approaches leveraging auxiliary modalities have demonstrated performance gains ov…

2025

GenHMR: Generative Human Mesh Recovery

AAAI 2025technical

Human mesh recovery (HMR) is crucial in many computer vision applications; from health to arts and entertainment. HMR from monocular images has predominantly been addressed by deterministic methods that output a single prediction for a given 2D image. However, HMR from a single image is an ill-posed…

Cited by 0SourcePDFScholar
2025

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

CVPR 2025poster

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation s…

2025

MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild

ICCV 2025poster

Reconstructing a 3D hand mesh from a single RGB image is challenging due to complex articulations, self-occlusions, and depth ambiguities. Traditional discriminative methods, which learn a deterministic mapping from a 2D image to a single 3D mesh, often struggle with the inherent ambiguities in 2D-t…

2025

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living

AAAI 2025technical

The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, these models are typically trained on web videos, which often fail to capture the challenges present in Activities of Dai…

2024

BAMM: Bidirectional Autoregressive Motion Model

ECCV 2024poster

"Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion length. Conversely, autoregressive motion models address this…

2024

Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier

ECCV 2024poster

"Incorporating pixel contextual information is critical for accurate segmentation. In this paper, we show that an effective way to incorporate contextual information is through a patch-based classifier. This patch classifier is trained to identify classes present within an image region, which facili…

2024

Just Add ?! Pose Induced Video Transformers for Understanding Activities of Daily Living

CVPR 2024poster

Video transformers have become the de facto standard for human action recognition yet their exclusive reliance on the RGB modality still limits their adoption in certain domains. One such domain is Activities of Daily Living (ADL) where RGB alone is not sufficient to distinguish between visually sim…

2024

Multiview Aerial Visual RECognition (MAVREC): Can Multi-view Improve Aerial Visual Perception?

CVPR 2024poster

Despite the commercial abundance of UAVs aerial data acquisition remains challenging and the existing Asia and North America-centric open-source UAV datasets are small-scale or low-resolution and lack diversity in scene contextuality. Additionally the color content of the scenes solar zenith angle a…

Cited by 4SourcePDFScholar
2024

SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel Histopathology

CVPR 2024poster

Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging given the complexity of gigapixel slides. Traditionally MIL interpretability is limited to identifying salient regions deemed pertinent for downstream tasks of…

2022

Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?

NeurIPS 2022accept

We investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervis…

Cited by 35SourcePDFScholar
2022

Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space

NeurIPS 2022accept

Humans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a pool of 2D images often fail to generalize over novel camera viewpoints. Recently,…

2022

MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection

CVPR 2022poster

Action detection is an essential and challenging task, especially for densely labelled datasets of untrimmed videos. The temporal relation is complex in those datasets, including challenges like composite action, and co-occurring action. For detecting actions in those complex videos, efficiently cap…

Cited by 99PDFcodeScholar
2021

Learning an Augmented RGB Representation With Cross-Modal Knowledge Distillation for Action Detection

ICCV 2021poster

In video understanding, most cross-modal knowledge distillation (KD) methods are tailored for classification tasks, focusing on the discriminative representation of the trimmed videos. However, action detection requires not only categorizing actions, but also localizing them in untrimmed videos. The…

Cited by 49PDFScholar
2020

VPN: Learning Video-Pose Embedding for Activities of Daily Living

ECCV 2020poster

In this paper, we focus on the spatio-temporal aspect of recognizing Activities of Daily Living (ADL). ADL have two specific properties (i) subtle spatio-temporal patterns and (ii) similar visual patterns varying with time. Therefore, ADL may look very similar and often necessitate to look at their…

2019

Toyota Smarthome: Real-World Activities of Daily Living

ICCV 2019poster

The performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce…

Cited by 204PDFScholar