← Search

Arkaprava Sinha

3 accepted papers

2026

MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos

CVPR 2026

Temporal Action Detection (TAD) in untrimmed videos poses significant challenges, particularly for Activities of Daily Living (ADL) requiring models to (1) process long-duration videos, (2) capture temporal variations in actions, and (3) simultaneously detect dense overlapping actions. Existing CNN

Cited by 0SourcecodeScholar
2025

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

CVPR 2025poster

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation s…

2025

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living

AAAI 2025technical

The introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, these models are typically trained on web videos, which often fail to capture the challenges present in Activities of Dai…