← Search

Ishan Rajendrakumar Dave

10 accepted papers

2026

Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding

ICLR 2026poster

We introduce a novel formulation of visual privacy preservation for video foundation models that operates entirely in the latent space. While spatio-temporal features learned by foundation models have deepened general understanding of video content, sharing or storing these extracted visual features…

Cited by 0SourceScholar
2025

ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition

ICLR 2025poster

Bias in machine learning models can lead to unfair decision making, and while it has been well-studied in the image and text domains, it remains underexplored in action recognition. Action recognition models often suffer from background bias (i.e., inferring actions based on background cues) and for…

Cited by 0SourcePDFScholar
2025

From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos

NeurIPS 2025poster

Composed Video Retrieval (CoVR) retrieves a target video given a query video and a modification text describing the intended change. Existing CoVR benchmarks emphasize appearance shifts or coarse event changes and therefore do not test the ability to capture subtle, fast-paced temporal differences.…

Cited by 0SourcecodeScholar
2025

GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space

ICCV 2025poster

Timestamp prediction aims to determine when an image was captured using only visual information, supporting applications such as metadata correction, retrieval, and digital forensics. In outdoor scenarios, hourly estimates rely on cues like brightness, hue, and shadow positioning, while seasonal cha…

Cited by 0SourcePDFScholar
2024

No More Shortcuts: Realizing the Potential of Temporal Self-Supervision

AAAI 2024technical

Self-supervised approaches for video have shown impressive results in video understanding tasks. However, unlike early works that leverage temporal self-supervision, current state-of-the-art methods primarily rely on tasks from the image domain (e.g., contrastive learning) that do not explicitly pro…

2023

EventTransAct: A Video Transformer-Based Framework for Event-Camera Based Action Recognition

IROS 2023poster

Recognizing and comprehending human actions and gestures is a crucial perception requirement for robots to interact with humans and carry out tasks in diverse domains, including service robotics, healthcare, and manufacturing. Event cameras, with their ability to capture fast-moving objects at a hig…

Cited by 15SourceScholar
2023

TeD-SPAD: Temporal Distinctiveness for Self-Supervised Privacy-Preservation for Video Anomaly Detection

ICCV 2023poster

Video anomaly detection (VAD) without human monitoring is a complex computer vision task that can have a positive impact on society if implemented successfully. While recent advances have made significant progress in solving this task, most existing approaches overlook a critical real-world concern:…

Cited by 28PDFcodeScholar
2023

TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition

CVPR 2023poster

Semi-Supervised Learning can be more beneficial for the video domain compared to images because of its higher annotation cost and dimensionality. Besides, any video understanding task requires reasoning over both spatial and temporal dimensions. In order to learn both the static and motion related f…

2023

TransVisDrone: Spatio-Temporal Transformer for Vision-based Drone-to-Drone Detection in Aerial Videos

ICRA 2023poster

Drone-to-drone detection using visual feed has crucial applications, such as detecting drone collisions, detecting drone attacks, or coordinating flight with other drones. However, existing methods are computationally costly, follow non-end-to-end optimization, and have complex multi-stage pipelines…

Cited by 36SourcecodeScholar
2022

SPAct: Self-Supervised Privacy Preservation for Action Recognition

CVPR 2022poster

Visual private information leakage is an emerging key issue for the fast growing applications of video understanding like activity recognition. Existing approaches for mitigating privacy leakage in action recognition require privacy labels along with the action labels from the video dataset. However…

Cited by 69PDFcodeScholar