← Search

Deepu Rajan

10 accepted papers

2026

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

AAAI 2026technical

Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowl

Cited by 0SourcePDFScholar
2025

Chebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute Loss

CVPR 2025poster

Despite recent advances in deep texture recognition, existing methods still lack representational diversity and struggle to capture and preserve discriminative cues across stages of representation hierarchies. Moreover, many rely on loss formulations that prioritize recognition accuracy while overlo…

2025

IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A

ICCV 2025poster

Existing human motion Q&A methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To overcome this, we propose an implicit program-guided motion reasoning (IMoRe) framework that unifies reasoning across mul…

2025

Intra-modal and Cross-modal Synchronization for Audio-visual Deepfake Detection and Temporal Localization

ICCV 2025poster

Recent deepfake detection algorithms focus solely on uni-modal or cross-modal inconsistencies. While the former disregards audio-visual correspondence entirely rendering them less effective against multimodal attacks, the latter overlooks inconsistencies in a particular modality. Moreover, many mode…

2024

Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake Detection

ICASSP 2024accepted

Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and inconsistencies in learned representations caused by independent modali…

Cited by 0SourceScholar
2023

UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning

ACL 2023findings

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality relationship, treat each modality equally, suffer sensor noise, and…

2022

Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information

ICASSP 2022accepted

Speech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition…

Cited by 0SourceScholar