← Search

Sirnam Swetha

5 accepted papers

2026

SMPRO: Self-Supervised Visual Preference Alignment via Differentiable Multi-Preference Multi-Group Ranking

AAAI 2026technical

Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning models with human preferences. However, existing DPO-based methods suffer from 3 key drawbacks: they rely on only a single positive-negative preference pair per question, restricting the diversity and ri

Cited by 0SourcePDFScholar
2026

VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

CVPR 2026

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable with

Cited by 0SourcecodeScholar
2025

GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space

ICCV 2025poster

Timestamp prediction aims to determine when an image was captured using only visual information, supporting applications such as metadata correction, retrieval, and digital forensics. In outdoor scenarios, hourly estimates rely on cues like brightness, hue, and shadow positioning, while seasonal cha…

Cited by 0SourcePDFScholar
2023

Preserving Modality Structure Improves Multi-Modal Learning

ICCV 2023poster

Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human annotations. These joint embeddings enable zero-shot cross-modal tasks like retrieval and classification. However, thes…

Cited by 6PDFcodeScholar