← Search

Deepti Ghadiyaram

16 accepted papers

2025

Revelio: Interpreting and leveraging semantic information in diffusion models

ICCV 2025poster

We study how rich visual semantic information is represented within various layers and denoising timesteps of different diffusion architectures. We uncover monosemantic interpretable features by leveraging k-sparse autoencoders (k-SAE). We substantiate our mechanistic interpretations via transfer le…

2025

What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization

ICCV 2025poster

Domain Generalization aims to develop models that can generalize to novel and unseen data distributions. In this work, we study how model architectures and pre-training objectives impact feature richness and propose a method to effectively leverage them for domain generalization. Specifically, given…

2023

GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition

NeurIPS 2023poster

Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North Amer…

Cited by 34SourcePDFScholar
2022

Making Heads or Tails: Towards Semantically Consistent Visual Counterfactuals

ECCV 2022poster

"A visual counterfactual explanation replaces image regions in a query image with regions from a distractor image such that the system’s decision on the transformed image changes to the distractor class. In this work, we present a novel framework for computing visual counterfactual explanations base…

2021

Generic Event Boundary Detection: A Benchmark for Event Segmentation

ICCV 2021poster

This paper presents a novel task together with a new benchmark for detecting generic, taxonomy-free event boundaries that segment a whole video into chunks. Conventional work in temporal video segmentation and action detection focuses on localizing pre-defined action categories and thus does not sca…

Cited by 84PDFcodeScholar
2021

How2Sign: A Large-Scale Multimodal Dataset for Continuous American Sign Language

CVPR 2021poster

One of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous American Sign Language (ASL) dataset, consisting of a paral…

Cited by 257PDFcodeScholar
2020

ClusterFit: Improving Generalization of Visual Representations

CVPR 2020poster

Pre-training convolutional neural networks with weakly-supervised and self-supervised strategies is becoming increasingly popular for several computer vision tasks. However, due to the lack of strong discriminative signals, these learned representations may overfit to the pre-training objective (e.g…

Cited by 167PDFcodeScholar
2020

Don't Judge an Object by Its Context: Learning to Overcome Contextual Bias

CVPR 2020oral

Existing models often leverage co-occurrences between objects and their context to improve recognition accuracy. However, strongly relying on context risks a model's generalizability, especially when typical co-occurrence patterns are absent. This work focuses on addressing such contextual biases to…

Cited by 141PDFScholar
2020

From Patches to Pictures (PaQ-2-PiQ): Mapping the Perceptual Space of Picture Quality

CVPR 2020poster

Blind or no-reference (NR) perceptual picture quality prediction is a difficult, unsolved problem of great consequence to the social and streaming media industries that impacts billions of viewers daily. Unfortunately, popular NR prediction models perform poorly on real-world distorted pictures. To…

Cited by 397PDFScholar
2019

Activity Driven Weakly Supervised Object Detection

CVPR 2019poster

Weakly supervised object detection aims at reducing the amount of supervision required to train detection models. Such models are traditionally learned from images/videos labelled only with the object class and not the object bounding box. In our work, we try to leverage not only the object class la…

Cited by 40PDFScholar
2019

Large-Scale Weakly-Supervised Pre-Training for Video Action Recognition

CVPR 2019poster

Current fully-supervised video datasets consist of only a few hundred thousand videos and fewer than a thousand domain-specific labels. This hinders the progress towards advanced video architectures. This paper presents an in-depth study of using large volumes of web videos for pre-training video mo…

Cited by 391PDFcodeScholar
2019

Less Is More: Learning Highlight Detection From Video Duration

CVPR 2019poster

Highlight detection has the potential to significantly ease video browsing, but existing methods often suffer from expensive supervision requirements, where human viewers must manually identify highlights in training videos. We propose a scalable unsupervised solution that exploits video duration as…

Cited by 159PDFScholar
2017

Subjective and objective quality assessment of Mobile Videos with In-Capture distortions

ICASSP 2017accepted

We designed and created a new video database that models a variety of complex distortions generated during the video capturing process on hand-held mobile capturing devices. We describe the content and characteristics of the new database, which we call the LIVE Mobile In-Capture Video Quality Databa…

Cited by 0SourceScholar