← Search

Srikrishna Karanam

23 accepted papers

2026

Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach

CVPR 2026

Contrastive vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition yet remain vulnerable to spurious correlations, particularly background over-reliance. We introduce Cluster-based Concept Importance (CCI), a novel interpretability method that uses CLIP's own patch embedding

Cited by 0SourceScholar
2026

Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts

AAAI 2026technical

Existing methods for human parsing into body parts and clothing often use fixed mask categories with broad labels that obscure fine-grained clothing types. Recent open-vocabulary segmentation approaches leverage pretrained text-to-image (T2I) diffusion model features for strong zero-shot transfer, b

Cited by 0SourcePDFScholar
2025

Composing Parts for Expressive Object Generation

CVPR 2025poster

Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, like Stable Diffusion, cannot handle fine-grained part-level attributes in the text prompts. Specifically, when additiona…

Cited by 0SourcePDFScholar
2025

TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction

CVPR 2025highlight

We consider the problem of single-source domain generalization. Existing methods typically rely on extensive augmentations to synthetically cover diverse domains during training. However, they struggle with semantic shifts (e.g., background and viewpoint changes), as they often learn global features…

Cited by 2SourcePDFScholar
2024

CoPL: Contextual Prompt Learning for Vision-Language Understanding

AAAI 2024technical

Recent advances in multimodal learning has resulted in powerful vision-language models, whose representations are generalizable across a variety of downstream tasks. Recently, their generalization ability has been further extended by incorporating trainable prompts, borrowed from the natural languag…

Cited by 7SourcePDFScholar
2024

SAFARI: Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation

ECCV 2024poster

"Referring Expression Segmentation (RES) aims to provide a segmentation mask of the target object in an image referred to by the text (i.e., referring expression). Existing methods require large-scale mask annotations. Moreover, such approaches do not generalize well to unseen/zero-shot scenarios. T…

2023

A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis

ICCV 2023poster

While recent developments in text-to-image generative models have led to a suite of high-performing methods capable of producing creative imagery from free-form text, there are several limitations. By analyzing the cross-attention representations of these models, we notice two key issues. First, for…

Cited by 44PDFScholar
2022

PseudoClick: Interactive Image Segmentation with Click Imitation

ECCV 2022poster

"The goal of click-based interactive image segmentation is to obtain precise object segmentation masks with limited user interaction, i.e., by a minimal number of user clicks. Existing methods require users to provide all the clicks: by first inspecting the segmentation mask and then providing point…

Cited by 69SourcePDFScholar
2022

SMPL-A: Modeling Person-Specific Deformable Anatomy

CVPR 2022poster

A variety of diagnostic and therapeutic protocols rely on locating in vivo target anatomical structures, which can be obtained from medical scans. However, organs move and deform as the patient changes his/her pose. In order to obtain accurate target location information, clinicians have to either c…

Cited by 12PDFScholar
2022

Self-supervised Human Mesh Recovery with Cross-Representation Alignment

ECCV 2022poster

"Fully supervised human mesh recovery methods are data-hungry and have poor generalizability due to the limited availability and diversity of 3D-annotated benchmark datasets. Recent progress in self-supervised human mesh recovery has been made using synthetic-data-driven training paradigms where the…

Cited by 15SourcePDFScholar
2021

A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual Concepts

CVPR 2021poster

Despite substantial progress in applying neural networks (NN) to a wide variety of areas, they still largely suffer from a lack of transparency and interpretability. While recent developments in explainable artificial intelligence attempt to bridge this gap (e.g., by visualizing the correlation betw…

Cited by 60PDFScholar
2021

Ensemble Attention Distillation for Privacy-Preserving Federated Learning

ICCV 2021poster

We consider the problem of Federated Learning (FL) where numerous decentralized computational nodes collaborate with each other to train a centralized machine learning model without explicitly sharing their local data samples. Such decentralized training naturally leads to issues of imbalanced or di…

Cited by 148PDFScholar
2021

Spatio-Temporal Representation Factorization for Video-Based Person Re-Identification

ICCV 2021poster

Despite much recent progress in video-based person re-identification (re-ID), the current state-of-the-art still suffers from common real-world challenges such as appearance similarity among various people, occlusions, and frame misalignment. To alleviate these problems, we propose Spatio-Temporal R…

Cited by 93PDFScholar
2020

Hierarchical Kinematic Human Mesh Recovery

ECCV 2020poster

We consider the problem of estimating a parametric model of 3D human mesh from a single image. While there has been substantial recent progress in this area with direct regression of model parameters, these methods only implicitly exploit the human body kinematic structure, leading to sub-optimal us…

Cited by 128SourcePDFScholar
2020

Towards Visually Explaining Variational Autoencoders

CVPR 2020oral

Recent advances in Convolutional Neural Network (CNN) model interpretability have led to impressive progress in visualizing and understanding model predictions. In particular, gradient-based visual attention methods have driven much recent effort in using visual attention maps as a means for visual…

Cited by 301PDFcodeScholar
2019

Incremental Scene Synthesis

NeurIPS 2019poster

We present a method to incrementally generate complete 2D or 3D scenes with the following properties: (a) it is globally consistent at each step according to a learned scene prior, (b) real observations of a scene can be incorporated while observing global consistency, (c) unobserved regions can be…

Cited by 9SourcePDFScholar
2019

Learning Local RGB-to-CAD Correspondences for Object Pose Estimation

ICCV 2019poster

We consider the problem of 3D object pose estimation. While much recent work has focused on the RGB domain, the reliance on accurately annotated images limits generalizability and scalability. On the other hand, the easily available object CAD models are rich sources of data, providing a large numbe…

Cited by 30PDFScholar
2019

Sharpen Focus: Learning With Attention Separability and Consistency

ICCV 2019poster

Recent developments in gradient-based attention modeling have seen attention maps emerge as a powerful tool for interpreting convolutional neural networks. Despite good localization for an individual class of interest, these techniques produce attention maps with substantially overlapping responses…

Cited by 41PDFScholar
2018

End-to-End Learning of Keypoint Detector and Descriptor for Pose Invariant 3D Matching

CVPR 2018poster

Finding correspondences between images or 3D scans is at the heart of many computer vision and image retrieval applications and is often enabled by matching local keypoint descriptors. Various learning approaches have been applied in the past to different stages of the matching pipeline, considering…

Cited by 71SourcePDFScholar
2018

Learning Compositional Visual Concepts With Mutual Consistency

CVPR 2018poster

Compositionality of semantic concepts in image synthesis and analysis is appealing as it can help in decomposing known and generatively recomposing unknown data. For instance, we may learn concepts of changing illumination, geometry or albedo of a scene, and try to recombine them to generate physica…

Cited by 17SourcePDFScholar
2015

Person Re-Identification With Discriminatively Trained Viewpoint Invariant Dictionaries

ICCV 2015poster

This paper introduces a new approach to address the person re-identification problem in cameras with non-overlapping fields of view. Unlike previous approaches that learn Mahalanobis-like distance metrics in some transformed feature space, we propose to learn a dictionary that is capable of discrimi…

Cited by 259PDFScholar