← Search

Anirban Chakraborty

12 accepted papers

2026

O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model

AAAI 2026technical

While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, they struggle to comprehend hand-drawn sketches, a modality that offers an intuitive means of expressing concepts that are

Cited by 0SourcePDFScholar
2026

Probabilistically-routed Bayesian Additive Spanning Trees for Learning on Constrained Domains

ICML 2026poster

Bayesian additive spanning tree (BAST) is an useful tool for interpretable, non-parametric regression on complex constrained domains. It improves upon the performance of Bayesian additive regression trees (BART) by replacing axis-aligned splits through binary tree components by cuts on a spanning tr…

Cited by 0SourceScholar
2026

Prompt Estimation from Prototypes for Federated Prompt Tuning of Vision Transformers

ICML 2026poster

Visual Prompt Tuning (VPT) of pre-trained Vision Transformers (ViTs) has proven highly effective as a parameter-efficient fine-tuning technique for adapting large models to downstream tasks with limited data. Its parameter efficiency makes it particularly suitable for Federated Learning (FL), where …

Cited by 0SourceScholar
2024

Dynamic Data Selection for Efficient SSL via Coarse-to-Fine Refinement

ECCV 2024poster

"Self-supervised learning (SSL) is critical for learning high-quality representations from unlabeled images at scale. Earlier efforts at reducing the compute requirements of SSL have focused on identifying subsets of training data that are sufficient for training. In addition to using a static repre…

Cited by 0SourcePDFScholar
2023

Edges to Shapes to Concepts: Adversarial Augmentation for Robust Vision

CVPR 2023poster

Recent work has shown that deep vision models tend to be overly dependent on low-level or "texture" features, leading to poor generalization. Various data augmentation strategies have been proposed to overcome this so-called texture bias in DNNs. We propose a simple, lightweight adversarial augmenta…

Cited by 7SourcePDFScholar
2022

Uncertainty-Aware Adaptation for Self-Supervised 3D Human Pose Estimation

CVPR 2022poster

The advances in monocular 3D human pose estimation are dominated by supervised techniques that require large-scale 2D/3D pose annotations. Such methods often behave erratically in the absence of any provision to discard unfamiliar out-of-distribution data. To this end, we cast the 3D human pose lear…

Cited by 45PDFScholar
2021

Non-local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation

NeurIPS 2021poster

Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target environment is highly inconvenient. To this end, we cast 3D pose lear…

Cited by 12SourcePDFScholar
2020

Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis

CVPR 2020oral

Camera captured human pose is an outcome of several sources of variation. Performance of supervised 3D pose estimation approaches comes at the cost of dispensing with variations, such as shape and appearance, that may be useful for solving other related tasks. As a result, the learned model not only…

Cited by 109PDFScholar
2019

All for One: Frame-wise Rank Loss for Improving Video-based Person Re-identification

ICASSP 2019accepted

Person re-identification involves retrieving correct matches for a target image (query) from a set of gallery images, while video based re-identification extends this to the case of query and gallery videos. Typical video-based re-id methods ignore the temporal evolution of the intermediate represen…

Cited by 0SourceScholar
2019

From Strings to Things: Knowledge-Enabled VQA Model That Can Read and Reason

ICCV 2019oral

Text present in images are not merely strings, they provide useful cues about the image. Despite their utility in better image understanding, scene texts are not used in traditional visual question answering (VQA) models. In this work, we present a VQA model which can read scene texts and perform re…

Cited by 64PDFScholar