← Search

Ayan Kumar Bhunia

39 accepted papers

2025

Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection

ICCV 2025poster

Keypoint detection, integral to modern machine perception, faces challenges in few-shot learning, particularly when source data from the same distribution as the query is unavailable. This gap is addressed by leveraging sketches, a popular form of human expression, providing a source-free alternativ…

Cited by 0SourcePDFScholar
2025

Sketch Down the FLOPs: Towards Efficient Networks for Human Sketch

CVPR 2025poster

As sketch research has collectively matured over time, its adaptation for at-mass commercialisation emerges on the immediate horizon. Despite an already mature research endeavour for photos, there is no research on the efficient inference specifically designed for sketch data. In this paper, we firs…

Cited by 0SourcePDFScholar
2025

SketchFusion: Learning Universal Sketch Features through Fusing Foundation Models

CVPR 2025poster

While foundation models have revolutionised computer vision, their effectiveness for sketch understanding remains limited by the unique challenges of abstract, sparse visual inputs. Through systematic analysis, we uncover two fundamental limitations: Stable Diffusion (SD) struggles to extract meanin…

Cited by 0SourcePDFScholar
2024

DemoCaricature: Democratising Caricature Generation with a Rough Sketch

CVPR 2024poster

In this paper we democratise caricature generation empowering individuals to effortlessly craft personalised caricatures with just a photo and a conceptual sketch. Our objective is to strike a delicate balance between abstraction and identity while preserving the creativity and subjectivity inherent…

Cited by 9SourcePDFScholar
2024

Do Generalised Classifiers really work on Human Drawn Sketches?

ECCV 2024poster

"This paper, for the first time, marries large foundation models with human sketch understanding. We demonstrate what this brings – a paradigm shift in terms of generalised sketch representation learning (e.g., classification). This generalisation happens on two fronts: (i) generalisation across unk…

2024

Doodle Your 3D: From Abstract Freehand Sketches to Precise 3D Shapes

CVPR 2024poster

In this paper we democratise 3D content creation enabling precise generation of 3D shapes from abstract sketches while overcoming limitations tied to drawing skills. We introduce a novel part-level modelling and alignment framework that facilitates abstraction modelling and cross-modal correspondenc…

2024

Freeview Sketching: View-Aware Fine-Grained Sketch-Based Image Retrieval

ECCV 2024poster

"In this paper, we delve into the intricate dynamics of Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) by addressing a critical yet overlooked aspect – the choice of viewpoint during sketch creation. Unlike photo systems that seamlessly handle diverse views through extensive datasets, sketch sy…

Cited by 0SourcePDFScholar
2024

How to Handle Sketch-Abstraction in Sketch-Based Image Retrieval?

CVPR 2024poster

In this paper we propose a novel abstraction-aware sketch-based image retrieval framework capable of handling sketch abstraction at varied levels. Prior works had mainly focused on tackling sub-factors such as drawing style and order we instead attempt to model abstraction as a whole and propose fea…

Cited by 16SourcePDFScholar
2024

It's All About Your Sketch: Democratising Sketch Control in Diffusion Models

CVPR 2024poster

This paper unravels the potential of sketches for diffusion models addressing the deceptive promise of direct sketch control in generative AI. We importantly democratise the process enabling amateur sketches to generate precise images living up to the commitment of "what you sketch is what you get".…

2024

SketchINR: A First Look into Sketches as Implicit Neural Representations

CVPR 2024poster

We propose SketchINR to advance the representation of vector sketches with implicit neural models. A variable length vector sketch is compressed into a latent space of fixed dimension that implicitly encodes the underlying shape as a function of time and strokes. The learned function predicts the xy…

2024

Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers

CVPR 2024poster

This paper for the first time explores text-to-image diffusion models for Zero-Shot Sketch-based Image Retrieval (ZS-SBIR). We highlight a pivotal discovery: the capacity of text-to-image diffusion models to seamlessly bridge the gap between sketches and photos. This proficiency is underpinned by th…

Cited by 11SourcePDFScholar
2024

What Sketch Explainability Really Means for Downstream Tasks?

CVPR 2024poster

In this paper we explore the unique modality of sketch for explainability emphasising the profound impact of human strokes compared to conventional pixel-oriented studies. Beyond explanations of network behavior we discern the genuine implications of explainability across diverse downstream sketch-r…

Cited by 4SourcePDFScholar
2024

You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image Retrieval

CVPR 2024poster

Two primary input modalities prevail in image retrieval: sketch and text. While text is widely used for inter-category retrieval tasks sketches have been established as the sole preferred modality for fine-grained image retrieval due to their ability to capture intricate visual details. In this pape…

Cited by 16SourcePDFScholar
2023

CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or Not

CVPR 2023poster

In this paper, we leverage CLIP for zero-shot sketch based image retrieval (ZS-SBIR). We are largely inspired by recent advances on foundation models and the unparalleled generalisation ability they seem to offer, but for the first time tailor it to benefit the sketch community. We put forward novel…

2023

Democratising 2D Sketch to 3D Shape Retrieval Through Pivoting

ICCV 2023poster

This paper studies the problem of 2D sketch to 3D shape retrieval, but with a focus on democratising the process. We would like this democratisation to happen on two fronts: (i) to remove the need for large-scale specifically sourced 2D sketch and 3D shape datasets, and (ii) to remove restrictions o…

Cited by 6PDFScholar
2023

Exploiting Unlabelled Photos for Stronger Fine-Grained SBIR

CVPR 2023poster

This paper advances the fine-grained sketch-based image retrieval (FG-SBIR) literature by putting forward a strong baseline that overshoots prior state-of-the art by 11%. This is not via complicated design though, but by addressing two critical issues facing the community (i) the gold standard trip…

2023

Picture That Sketch: Photorealistic Image Generation From Abstract Sketches

CVPR 2023poster

Given an abstract, deformed, ordinary sketch from untrained amateurs like you and me, this paper turns it into a photorealistic image - just like those shown in Fig. 1(a), all non-cherry-picked. We differ significantly from prior art in that we do not dictate an edgemap-like sketch to start with, bu…

2023

SceneTrilogy: On Human Scene-Sketch and Its Complementarity With Photo and Text

CVPR 2023poster

In this paper, we extend scene understanding to include that of human sketch. The result is a complete trilogy of scene representation from three diverse and complementary modalities -- sketch, photo, and text. Instead of learning a rigid three-way embedding and be done with it, we focus on learning…

Cited by 32SourcePDFScholar
2023

Sketch2Saliency: Learning To Detect Salient Objects From Human Drawings

CVPR 2023poster

Human sketch has already proved its worth in various visual understanding tasks (e.g., retrieval, segmentation, image-captioning, etc). In this paper, we reveal a new trait of sketches -- that they are also salient. This is intuitive as sketching is a natural attentive process at its core. More spec…

Cited by 25SourcePDFScholar
2023

What Can Human Sketches Do for Object Detection?

CVPR 2023poster

Sketches are highly expressive, inherently capturing subjective and fine-grained visual cues. The exploration of such innate properties of human sketches has, however, been limited to that of image retrieval. In this paper, for the first time, we cultivate the expressiveness of sketches but for the…

2022

Adaptive Fine-Grained Sketch-Based Image Retrieval

ECCV 2022poster

"The recent focus on Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) has shifted towards generalising a model to new categories without any training data from them. In real-world applications, however, a trained FG-SBIR model is often applied to both new categories and different human sketchers,…

2022

Doodle It Yourself: Class Incremental Learning by Drawing a Few Sketches

CVPR 2022poster

The human visual system is remarkable in learning new visual concepts from just a few examples. This is precisely the goal behind few-shot class incremental learning (FSCIL), where the emphasis is additionally placed on ensuring the model does not suffer from "forgetting". In this paper, we push the…

Cited by 36PDFScholar
2022

FS-COCO: Towards Understanding of Freehand Sketches of Common Objects in Context

ECCV 2022poster

"We advance sketch research to scenes with the first dataset of freehand scene sketches, FSCOCO. With practical applications in mind, we collect sketches that convey well scene content but can be sketched within a few minutes by a person with any sketching skills. Our dataset comprises 10,000 freeha…

2022

Partially Does It: Towards Scene-Level FG-SBIR With Partial Input

CVPR 2022poster

We scrutinise an important observation plaguing scene-level sketch research -- that a significant portion of scene sketches are "partial". A quick pilot study reveals: (i) a scene sketch does not necessarily contain all objects in the corresponding photo, due to the subjective holistic interpretatio…

Cited by 29PDFScholar
2022

Sketch3T: Test-Time Training for Zero-Shot SBIR

CVPR 2022poster

Zero-shot sketch-based image retrieval typically asks for a trained model to be applied as is to unseen categories. In this paper, we question to argue that this setup by definition is not compatible with the inherent abstract and subjective nature of sketches -- the model might transfer well to new…

Cited by 60PDFScholar
2022

Sketching Without Worrying: Noise-Tolerant Sketch-Based Image Retrieval

CVPR 2022poster

Sketching enables many exciting applications, notably, image retrieval. The fear-to-sketch problem (i.e., "I can't sketch") has however proven to be fatal for its widespread adoption. This paper tackles this "fear" head on, and for the first time, proposes an auxiliary module for existing retrieval…

Cited by 68PDFcodeScholar
2021

Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition

ICCV 2021poster

Although text recognition has significantly evolved over the years, state-of the-art (SOTA) models still struggle in the wild scenarios due to complex backgrounds, varying fonts, uncontrolled illuminations, distortions and other artifacts. This is because such models solely depend on visual informat…

Cited by 80PDFScholar
2021

MetaHTR: Towards Writer-Adaptive Handwritten Text Recognition

CVPR 2021poster

Handwritten Text Recognition (HTR) remains a challenging problem to date, largely due to the varying writing styles that exist amongst us. Prior works however generally operate with the assumption that there is a limited number of styles, most of which have already been captured by existing datasets…

Cited by 43PDFScholar
2021

More Photos Are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image Retrieval

CVPR 2021poster

A fundamental challenge faced by existing Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) models is the data scarcity -- model performances are largely bottlenecked by the lack of sketch-photo pairs. Whilst the number of photos can be easily scaled, each corresponding sketch still needs to be in…

Cited by 85PDFScholar
2021

StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval

CVPR 2021poster

Sketch-based image retrieval (SBIR) is a cross-modal matching problem which is typically solved by learning a joint embedding space where the semantic content shared between photo and sketch modalities are preserved. However, a fundamental challenge in SBIR has been largely ignored so far, that is,…

Cited by 127PDFScholar
2021

Text Is Text, No Matter What: Unifying Text Recognition Using Knowledge Distillation

ICCV 2021poster

Text recognition remains a fundamental and extensively researched topic in computer vision, largely owing to its wide array of commercial applications. The challenging nature of the very problem however dictated a fragmentation of research efforts: Scene Text Recognition (STR) that deals with text i…

Cited by 35PDFScholar
2021

Towards the Unseen: Iterative Text Recognition by Distilling From Errors

ICCV 2021poster

Visual text recognition is undoubtedly one of the most extensively researched topics in computer vision. Great progress have been made to date, with the latest models starting to focus on the more practical "in-the-wild" setting. However, a salient problem still hinders practical deployment -- prior…

Cited by 23PDFScholar
2021

Vectorization and Rasterization: Self-Supervised Learning for Sketch and Handwriting

CVPR 2021poster

Self-supervised learning has gained prominence due to its efficacy at learning powerful representations from unlabelled data that achieve excellent performance on many challenging downstream tasks. However, supervision-free pre-text tasks are challenging to design and usually modality specific. Alth…

Cited by 69PDFScholar
2020

Fine-Grained Visual Classification via Progressive Multi-Granularity Training of Jigsaw Patches

ECCV 2020poster

Fine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works mainly tackle this problem by focusing on how to locate the most discriminative parts, more complementary parts, and parts o…

2020

Sketch Less for More: On-the-Fly Fine-Grained Sketch-Based Image Retrieval

CVPR 2020oral

Fine-grained sketch-based image retrieval (FG-SBIR) addresses the problem of retrieving a particular photo instance given a user's query sketch. Its widespread applicability is however hindered by the fact that drawing a sketch takes time, and most people struggle to draw a complete and faithful ske…

Cited by 135PDFScholar
2019

Facial Micro-expression Spotting and Recognition Using Time Contrasted Feature with Visual Memory

ICASSP 2019accepted

Facial micro-expressions are sudden involuntary minute muscle movements which reveal true emotions that people try to conceal. Spotting a micro-expression and recognizing it is a major challenge owing to its short duration and intensity. Many works pursued traditional and deep learning based approac…

Cited by 0SourceScholar
2019

Handwriting Recognition in Low-Resource Scripts Using Adversarial Learning

CVPR 2019poster

Handwritten Word Recognition and Spotting is a challenging field dealing with handwritten text possessing irregular and complex shapes. The design of deep neural network models makes it necessary to extend training datasets in order to introduce variations and increase the number of samples; word-re…

Cited by 84PDFScholar
2019

User Constrained Thumbnail Generation Using Adaptive Convolutions

ICASSP 2019accepted

Thumbnails are widely used all over the world as a preview for digital images. In this work we propose a deep neural framework to generate thumbnails of any size and aspect ratio, even for unseen values during training, with high accuracy and precision. We use Global Context Aggregation (GCA) and a…

Cited by 0SourceScholar