← Search

Yi-Zhe Song

107 accepted papers

2026

PerfGuard: A Performance-Aware Agent for Visual Content Generation

ICLR 2026poster

The advancement of Large Language Model (LLM)-powered agents has enabled automated task processing through reasoning and tool invocation capabilities. However, existing frameworks often operate under the idealized assumption that tool executions are invariably successful, relying solely on textual d…

Cited by 0SourcecodeScholar
2026

SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance

CVPR 2026

Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending this success to video reconstruction remains a significant c

Cited by 0SourcecodeScholar
2026

SketchEvo: Leveraging Drawing Dynamics for Enhanced Image Synthesis

ICLR 2026poster

Sketching represents humanity's most intuitive form of visual expression -- a universal language that transcends barriers. Although recent diffusion models integrate sketches with text, they often regard the complete sketch merely as a static visual constraint, neglecting the human preference inform…

Cited by 0SourceScholar
2025

CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth

CVPR 2025poster

We present CaricatureBooth, a system that transforms caricature creation into a simple interactive experience -- as easy as using a photo booth! A key challenge in caricature generation is two-fold: the scarcity of high-quality caricature data and the difficulty in enabling precise creative control…

2025

Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection

ICCV 2025poster

Keypoint detection, integral to modern machine perception, faces challenges in few-shot learning, particularly when source data from the same distribution as the query is unavailable. This gap is addressed by leveraging sketches, a popular form of human expression, providing a source-free alternativ…

Cited by 0SourcePDFScholar
2025

Fuse2Match: Training-Free Fusion of Flow, Diffusion, and Contrastive Models for Zero-Shot Semantic Matching

NeurIPS 2025poster

Recent work shows that features from Stable Diffusion (SD) and contrastively pretrained models like DINO can be directly used for zero-shot semantic correspondence via naive feature concatenation. In this paper, we explore the stronger potential of Stable Diffusion 3 (SD3), a rectified flow-based mo…

Cited by 0SourcecodeScholar
2025

NitroFusion: High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training

CVPR 2025poster

We introduce NitroFusion, a fundamentally different approach to single-step diffusion that achieves high-quality generation through a dynamic adversarial framework. While one-step methods offer dramatic speed advantages, they typically suffer from quality degradation compared to their multi-step cou…

2025

Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models

NeurIPS 2025poster

Negative guidance -- explicitly suppressing unwanted attributes -- remains a fundamental challenge in diffusion models, particularly in few-step sampling regimes. While Classifier-Free Guidance (CFG) works well in standard settings, it fails under aggressive sampling step compression due to divergen…

Cited by 0SourceScholar
2025

Sketch Down the FLOPs: Towards Efficient Networks for Human Sketch

CVPR 2025poster

As sketch research has collectively matured over time, its adaptation for at-mass commercialisation emerges on the immediate horizon. Despite an already mature research endeavour for photos, there is no research on the efficient inference specifically designed for sketch data. In this paper, we firs…

Cited by 0SourcePDFScholar
2025

SketchFusion: Learning Universal Sketch Features through Fusing Foundation Models

CVPR 2025poster

While foundation models have revolutionised computer vision, their effectiveness for sketch understanding remains limited by the unique challenges of abstract, sparse visual inputs. Through systematic analysis, we uncover two fundamental limitations: Stable Diffusion (SD) struggles to extract meanin…

Cited by 0SourcePDFScholar
2025

VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis

AAAI 2025technical

Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study rev…

2024

DemoCaricature: Democratising Caricature Generation with a Rough Sketch

CVPR 2024poster

In this paper we democratise caricature generation empowering individuals to effortlessly craft personalised caricatures with just a photo and a conceptual sketch. Our objective is to strike a delicate balance between abstraction and identity while preserving the creativity and subjectivity inherent…

Cited by 9SourcePDFScholar
2024

DemoFusion: Democratising High-Resolution Image Generation With No $$$

CVPR 2024poster

High-resolution image generation with Generative Artificial Intelligence (GenAI) has immense potential but due to the enormous capital investment required for training it is increasingly centralised to a few large corporations and hidden behind paywalls. This paper aims to democratise high-resolutio…

2024

Do Generalised Classifiers really work on Human Drawn Sketches?

ECCV 2024poster

"This paper, for the first time, marries large foundation models with human sketch understanding. We demonstrate what this brings – a paradigm shift in terms of generalised sketch representation learning (e.g., classification). This generalisation happens on two fronts: (i) generalisation across unk…

2024

Doodle Your 3D: From Abstract Freehand Sketches to Precise 3D Shapes

CVPR 2024poster

In this paper we democratise 3D content creation enabling precise generation of 3D shapes from abstract sketches while overcoming limitations tied to drawing skills. We introduce a novel part-level modelling and alignment framework that facilitates abstraction modelling and cross-modal correspondenc…

2024

Freeview Sketching: View-Aware Fine-Grained Sketch-Based Image Retrieval

ECCV 2024poster

"In this paper, we delve into the intricate dynamics of Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) by addressing a critical yet overlooked aspect – the choice of viewpoint during sketch creation. Unlike photo systems that seamlessly handle diverse views through extensive datasets, sketch sy…

Cited by 0SourcePDFScholar
2024

How to Handle Sketch-Abstraction in Sketch-Based Image Retrieval?

CVPR 2024poster

In this paper we propose a novel abstraction-aware sketch-based image retrieval framework capable of handling sketch abstraction at varied levels. Prior works had mainly focused on tackling sub-factors such as drawing style and order we instead attempt to model abstraction as a whole and propose fea…

Cited by 16SourcePDFScholar
2024

Improved Image Captioning Via Knowledge Graph-Augmented Models

ICASSP 2024accepted

Multimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address…

Cited by 0SourceScholar
2024

It's All About Your Sketch: Democratising Sketch Control in Diffusion Models

CVPR 2024poster

This paper unravels the potential of sketches for diffusion models addressing the deceptive promise of direct sketch control in generative AI. We importantly democratise the process enabling amateur sketches to generate precise images living up to the commitment of "what you sketch is what you get".…

2024

Modelling complex vector drawings with stroke-clouds

ICLR 2024poster

Vector drawings are innately interactive as they preserve creational cues. Despite this desirable property they remain relatively under explored due to the difficulties in modeling complex vector drawings. This is in part due to the primarily _sequential and auto-regressive nature_ of existing appro…

Cited by 3SourcePDFScholar
2024

SketchINR: A First Look into Sketches as Implicit Neural Representations

CVPR 2024poster

We propose SketchINR to advance the representation of vector sketches with implicit neural models. A variable length vector sketch is compressed into a latent space of fixed dimension that implicitly encodes the underlying shape as a function of time and strokes. The learned function predicts the xy…

2024

Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers

CVPR 2024poster

This paper for the first time explores text-to-image diffusion models for Zero-Shot Sketch-based Image Retrieval (ZS-SBIR). We highlight a pivotal discovery: the capacity of text-to-image diffusion models to seamlessly bridge the gap between sketches and photos. This proficiency is underpinned by th…

Cited by 11SourcePDFScholar
2024

What Sketch Explainability Really Means for Downstream Tasks?

CVPR 2024poster

In this paper we explore the unique modality of sketch for explainability emphasising the profound impact of human strokes compared to conventional pixel-oriented studies. Beyond explanations of network behavior we discern the genuine implications of explainability across diverse downstream sketch-r…

Cited by 4SourcePDFScholar
2024

Wired Perspectives: Multi-View Wire Art Embraces Generative AI

CVPR 2024poster

Creating multi-view wire art (MVWA) a static 3D sculpture with diverse interpretations from different viewpoints is a complex task even for skilled artists. In response we present DreamWire an AI system enabling everyone to craft MVWA easily. Users express their vision through text prompts or scribb…

Cited by 6SourcePDFScholar
2024

You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image Retrieval

CVPR 2024poster

Two primary input modalities prevail in image retrieval: sketch and text. While text is widely used for inter-category retrieval tasks sketches have been established as the sole preferred modality for fine-grained image retrieval due to their ability to capture intricate visual details. In this pape…

Cited by 16SourcePDFScholar
2023

3D VR Sketch Guided 3D Shape Prototyping and Exploration

ICCV 2023poster

3D shape modeling is labor-intensive, time-consuming, and requires years of expertise. To facilitate 3D shape modeling, we propose a 3D shape generation network that takes a 3D VR sketch as a condition. We assume that sketches are created by novices without art training and aim to reconstruct geomet…

Cited by 13PDFcodeScholar
2023

An Erudite Fine-Grained Visual Classification Model

CVPR 2023poster

Current fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of the FGVC algorithm in real-life scenarios. In this paper, we…

2023

Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image Classification

AAAI 2023technical

The main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting --…

2023

CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or Not

CVPR 2023poster

In this paper, we leverage CLIP for zero-shot sketch based image retrieval (ZS-SBIR). We are largely inspired by recent advances on foundation models and the unparalleled generalisation ability they seem to offer, but for the first time tailor it to benefit the sketch community. We put forward novel…

2023

ChiroDiff: Modelling chirographic data with Diffusion Models

ICLR 2023poster

Generative modelling over continuous-time geometric constructs, a.k.a $chirographic\ data$ such as handwriting, sketches, drawings etc., have been accomplished through autoregressive distributions. Such strictly-ordered discrete factorization however falls short of capturing key properties of chirog…

Cited by 16SourcePDFScholar
2023

Controllable Person Image Synthesis with Pose-Constrained Latent Diffusion

ICCV 2023poster

Controllable person image synthesis aims at rendering a source image based on user-specified changes in body pose or appearance. Prior art approaches leverage pixel-level denoising diffusion models conditioned on the coarse skeleton via cross-attention. This leads to two limitations: low efficiency…

Cited by 25PDFcodeScholar
2023

Democratising 2D Sketch to 3D Shape Retrieval Through Pivoting

ICCV 2023poster

This paper studies the problem of 2D sketch to 3D shape retrieval, but with a focus on democratising the process. We would like this democratisation to happen on two fronts: (i) to remove the need for large-scale specifically sourced 2D sketch and 3D shape datasets, and (ii) to remove restrictions o…

Cited by 6PDFScholar
2023

DiffTAD: Temporal Action Detection with Proposal Denoising Diffusion

ICCV 2023poster

We propose a new formulation of temporal action detection (TAD) with denoising diffusion, DiffTAD in short. Taking as input random temporal proposals, it can yield action proposals accurately given an untrimmed long video. This presents a generative modeling perspective, against previous discriminat…

Cited by 43PDFcodeScholar
2023

Exploiting Unlabelled Photos for Stronger Fine-Grained SBIR

CVPR 2023poster

This paper advances the fine-grained sketch-based image retrieval (FG-SBIR) literature by putting forward a strong baseline that overshoots prior state-of-the art by 11%. This is not via complicated design though, but by addressing two critical issues facing the community (i) the gold standard trip…

2023

FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks

CVPR 2023highlight

In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in each individual input/output format and dataset size. It has been common to des…

2023

HeadSculpt: Crafting 3D Head Avatars with Text

NeurIPS 2023poster

Recently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models. However, existing methods still struggle to create high-fidelity 3D head avatars in t…

Cited by 51SourcePDFScholar
2023

On-the-Fly Category Discovery

CVPR 2023poster

Although machines have surpassed humans on visual recognition problems, they are still limited to providing closed-set answers. Unlike machines, humans can cognize novel categories at the first observation. Novel category discovery (NCD) techniques, transferring knowledge from seen categories to dis…

2023

Picture That Sketch: Photorealistic Image Generation From Abstract Sketches

CVPR 2023poster

Given an abstract, deformed, ordinary sketch from untrained amateurs like you and me, this paper turns it into a photorealistic image - just like those shown in Fig. 1(a), all non-cherry-picked. We differ significantly from prior art in that we do not dictate an edgemap-like sketch to start with, bu…

2023

SceneTrilogy: On Human Scene-Sketch and Its Complementarity With Photo and Text

CVPR 2023poster

In this paper, we extend scene understanding to include that of human sketch. The result is a complete trilogy of scene representation from three diverse and complementary modalities -- sketch, photo, and text. Instead of learning a rigid three-way embedding and be done with it, we focus on learning…

Cited by 32SourcePDFScholar
2023

Sketch2Saliency: Learning To Detect Salient Objects From Human Drawings

CVPR 2023poster

Human sketch has already proved its worth in various visual understanding tasks (e.g., retrieval, segmentation, image-captioning, etc). In this paper, we reveal a new trait of sketches -- that they are also salient. This is intuitive as sketching is a natural attentive process at its core. More spec…

Cited by 25SourcePDFScholar
2023

SketchKnitter: Vectorized Sketch Generation with Diffusion Models

ICLR 2023top-25%

We show vectorized sketch generation can be identified as a reversal of the stroke deformation process. This relationship was established by means of a diffusion model that learns data distributions over the stroke-point locations and pen states of real human sketches. Given randomly scattered strok…

Cited by 31SourcePDFScholar
2023

SketchXAI: A First Look at Explainability for Human Sketches

CVPR 2023poster

This paper, for the very first time, introduces human sketches to the landscape of XAI (Explainable Artificial Intelligence). We argue that sketch as a "human-centred" data form, represents a natural interface to study explainability. We focus on cultivating sketch-specific explainability designs. T…

2023

Task-aware Adaptive Learning for Cross-domain Few-shot Learning

ICCV 2023poster

Although existing few-shot learning works yield promising results for in-domain queries, they still suffer from weak cross-domain generalization. Limited support data requires effective knowledge transfer, but domain-shift makes this harder. Towards this emerging challenge, researchers improved adap…

Cited by 15PDFcodeScholar
2023

What Can Human Sketches Do for Object Detection?

CVPR 2023poster

Sketches are highly expressive, inherently capturing subjective and fine-grained visual cues. The exploration of such innate properties of human sketches has, however, been limited to that of image retrieval. In this paper, for the first time, we cultivate the expressiveness of sketches but for the…

2023

Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable Style

CVPR 2023highlight

This paper studies the problem of zero-short sketch-based image retrieval (ZS-SBIR), however with two significant differentiators to prior art (i) we tackle all variants (inter-category, intra-category, and cross datasets) of ZS-SBIR with just one network ("everything"), and (ii) we would really lik…

2022

Adaptive Fine-Grained Sketch-Based Image Retrieval

ECCV 2022poster

"The recent focus on Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) has shifted towards generalising a model to new categories without any training data from them. In real-world applications, however, a trained FG-SBIR model is often applied to both new categories and different human sketchers,…

2022

Doodle It Yourself: Class Incremental Learning by Drawing a Few Sketches

CVPR 2022poster

The human visual system is remarkable in learning new visual concepts from just a few examples. This is precisely the goal behind few-shot class incremental learning (FSCIL), where the emphasis is additionally placed on ensuring the model does not suffer from "forgetting". In this paper, we push the…

Cited by 36PDFScholar
2022

FS-COCO: Towards Understanding of Freehand Sketches of Common Objects in Context

ECCV 2022poster

"We advance sketch research to scenes with the first dataset of freehand scene sketches, FSCOCO. With practical applications in mind, we collect sketches that convey well scene content but can be sketched within a few minutes by a person with any sketching skills. Our dataset comprises 10,000 freeha…

2022

FashionViL: Fashion-Focused Vision-and-Language Representation Learning

ECCV 2022poster

"Large-scale Vision-and-Language (V+L) pre-training for representation learning has proven to be effective in boosting various downstream V+L tasks. However, when it comes to the fashion domain, existing V+L methods are inadequate as they overlook the unique characteristics of both fashion V+L data…

2022

Partially Does It: Towards Scene-Level FG-SBIR With Partial Input

CVPR 2022poster

We scrutinise an important observation plaguing scene-level sketch research -- that a significant portion of scene sketches are "partial". A quick pilot study reveals: (i) a scene sketch does not necessarily contain all objects in the corresponding photo, due to the subjective holistic interpretatio…

Cited by 29PDFScholar
2022

Proposal-Free Temporal Action Detection via Global Segmentation Mask Learning

ECCV 2022poster

"Existing temporal action detection (TAD) methods rely on generating an overwhelmingly large number of proposals per video. This leads to complex model designs due to proposal generation and/or per-proposal action instance evaluation and the resultant high computational cost. In this work, for the f…

2022

Semi-Supervised Temporal Action Detection with Proposal-Free Masking

ECCV 2022poster

"Existing temporal action detection (TAD) methods rely on a large number of training data with segment-level annotations. Collecting and annotating such a training set is thus highly expensive and unscalable. Semi-supervised TAD (SS-TAD) alleviates this problem by leveraging unlabeled videos freely…

2022

Sketch3T: Test-Time Training for Zero-Shot SBIR

CVPR 2022poster

Zero-shot sketch-based image retrieval typically asks for a trained model to be applied as is to unseen categories. In this paper, we question to argue that this setup by definition is not compatible with the inherent abstract and subjective nature of sketches -- the model might transfer well to new…

Cited by 60PDFScholar
2022

SketchODE: Learning neural sketch representation in continuous time

ICLR 2022poster

Learning meaningful representations for chirographic drawing data such as sketches, handwriting, and flowcharts is a gateway for understanding and emulating human creative expression. Despite being inherently continuous-time data, existing works have treated these as discrete-time sequences, disrega…

Cited by 18SourcePDFScholar
2022

SketchSampler: Sketch-Based 3D Reconstruction via View-Dependent Depth Sampling

ECCV 2022poster

"Reconstructing a 3D shape based on a single sketch image is challenging due to the large domain gap between a sparse, irregular sketch and a regular, dense 3D shape. Existing works try to employ the global feature extracted from sketch to directly predict the 3D coordinates, but they usually suffer…

2022

Sketching Without Worrying: Noise-Tolerant Sketch-Based Image Retrieval

CVPR 2022poster

Sketching enables many exciting applications, notably, image retrieval. The fear-to-sketch problem (i.e., "I can't sketch") has however proven to be fatal for its widespread adoption. This paper tackles this "fear" head on, and for the first time, proposes an auxiliary module for existing retrieval…

Cited by 68PDFcodeScholar
2022

Zero-Shot Temporal Action Detection via Vision-Language Prompting

ECCV 2022poster

"Existing temporal action detection (TAD) methods rely on large training data including segment-level annotations, limited to recognizing previously seen classes alone during inference. Collecting and annotating a large training set for each class of interest is costly and hence unscalable. Zero-sho…

2021

Cloud2Curve: Generation and Vectorization of Parametric Sketches

CVPR 2021poster

Analysis of human sketches in deep learning has advanced immensely through the use of waypoint-sequences rather than raster-graphic representations. We further aim to model sketches as a sequence of low-dimensional parametric curves. To this end, we propose an inverse graphics framework capable of a…

Cited by 42PDFScholar
2021

Context-Aware Layout to Image Generation With Enhanced Object Appearance

CVPR 2021poster

A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout. Built upon the recent advances in generative adversarial networks (GANs), recent L2I models have made great progress.…

Cited by 65PDFcodeScholar
2021

Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition

ICCV 2021poster

Although text recognition has significantly evolved over the years, state-of the-art (SOTA) models still struggle in the wild scenarios due to complex backgrounds, varying fonts, uncontrolled illuminations, distortions and other artifacts. This is because such models solely depend on visual informat…

Cited by 80PDFScholar
2021

MetaHTR: Towards Writer-Adaptive Handwritten Text Recognition

CVPR 2021poster

Handwritten Text Recognition (HTR) remains a challenging problem to date, largely due to the varying writing styles that exist amongst us. Prior works however generally operate with the assumption that there is a limited number of styles, most of which have already been captured by existing datasets…

Cited by 43PDFScholar
2021

More Photos Are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image Retrieval

CVPR 2021poster

A fundamental challenge faced by existing Fine-Grained Sketch-Based Image Retrieval (FG-SBIR) models is the data scarcity -- model performances are largely bottlenecked by the lack of sketch-photo pairs. Whilst the number of photos can be easily scaled, each corresponding sketch still needs to be in…

Cited by 85PDFScholar
2021

One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning Objective

NeurIPS 2021poster

A deep hashing model typically has two main learning objectives: to make the learned binary hash codes discriminative and to minimize a quantization error. With further constraints such as bit balance and code orthogonality, it is not uncommon for existing models to employ a large number (>4) of los…

2021

Simpler Is Better: Few-Shot Semantic Segmentation With Classifier Weight Transformer

ICCV 2021poster

A few-shot semantic segmentation model is typically composed of a CNN encoder, a CNN decoder and a simple classifier (separating foreground and background pixels). Most existing methods meta-learn all three model components for fast adaptation to a new class. However, given that as few as a single s…

Cited by 229PDFcodeScholar
2021

SketchLattice: Latticed Representation for Sketch Manipulation

ICCV 2021poster

The key challenge in designing a sketch representation lies with handling the abstract and iconic nature of sketches. Existing work predominantly utilizes either, (i) a pixelative format that treats sketches as natural images employing off-the-shelf CNN-based networks, or (ii) an elaborately designe…

Cited by 24PDFScholar
2021

StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval

CVPR 2021poster

Sketch-based image retrieval (SBIR) is a cross-modal matching problem which is typically solved by learning a joint embedding space where the semantic content shared between photo and sketch modalities are preserved. However, a fundamental challenge in SBIR has been largely ignored so far, that is,…

Cited by 127PDFScholar
2021

Text Is Text, No Matter What: Unifying Text Recognition Using Knowledge Distillation

ICCV 2021poster

Text recognition remains a fundamental and extensively researched topic in computer vision, largely owing to its wide array of commercial applications. The challenging nature of the very problem however dictated a fragmentation of research efforts: Scene Text Recognition (STR) that deals with text i…

Cited by 35PDFScholar
2021

Towards the Unseen: Iterative Text Recognition by Distilling From Errors

ICCV 2021poster

Visual text recognition is undoubtedly one of the most extensively researched topics in computer vision. Great progress have been made to date, with the latest models starting to focus on the more practical "in-the-wild" setting. However, a salient problem still hinders practical deployment -- prior…

Cited by 23PDFScholar
2021

Vectorization and Rasterization: Self-Supervised Learning for Sketch and Handwriting

CVPR 2021poster

Self-supervised learning has gained prominence due to its efficacy at learning powerful representations from unlabelled data that achieve excellent performance on many challenging downstream tasks. However, supervision-free pre-text tasks are challenging to design and usually modality specific. Alth…

Cited by 69PDFScholar
2021

Your "Flamingo" is My "Bird": Fine-Grained, or Not

CVPR 2021poster

Whether what you see in Figure 1 is a "flamingo" or a "bird", is the question we ask in this paper. While fine-grained visual classification (FGVC) strives to arrive at the former, for the majority of us non-experts just "bird" would probably suffice. The real question is therefore -- how can we tai…

Cited by 145PDFcodeScholar
2020

A Tree-Structured Decoder for Image-to-Markup Generation

ICML 2020poster

Recent encoder-decoder approaches typically employ string decoders to convert images into serialized strings for image-to-markup. However, for tree-structured representational markup, string representations can hardly cope with the structural complexity. In this work, we first show via a set of toy…

Cited by 89SourcePDFScholar
2020

BézierSketch: A generative model for scalable vector sketches

ECCV 2020poster

The study of neural generative models of human sketches is a fascinating contemporary modeling problem due to the links between sketch image generation and the human drawing process. The landmark SketchRNN provided breakthrough by sequentially generating sketches as a sequence of waypoints. However…

2020

Fine-Grained Visual Classification via Progressive Multi-Granularity Training of Jigsaw Patches

ECCV 2020poster

Fine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works mainly tackle this problem by focusing on how to locate the most discriminative parts, more complementary parts, and parts o…

2020

Sketch Less for More: On-the-Fly Fine-Grained Sketch-Based Image Retrieval

CVPR 2020oral

Fine-grained sketch-based image retrieval (FG-SBIR) addresses the problem of retrieving a particular photo instance given a user's query sketch. Its widespread applicability is however hindered by the fact that drawing a sketch takes time, and most people struggle to draw a complete and faithful ske…

Cited by 135PDFScholar
2020

Solving Mixed-Modal Jigsaw Puzzle for Fine-Grained Sketch-Based Image Retrieval

CVPR 2020poster

ImageNet pre-training has long been considered crucial by the fine-grained sketch-based image retrieval (FG-SBIR) community due to the lack of large sketch-photo paired datasets for FG-SBIR training. In this paper, we propose a self-supervised alternative for representation pre-training. Specificall…

Cited by 107PDFScholar
2020

Stochastic Classifiers for Unsupervised Domain Adaptation

CVPR 2020poster

A common strategy adopted by existing state-of-the-art unsupervised domain adaptation (UDA) methods is to employ two classifiers to identify the misaligned local regions between source and target domain. Following the 'wisdom of the crowd' principle, one has to ask: why stop at two? Indeed, we find…

Cited by 205PDFcodeScholar
2019

Doodle to Search: Practical Zero-Shot Sketch-Based Image Retrieval

CVPR 2019oral

In this paper, we investigate the problem of zero-shot sketch-based image retrieval (ZS-SBIR), where human sketches are used as queries to conduct retrieval of photos from unseen categories. We importantly advance prior arts by proposing a novel ZS-SBIR scenario that represents a firm step forward i…

Cited by 224PDFScholar
2019

Episodic Training for Domain Generalization

ICCV 2019oral

Domain generalization (DG) is the challenging and topical problem of learning models that generalize to novel testing domains with different statistics than a set of known training domains. The simple approach of aggregating data from all source domains and training a single deep neural network end-…

Cited by 565PDFcodeScholar
2019

Generalising Fine-Grained Sketch-Based Image Retrieval

CVPR 2019poster

Fine-grained sketch-based image retrieval (FG-SBIR) addresses matching specific photo instance using free-hand sketch as a query modality. Existing models aim to learn an embedding space in which sketch and photo can be directly compared. While successful, they require instance-level pairing within…

Cited by 118PDFScholar
2019

Generalizable Person Re-Identification by Domain-Invariant Mapping Network

CVPR 2019poster

We aim to learn a domain generalizable person re-identification (ReID) model. When such a model is trained on a set of source domains (ReID datasets collected from different camera networks), it can be directly applied to any new unseen dataset for effective ReID without any model updating. Despite…

Cited by 301PDFScholar
2018

Learning Deep Sketch Abstraction

CVPR 2018poster

Human free-hand sketches have been studied in various contexts including sketch recognition, synthesis and fine-grained sketch-based image retrieval (FG-SBIR). A fundamental challenge for sketch analysis is to deal with drastically different human drawing styles, particularly in terms of abstraction…

Cited by 117SourcePDFScholar
2018

Learning to Sketch With Shortcut Cycle Consistency

CVPR 2018poster

To see is to sketch -- free-hand sketching naturally builds ties between human and machine vision. In this paper, we present a novel approach for translating an object photo to a sketch, mimicking the human sketching process. This is an extremely challenging task because the photo and sketch domains…

Cited by 139SourcePDFScholar
2018

Sketch-a-Classifier: Sketch-Based Photo Classifier Generation

CVPR 2018poster

Contemporary deep learning techniques have made image recognition a reasonably reliable technology. However training effective photo classifiers typically takes numerous examples which limits image recognition's scalability and applicability to scenarios where images may not be available. This has m…

Cited by 27SourcePDFScholar
2018

SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval

CVPR 2018poster

We propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset.Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventiona…

Cited by 149SourcePDFScholar
2018

SketchyScene: Richly-Annotated Scene Sketches

ECCV 2018poster

We contribute the rst large-scale dataset of scene sketches, SketchyScene, with the goal of advancing research on sketch understanding at both the object and scene level. The dataset is created through a novel and carefully designed crowdsourcing pipeline, enabling users to eciently generate large q…

2018

Universal Sketch Perceptual Grouping

ECCV 2018poster

In this work we aim to develop a universal sketch grouper. That is, a grouper that can be applied to sketches of any category in any domain to group constituent strokes/segments into semantically meaningful object parts. The first obstacle to this goal is the lack of large-scale datasets with groupi…

Cited by 58SourcePDFScholar
2017

Deep Spatial-Semantic Attention for Fine-Grained Sketch-Based Image Retrieval

ICCV 2017poster

Human sketches are unique in being able to capture both the spatial topology of a visual object, as well as its subtle appearance details. Fine-grained sketch-based image retrieval (FG-SBIR) importantly leverages on such fine-grained characteristics of sketches to conduct instance-level retrieval of…

Cited by 318PDFScholar
2015

Making Better Use of Edges via Perceptual Grouping

CVPR 2015poster

We propose a perceptual grouping framework that organizes image edges into meaningful structures and demonstrate its usefulness on various computer vision tasks. Our grouper formulates edge grouping as a graph partition problem, where a learning to rank method is developed to encode probabilities of…

Cited by 105SourcePDFScholar