← Search

Zhaowen Wang

37 accepted papers

2026

SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons

CVPR 2026

Graphic icons are a cornerstone of modern design workflows, yet they are often distributed as flattened single-path or compound-path graphics, where the original semantic layering is lost. This absence of semantic decomposition hinders downstream tasks such as editing, restyling, and animation. We f

Cited by 0SourcecodeScholar
2026

ULW-SLEEPNET: AN ULTRA-LIGHTWEIGHT NETWORK FOR MULTIMODAL SLEEP STAGE SCORING

ICASSP 2026poster

Automatic sleep stage scoring is crucial for the diagnosis and treatment of sleep disorders. Although deep learning models have advanced the field, many existing models are computationally demanding and designed for single-channel electroencephalography (EEG), limiting their practicality for multimo…

Cited by 0SourcePDFScholar
2025

CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long Contexts

ACL 2025long

Large Language Models (LLMs) have achieved remarkable success in natural language processing (NLP), particularly in single-turn question answering (QA) on short-text. However, their performance significantly declines when applied to multi-turn QA over extra-long context (ELC), as they struggle to ca…

Cited by 0SourcePDFScholar
2025

Rethinking Layered Graphic Design Generation with a Top-Down Approach

ICCV 2025poster

Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphi…

Cited by 0SourcePDFScholar
2024

Scaling Up Video Summarization Pretraining with Large Language Models

CVPR 2024poster

Long-form video content constitutes a significant portion of internet traffic making automated video summarization an essential research problem. However existing video summarization datasets are notably limited in their size constraining the effectiveness of state-of-the-art methods for generalizat…

Cited by 13SourcePDFScholar
2024

Visual Layout Composer: Image-Vector Dual Diffusion Model for Design Layout Generation

CVPR 2024poster

This paper proposes an image-vector dual diffusion model for generative layout design. Distinct from prior efforts that mostly ignore element-level visual information our approach integrates the power of a pre-trained large image diffusion model to guide layout composition in a vector diffusion mode…

Cited by 5SourcePDFScholar
2024

WAS: Dataset and Methods for Artistic Text Segmentation

ECCV 2024poster

"Accurate text segmentation results are crucial for text-related generative tasks, such as text image generation, text editing, text removal, and text style transfer. Recently, some scene text segmentation methods have made significant progress in segmenting regular text. However, these methods perf…

2023

Align and Attend: Multimodal Summarization With Dual Contrastive Losses

CVPR 2023poster

The goal of multimodal summarization is to extract the most important information from different modalities to form summaries. Unlike unimodal summarization, the multimodal summarization task explicitly leverages cross-modal information to help generate more reliable and high-quality summaries. Howe…

2023

DualVector: Unsupervised Vector Font Synthesis With Dual-Part Representation

CVPR 2023poster

Automatic generation of fonts can be an important aid to typeface design. Many current approaches regard glyphs as pixelated images, which present artifacts when scaling and inevitable quality losses after vectorization. On the other hand, existing vector font synthesis methods either fail to repres…

2023

Layout Representation Learning with Spatial and Structural Hierarchies

AAAI 2023technical

We present a novel hierarchical modeling method for layout representation learning, the core of design documents (e.g., user interface, poster, template). Existing works on layout representation often ignore element hierarchies, which is an important facet of layouts, and mainly rely on the spatial…

2023

Moment Detection in Long Tutorial Videos

ICCV 2023poster

Tutorial videos play an increasingly important role in professional development and self-directed education. For users to realise the full benefits of this medium, tutorial videos must be efficiently searchable. In this work, we focus on the task of moment detection, in which the goal is to localise…

Cited by 4PDFcodeScholar
2023

SCCS: Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment

ACL 2023findings

Multimedia summarization with multimodal output (MSMO) is a recently explored application in language grounding. It plays an essential role in real-world applications, i.e., automatically generating cover images and titles for news articles or providing introductions to online videos. However, exist…

Cited by 9SourcePDFScholar
2023

SVGformer: Representation Learning for Continuous Vector Graphics Using Transformers

CVPR 2023poster

Advances in representation learning have led to great success in understanding and generating data in various domains. However, in modeling vector graphics data, the pure data-driven approach often yields unsatisfactory results in downstream tasks as existing deep learning methods often require the…

Cited by 9SourcePDFScholar
2022

Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition

ECCV 2022poster

"Artistic text recognition is an extremely challenging task with a wide range of applications. However, current scene text recognition methods mainly focus on irregular text while have not explored artistic text specifically. The challenges of artistic text recognition include the various appearance…

2021

A Multi-Implicit Neural Representation for Fonts

NeurIPS 2021poster

Fonts are ubiquitous across documents and come in a variety of styles. They are either represented in a native vector format or rasterized to produce fixed resolution images. In the first case, the non-standard representation prevents benefiting from latest network architectures for neural represen…

Cited by 29SourcePDFScholar
2021

Rethinking Text Segmentation: A Novel Dataset and a Text-Specific Refinement Approach

CVPR 2021poster

Text segmentation is a prerequisite in many real-world text-related tasks, e.g., text style transfer, and scene text removal. However, facing the lack of high-quality datasets and dedicated investigations, this critical prerequisite has been left as an assumption in many works, and has been largely…

Cited by 85PDFcodeScholar
2020

Texture Hallucination for Large-Factor Painting Super-Resolution

ECCV 2020poster

We aim to super-resolve digital paintings, synthesizing realistic details from high-resolution reference painting materials for very large scaling factors (g 8$ imes$, 16$ imes$). However, previous single image super-resolution (SISR) methods would either lose textural details or introduce unpleasin…

Cited by 29SourcePDFScholar
2019

An Internal Learning Approach to Video Inpainting

ICCV 2019poster

We propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP…

Cited by 101PDFcodeScholar
2019

Controllable Artistic Text Style Transfer via Shape-Matching GAN

ICCV 2019oral

Artistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an…

Cited by 129PDFcodeScholar
2019

Large-Scale Tag-Based Font Retrieval With Generative Feature Learning

ICCV 2019poster

Font selection is one of the most important steps in a design workflow. Traditional methods rely on ordered lists which require significant domain knowledge and are often difficult to use even for trained professionals. In this paper, we address the problem of large-scale tag-based font retrieval wh…

Cited by 36PDFScholar
2018

Flow-Grounded Spatial-Temporal Video Prediction from Still Images

ECCV 2018poster

Existing video prediction methods mainly rely on observing multiple historical frames or focus on predicting the next one-frame. In this work, we study the problem of generating consecutive multiple future frames by observing one single still image only. We formulate the multi-frame prediction task…

Cited by 159SourcePDFScholar
2018

Multi-Content GAN for Few-Shot Font Style Transfer

CVPR 2018poster

In this work, we focus on the challenge of taking partial observations of highly-stylized text and generalizing the observations to generate unobserved glyphs in the ornamented typeface. To generate a set of multi-content images following a consistent style from very few examples, we propose an end-…

2018

Multi-Task Adversarial Network for Disentangled Feature Learning

CVPR 2018poster

We address the problem of image feature learning for the applications where multiple factors exist in the image generation process and only some factors are of our interest. We present a novel multi-task adversarial network based on an encoder-discriminator-generator architecture. The encoder extrac…

Cited by 77SourcePDFScholar
2018

Re-Weighted Adversarial Adaptation Network for Unsupervised Domain Adaptation

CVPR 2018poster

Unsupervised Domain Adaptation (UDA) aims to transfer domain knowledge from existing well-defined tasks to new ones where labels are unavailable. In the real-world applications, as the domain (task) discrepancies are usually uncontrollable, it is significantly motivated to match the feature distribu…

Cited by 174SourcePDFScholar
2018

Synthetically Supervised Feature Learning for Scene Text Recognition

ECCV 2018poster

We address the problem of image feature learning for scene text recognition. The image features in the state-of-the-art methods are learned from large-scale synthetic image datasets. However, most methods only rely on outputs of the synthetic data generation process, namely realistically looking ima…

Cited by 109SourcePDFScholar
2018

Towards Privacy-Preserving Visual Recognition via Adversarial Training: A Pilot Study

ECCV 2018poster

This paper aims to improve privacy-preserving visual recognition, an increasingly demanded feature in smart camera applications, by formulating a unique adversarial training framework. The proposed framework explicitly learns a degradation transform for the original video inputs, in order to optimiz…

2018

Visual to Sound: Generating Natural Sound for Videos in the Wild

CVPR 2018poster

As two of the five traditional human senses (sight, hearing, taste, smell, and touch), vision and sound are basic sources through which humans understand the world. Often correlated during natural events, these two modalities combine to jointly affect human perception. In this paper, we pose the tas…

Cited by 261SourcePDFScholar
2018

``Factual'' or ``Emotional'': Stylized Image Captioning with Adaptive Learning and Attention

ECCV 2018poster

Generating stylized captions for an image is an emerging topic in image captioning. Given an image as input, it requires the system to generate a caption that has a specific style (e.g., humorous, romantic, positive, and negative) while describing the image content semantically accurately. In this p…

Cited by 95SourcePDFScholar
2017

AMC: Attention guided Multi-modal Correlation Learning for Image Search

CVPR 2017poster

Given a user's query, traditional image search systems rank images according to its relevance to a single modality (e.g., image content or surrounding text). Nowadays, an increasing number of images on the Internet are available with associated meta data in rich modalities (e.g., titles, keywords, t…

Cited by 49PDFcodeScholar
2017

Diversified Texture Synthesis With Feed-Forward Networks

CVPR 2017spotlight

Recent progresses on deep discriminative and generative modeling have shown promising results on texture synthesis. However, existing feed-forward based methods trade off generality for efficiency, which suffer from many issues, such as shortage of generality (i.e., build one network per texture), l…

Cited by 341PDFScholar
2017

Robust Video Super-Resolution With Learned Temporal Dynamics

ICCV 2017poster

Video super-resolution (SR) aims to generate a high-resolution (HR) frame from multiple low-resolution (LR) frames. The inter-frame temporal relation is as crucial as the intra-frame spatial relation for tackling this problem. However, how to utilize temporal information efficiently and effectively…

Cited by 289PDFScholar
2017

Universal Style Transfer via Feature Transforms

NeurIPS 2017poster

Universal style transfer aims to transfer arbitrary visual styles to content images. Existing feed-forward based methods, while enjoying the inference efficiency, are mainly limited by inability of generalizing to unseen styles or compromised visual quality. In this paper, we present a simple yet ef…