← Search

Zhiyuan Fang

11 accepted papers

2026

HIVE-3D: Hierarchical Voxel Enhancement for High-Quality 3D Scene Generation

ICML 2026poster

Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation. In this work, we introduce \name, a novel method for high-quality 3D scene generation based on hierarchical …

Cited by 0SourceScholar
2025

A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian Splatting

ICCV 2025poster

Recently, the field of 3D scene stylization has attracted considerable attention, particularly for applications in the metaverse. A key challenge is rapidly transferring the style of an arbitrary reference image to a 3D scene while faithfully preserving its content structure and spatial layout. Work…

Cited by 0SourcePDFScholar
2024

Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation

ECCV 2024poster

"The literature on text-to-image generation is plagued by issues of faithfully composing entities with relations. But there lacks a formal understanding of how entity-relation compositions can be effectively learned. Moreover, the underlying phenomenon space that meaningfully reflects the problem st…

2023

End-to-end Knowledge Retrieval with Multi-modal Queries

ACL 2023long

We investigate knowledge retrieval with multi-modal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called ReMuQ for benchmarking progress on this task. ReMuQ require…

2022

Injecting Semantic Concepts Into End-to-End Image Captioning

CVPR 2022poster

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more fl…

Cited by 137PDFcodeScholar
2022

Mining Unseen Classes via Regional Objectness: A Simple Baseline for Incremental Segmentation

NeurIPS 2022accept

Incremental or continual learning has been extensively studied for image classification tasks to alleviate catastrophic forgetting, a phenomenon in which earlier learned knowledge is forgotten when learning new concepts. For class incremental semantic segmentation, such a phenomenon often becomes mu…

2021

Compressing Visual-Linguistic Model via Knowledge Distillation

ICCV 2021poster

Despite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation(KD) to effectively compress a transformer-based large VL model into a small VL model. The major challenge arises from the inconsis…

Cited by 100PDFcodeScholar
2021

SEED: Self-supervised Distillation For Visual Representation

ICLR 2021poster

This paper is concerned with self-supervised learning for small models. The problem is motivated by our empirical studies that while the widely used contrastive self-supervised learning method has shown great progress on large model training, it does not work well for small models. To address this p…

2020

ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language

ECCV 2020poster

Person search by natural language aims at retrieving a specific person in a large-scale image pool that matches given textual descriptions. While most of the current methods treat the task as a holistic visual and textual feature matching one, we approach it from an attribute-aligning perspective th…

2017

Range Loss for Deep Face Recognition With Long-Tailed Training Data

ICCV 2017poster

Deep convolutional neural networks have achieved significant improvements on face recognition task due to their ability to learn highly discriminative features from tremendous amounts of face images. Many large scale face datasets exhibit long-tail distribution where a small number of entities (pers…

Cited by 512PDFScholar