← Search

Hongyang Chao

17 accepted papers

2024

Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning

ECCV 2024poster

"Recently, zero-shot image captioning has gained increasing attention, where only text data is available for training. The remarkable progress in text-to-image diffusion model presents the potential to resolve this task by employing synthetic image-caption pairs generated by this pre-trained prior.…

Cited by 0SourcePDFScholar
2023

Semantic-Conditional Diffusion Networks for Image Captioning

CVPR 2023poster

Recent advances on text-to-image generation have witnessed the rise of diffusion models which act as powerful generative models. Nevertheless, it is not trivial to exploit such latent variable models to capture the dependency among discrete words and meanwhile pursue complex visual-language alignmen…

2023

TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance

ICCV 2023poster

In this paper, we propose a novel cross-modal distillation method, called TinyCLIP, for large-scale language-image pre-trained models. The method introduces two core techniques: affinity mimicking and weight inheritance. Affinity mimicking explores the interaction between modalities during distillat…

Cited by 65PDFcodeScholar
2021

Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers

NeurIPS 2021poster

We present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a full image from a single input (e.g., a latent code), the new formulation enables a flexible local manipulation for differ…

Cited by 33SourcePDFScholar
2021

Rethinking and Improving Relative Position Encoding for Vision Transformer

ICCV 2021poster

Relative position encoding (RPE) is important for transformer to capture sequence ordering of input tokens. General efficacy has been proven in natural language processing. However, in computer vision, its efficacy is not well studied and even remains controversial, e.g., whether relative position e…

Cited by 451PDFcodeScholar
2021

Searching the Search Space of Vision Transformer

NeurIPS 2021poster

Vision Transformer has shown great visual representation power in substantial vision tasks such as recognition and detection, and thus been attracting fast-growing efforts on manually designing more effective architectures. In this paper, we propose to use neural architecture search to automate this…

2020

Learning Joint Spatial-Temporal Transformations for Video Inpainting

ECCV 2020poster

High-quality video inpainting that completes missing regions in video frames is a promising yet challenging task. State-of-the-art approaches adopt attention models to complete a frame by searching missing contents from reference frames, and further complete whole videos frame by frame. However, the…

2019

Learning Pyramid-Context Encoder Network for High-Quality Image Inpainting

CVPR 2019poster

High-quality image inpainting requires filling missing regions in a damaged image with plausible content. Existing works either fill the regions by copying high-resolution patches or generating semantically-coherent patches from region context, while neglecting the fact that both visual and semantic…

Cited by 609PDFcodeScholar
2019

WSOD2: Learning Bottom-Up and Top-Down Objectness Distillation for Weakly-Supervised Object Detection

ICCV 2019poster

We study on weakly-supervised object detection (WSOD) which plays a vital role in relieving human involvement from object-level annotations. Predominant works integrate region proposal mechanisms with convolutional neural networks (CNN). Although CNN is proficient in extracting discriminative local…

Cited by 185PDFcodeScholar
2018

Image Blind Denoising With Generative Adversarial Network Based Noise Modeling

CVPR 2018poster

In this paper, we consider a typical image blind denoising problem, which is to remove unknown noise from noisy images. As we all know, discriminative learning based methods, such as DnCNN, can achieve state-of-the-art denoising results, but they are not applicable to this problem due to the lack of…

Cited by 728SourcePDFScholar
2018

Jointly Localizing and Describing Events for Dense Video Captioning

CVPR 2018poster

Automatically describing a video with natural language is regarded as a fundamental challenge in computer vision. The problem nevertheless is not trivial especially when a video contains multiple events to be worthy of mention, which often happens in real videos. A valid question is how to temporall…

2017

A fast intra-prediction decision algorithm in inter-frame based on a novel feature of HEVC

ICASSP 2017accepted

With the quad-tree based coding structure and more flexible intramodes, the coding efficiency provided by intra-technique in interframes of HEVC is much higher than the preceding standard H.264/AVC. However, the computing complexity is also significantly increased. Although only a few CUs are encode…

Cited by 0SourceScholar
2016

Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette Consistency

CVPR 2016poster

In this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouett…

Cited by 7PDFScholar
2015

MeshStereo: A Global Stereo Model With Mesh Alignment Regularization for View Interpolation

ICCV 2015oral

We present a novel global stereo model designed for view interpolation. Unlike existing stereo models which only output a disparity map, our model is able to output a 3D triangular mesh, which can be directly used for view interpolation. To this aim, we partition the input stereo images into 2D tria…

Cited by 206PDFScholar