← Search

Huaizhe Xu

4 accepted papers

2024

Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation

ECCV 2024poster

"Text-to-image generation has made significant advancements with the introduction of text-to-image diffusion models. These models typically consist of a language model that interprets user prompts and a vision model that generates corresponding images. As language and vision models continue to progr…

2023

MP-Former: Mask-Piloted Transformer for Image Segmentation

CVPR 2023poster

We present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and…

2023

Mask DINO: Towards a Unified Transformer-Based Framework for Object Detection and Segmentation

CVPR 2023poster

In this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query e…