← Search

Zhongang Qi

28 accepted papers

2026

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

ICLR 2026poster

Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexplored. In this paper, we introduce UniUGG, the first unified understanding and generation framework for 3D modalities. Ou…

Cited by 0SourcecodeScholar
2025

CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities

AAAI 2025technical

Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of video diffusion models (VDMs) to combine concepts and generate…

2025

DOGR: Towards Versatile Visual Document Grounding and Referring

ICCV 2025poster

With recent advances in Multimodal Large Language Models (MLLMs), grounding and referring capabilities have gained increasing attention for achieving detailed understanding and flexible user interaction. However, these capabilities still remain underdeveloped in visual document understanding due to…

2025

Less is More: Empowering GUI Agent with Context-Aware Simplification

ICCV 2025poster

The research focus of GUI agents is shifting from text-dependent to pure-vision-based approaches, which, though promising, prioritize comprehensive pre-training data collection while neglecting contextual modeling challenges. We probe the characteristics of element and history contextual modeling in…

2025

Mamba-3VL: Taming State Space Model for 3D Vision Language Learning

ICCV 2025poster

3D vision-language (3D-VL) reasoning, connecting natural language with 3D physical world, represents a milestone in advancing spatial intelligence. While transformer-based methods dominate 3D-VL research, their quadratic complexity and simplistic positional embedding mechanisms severely limits effec…

2025

Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion

CVPR 2025poster

With the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models (DMs) into this task. However, due to the scarcity of large-scale training data and comprehensive benchmarks, the optima…

Cited by 0SourcePDFScholar
2025

Taming Rectified Flow for Inversion and Editing

ICML 2025poster

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often struggle with inversion inaccuracies, which could further limit their effectivenes…

2025

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

NeurIPS 2025poster

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention has been given to scaling fine-grained pixel-level understa…

Cited by 0SourceScholar
2025

VisionMath: Vision-Form Mathematical Problem-Solving

ICCV 2025poster

Mathematical problems in real-world scenarios are often presented in a purely vision-form, where textual problem statement and accompanying math figures, e.g., geometry figures and functional graphs, are integrated into a single image. This vision-form problem-solving task requires precise comprehen…

2024

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

NeurIPS 2024poster

Recent advances in Video Large Language Models (Video-LLMs) have demonstrated their great potential in general-purpose video understanding. To verify the significance of these models, a number of benchmarks have been proposed to diagnose their capabilities in different scenarios. However, existing b…

2024

EA-VTR: Event-Aware Video-Text Retrieval

ECCV 2024poster

"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to…

Cited by 3SourcePDFScholar
2024

How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?

CVPR 2024poster

Dominant dual-encoder models enable efficient image-text retrieval but suffer from limited accuracy while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harnes…

Cited by 2SourcePDFScholar
2024

PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

CVPR 2024poster

Recent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However existing personalized generation methods cannot simultaneously satisfy the requirements of high efficiency promising identity (ID) fidelity and…

2024

SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model

AAAI 2024technical

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduc…

Cited by 8SourcePDFScholar
2024

T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models

AAAI 2024technical

The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and a…

2023

ERBNet: An Effective Representation Based Network for Unbiased Scene Graph Generation

ICASSP 2023accepted

The scene graph generation (SGG) task has attracted increasing attention in recent years. The goal of SGG is to predict relations between pairs of objects within an image. Due to the long-tailed distribution of the dataset annotations, the performance of SGG is still far from satisfactory. To addres…

Cited by 0SourceScholar
2023

Exploiting Contextual Objects and Relations for 3D Visual Grounding

NeurIPS 2023poster

3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information…

2023

LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image Generation

CVPR 2023poster

Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging…

2023

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

ICCV 2023poster

Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with…

Cited by 455PDFcodeScholar
2023

Order-Prompted Tag Sequence Generation for Video Tagging

ICCV 2023poster

Video Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multi-label class…

Cited by 4PDFScholar
2023

SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation

IJCAI 2023poster

As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D p…

2023

Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text Retrieval

AAAI 2023technical

Vision-language alignment learning for video-text retrieval arouses a lot of attention in recent years. Most of the existing methods either transfer the knowledge of image-text pretraining model to video-text retrieval task without fully exploring the multi-modal information of videos, or simply fus…

Cited by 25SourcePDFScholar
2023

ViLEM: Visual-Language Error Modeling for Image-Text Retrieval

CVPR 2023poster

Dominant pre-training works for image-text retrieval adopt "dual-encoder" architecture to enable high efficiency, where two encoders are used to extract image and text representations and contrastive learning is employed for global alignment. However, coarse-grained global alignment ignores detailed…

Cited by 14SourcePDFScholar
2022

BTS: A Bi-Lingual Benchmark for Text Segmentation in the Wild

CVPR 2022oral

As a prerequisite of many text-related tasks such as text erasing and text style transfer, text segmentation arouses more and more attention recently. Current researches mainly focus on only English characters and digits, while few work studies Chinese characters due to the lack of public large-scal…

Cited by 20PDFScholar
2021

Finding Discriminative Filters for Specific Degradations in Blind Super-Resolution

NeurIPS 2021spotlight

Recent blind super-resolution (SR) methods typically consist of two branches, one for degradation prediction and the other for conditional restoration. However, our experiments show that a one-branch network can achieve comparable performance to the two-branch scheme. Then we wonder: how can one-bra…

Cited by 43SourcePDFScholar
2021

Open-Book Video Captioning With Retrieve-Copy-Generate Network

CVPR 2021poster

In this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a…

Cited by 125PDFScholar