← Search

Guo-jun Qi

41 accepted papers

2026

C-Evolve: Consensus-based Evolution for Prompt Groups

ICLR 2026poster

Prompt evolution algorithms offer a powerful paradigm for enhancing AI systems based on closed-source models, while few work explores whether aggregating results from multiple prompts to reach a consensus can further advance the system capability boundary. In this paper, we introduce Consensus-Evol…

Cited by 0SourceScholar
2026

Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models

ICLR 2026poster

Mask-based Diffusion Language Models (DLMs) struggle to revise incorrect tokens: once a token is generated, it typically remains fixed. The key challenge is to identify potential errors in the inputs. In this paper, we propose Remasking-enabled Diffusion Language Model (RemeDi), a mask-based DLM tha…

Cited by 0SourceScholar
2025

DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making

ICML 2025poster

Integrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extra…

Cited by 0SourcePDFScholar
2025

Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models

NeurIPS 2025poster

We introduce the Diffusion Chain of Lateral Thought (DCoLT), a reasoning framework for diffusion language models. DCoLT treats each intermediate step in the reverse diffusion process as a latent "thinking" action and optimizes the entire reasoning trajectory to maximize the reward on the correctness…

Cited by 0SourceScholar
2025

Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation

CVPR 2025poster

Diffusion and flow matching models have achieved remarkable success in text-to-image generation. However, these models typically rely on the predetermined denoising schedules for all prompts. The multi-step reverse diffusion process can be regarded as a kind of chain-of-thought for generating high-q…

2024

BARET: Balanced Attention Based Real Image Editing Driven by Target-Text Inversion

AAAI 2024technical

Image editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple conditions (e.g., mask, sketch, caption), and time consuming fine-t…

Cited by 5SourcePDFScholar
2024

OmniMotionGPT: Animal Motion Generation with Limited Data

CVPR 2024poster

Our paper aims to generate diverse and realistic animal motion sequences from textual descriptions without a large-scale animal text-motion dataset. While the task of text-driven human motion synthesis is already extensively studied and benchmarked it remains challenging to transfer this success to…

Cited by 7SourcePDFScholar
2024

One-Step Diffusion Distillation through Score Implicit Matching

NeurIPS 2024poster

Despite their strong performances on many generative tasks, diffusion models require a large number of sampling steps in order to generate realistic samples. This has motivated the community to develop effective methods to distill pre-trained diffusion models into more efficient models, but these m…

2024

Towards Open Domain Text-Driven Synthesis of Multi-Person Motions

ECCV 2024poster

"This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the…

Cited by 10SourcePDFScholar
2024

Zero-shot High-fidelity and Pose-controllable Character Animation

IJCAI 2024poster

Image-to-video (I2V) generation aims to create a video sequence from a single image, which requires high temporal coherence and visual fidelity. However, existing approaches suffer from inconsistency of character appearances and poor preservation of fine details. Moreover, they require a large amoun…

Cited by 4SourcePDFScholar
2023

High-Fidelity Clothed Avatar Reconstruction From a Single Image

CVPR 2023poster

This paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR)…

2023

Monocular 3D Object Detection with Bounding Box Denoising in 3D by Perceiver

ICCV 2023poster

The main challenge of monocular 3D object detection is the accurate localization of 3D center. Motivated by a new and strong observation that this challenge can be remedied by a 3D-space local-grid search scheme in an ideal case, we propose a stage-wise approach, which combines the information flow…

Cited by 14PDFScholar
2023

OTAvatar: One-Shot Talking Face Avatar With Controllable Tri-Plane Rendering

CVPR 2023poster

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the represe…

2023

Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person Search

ICCV 2023poster

Weakly supervised person search aims to jointly detect and match persons with only bounding box annotations. Existing approaches typically focus on improving the features by exploring the relations of persons. However, scale variation problem is a more severe obstacle and under-studied that a person…

Cited by 14PDFcodeScholar
2022

Exploring Resolution and Degradation Clues As Self-Supervised Signal for Low Quality Object Detection

ECCV 2022poster

"Image restoration algorithms such as super resolution (SR) are indispensable pre-processing modules for object detection in low qual-ity images. Most of these algorithms assume the degradation is fixed andknown a priori. However, in pratical, either the real degrdation or optimalup-sampling ratio r…

2021

A Unified Multi-Scenario Attacking Network for Visual Object Tracking

AAAI 2021technical

Existing methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking…

Cited by 19SourcePDFScholar
2021

AdCo: Adversarial Contrast for Efficient Learning of Unsupervised Representations From Self-Trained Negative Adversaries

CVPR 2021poster

Contrastive learning relies on constructing a collection of negative examples that are sufficiently hard to discriminate against positive queries when their representations are self-trained. Existing contrastive learning methods either maintain a queue of negative samples over mini-batches while onl…

Cited by 178PDFcodeScholar
2021

Auto-Encoding Transformations in Reparameterized Lie Groups for Unsupervised Learning

AAAI 2021technical

Unsupervised training of deep representations has demonstrated remarkable potentials in mitigating the prohibitive expenses on annotating labeled data recently. Among them is predicting transformations as a pretext task to self-train representations, which has shown great potentials for unsupervised…

Cited by 5SourcePDFScholar
2021

Multitask AET With Orthogonal Tangent Regularity for Dark Object Detection

ICCV 2021poster

Dark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noises. Most of the existing studies tackle this by either targeting human vision for better visual perception or improving the machine vision for specific high-level tasks. In addition,…

Cited by 154PDFcodeScholar
2020

GraphTER: Unsupervised Learning of Graph Transformation Equivariant Representations via Auto-Encoding Node-Wise Transformations

CVPR 2020poster

Recent advances in Graph Convolutional Neural Networks (GCNNs) have shown their efficiency for nonEuclidean data on graphs, which often require a large amount of labeled data with high cost. It it thus critical to learn graph feature representations in an unsupervised manner in practice. To this end…

Cited by 56PDFcodeScholar
2020

PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search

ICLR 2020spotlight

Differentiable architecture search (DARTS) provided a fast solution in finding effective network architectures, but suffered from large memory and computing overheads in jointly training a super-net and searching for an optimal architecture. In this paper, we present a novel approach, namely Partia…

Cited by 920SourcecodeScholar
2020

Rotation Equivariant Graph Convolutional Network for Spherical Image Classification

CVPR 2020poster

Convolutional neural networks (CNNs) designed for low-dimensional regular grids will unfortunately lead to non-optimal solutions for analyzing spherical images, due to their different geometrical properties from planar images. In this paper, we generalize the grid-based CNNs to a non-Euclidean space…

Cited by 46PDFcodeScholar
2020

Transformation GAN for Unsupervised Image Synthesis and Representation Learning

CVPR 2020poster

Generative Adversarial Networks (GAN) have shown promising performance in image synthesis and unsupervised learning (USL). In most cases, however, the representations extracted from unsupervised GAN are usually unsatisfactory in other computer vision tasks. By using conditional GAN (CGAN), this prob…

Cited by 31PDFScholar
2019

AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations Rather Than Data

CVPR 2019oral

The success of deep neural networks often relies on a large amount of labeled examples, which can be difficult to obtain in many real scenarios. To address this challenge, unsupervised methods are strongly preferred for training neural networks without using any labeled data. In this paper, we prese…

Cited by 265PDFcodeScholar
2019

AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations

ICCV 2019poster

The learning of Transformation-Equivariant Representations (TERs), which is introduced by Hinton et al. [??], has been considered as a principle to reveal visual structures under various transformations. It contains the celebrated Convolutional Neural Networks (CNNs) as a special case that only equi…

Cited by 49PDFScholar
2018

An Adversarial Approach to Hard Triplet Generation

ECCV 2018poster

While deep neural networks have demonstrated competitive results for many visual recognition and image retrieval tasks, the major challenge lies in distinguishing similar images from different categories (i.e., hard negative examples) while clustering images with large variations from the same categ…

Cited by 121SourcePDFScholar
2018

CapProNet: Deep Feature Learning via Orthogonal Projections onto Capsule Subspaces

NeurIPS 2018poster

In this paper, we formalize the idea behind capsule nets of using a capsule vector rather than a neuron activation to predict the label of samples. To this end, we propose to learn a group of capsule subspaces onto which an input feature vector is projected. Then the lengths of resultant capsules ar…

Cited by 89SourcePDFScholar
2018

Global Versus Localized Generative Adversarial Nets

CVPR 2018poster

In this paper, we present a novel localized Generative Adversarial Net (GAN) to learn on the manifold of real data. Compared with the classic GAN that {em globally} parameterizes a manifold, the Localized GAN (LGAN) uses local coordinate charts to parameterize distinct local geometry of how data poi…

Cited by 97SourcePDFScholar
2018

Interleaved Structured Sparse Convolutional Neural Networks

CVPR 2018poster

In this paper, we study the problem of designing efficient convolutional neural network architectures with the interest in eliminating the redundancy in convolution kernels. In addition to structured sparse kernels, low-rank kernels and the product of low-rank kernels,the product of structured spars…

Cited by 160SourcePDFScholar