← Search

Guanglu Song

37 accepted papers

2026

High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning

CVPR 2026

Face swapping aims to seamlessly transfer a source facial identity onto a target while preserving target attributes such as pose and expression. Diffusion models, known for their superior generative capabilities, have recently shown promise in advancing face-swapping quality. This paper addresses tw

Cited by 0SourceScholar
2026

Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models

CVPR 2026

Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its deterministic sampling paradigm. Current methods address this issue by converting Ordina

Cited by 0SourceScholar
2025

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

ICML 2025poster

Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods mostly encode multiple reference images by averaging or concatenating their image embeddings as the injection condition, but such an image-independent operation cannot perform intera…

Cited by 6SourcePDFScholar
2025

MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines

ICLR 2025poster

The advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleav…

Cited by 0SourcePDFScholar
2025

Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning

ICRA 2025

Multimodal task specification is essential for enhanced robotic performance, where Cross-modality Alignment enables the robot to holistically understand complex task instructions. Directly annotating multimodal instructions for model training proves impractical, due to the sparsity of paired multimo

Cited by 5SourceScholar
2025

See Further When Clear: Curriculum Consistency Model

CVPR 2025poster

Significant advances have been made in the sampling efficiency of diffusion and flow matching models, driven by Consistency Distillation (CD), which trains a student model to mimic the output of a teacher model at a later timestep. However, we found that the knowledge discrepancy between student and…

Cited by 0SourcePDFScholar
2025

VividFace: A Robost and High-Fidelity Video Face Swapping Framework

NeurIPS 2025poster

Video face swapping has seen increasing adoption in diverse applications, yet existing methods primarily trained on static images struggle to address temporal consistency and complex real-world scenarios. To overcome these limitations, we propose the first video face swapping framework, VividFace,…

Cited by 0SourceScholar
2024

Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation

ECCV 2024poster

"Video outpainting is a challenging task, aiming at generating video content outside the viewport of the input video while maintaining inter-frame and intra-frame consistency. Existing methods fall short in either generation quality or flexibility. We introduce (Mastering Video Outpainting Through I…

2024

CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

NeurIPS 2024poster

Diffusion models have demonstrated great success in the field of text-to-image generation. However, alleviating the misalignment between the text prompts and images is still challenging. We break down the problem into two causes: concept ignorance and concept mismapping. To tackle the two challenges…

2024

Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models

ECCV 2024poster

"Optimizing a text-to-image diffusion model with a given reward function is an important but underexplored research area. In this study, we propose Deep Reward Tuning (DRTune), an algorithm that directly supervises the final output image of a text-to-image diffusion model and back-propagates through…

Cited by 14SourcePDFScholar
2024

Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

NeurIPS 2024poster

Large language models based on decoder-only transformers have demonstrated superior text understanding capabilities compared to CLIP and T5-series models. However, the paradigm for utilizing current advanced LLMs in text-to-image diffusion models remains to be explored. We observed an unusual phenom…

Cited by 17SourcePDFScholar
2024

FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis

ECCV 2024poster

"In this study, we delve into the generation of high-resolution images from pre-trained diffusion models, addressing persistent challenges, such as repetitive patterns and structural distortions, that emerge when models are applied beyond their trained resolutions. To address this issue, we introduc…

2024

LMDrive: Closed-Loop End-to-End Driving with Large Language Models

CVPR 2024poster

Despite significant recent progress in the field of autonomous driving modern methods still struggle and can incur serious accidents when encountering long-tail unforeseen events and challenging urban scenarios. On the one hand large language models (LLM) have shown impressive reasoning capabilities…

2024

MoVA: Adapting Mixture of Vision Experts to Multimodal Context

NeurIPS 2024poster

As the key component in multimodal large language models (MLLMs), the ability of the visual encoder greatly affects MLLM's understanding on diverse image content. Although some large-scale pretrained vision encoders such as vision encoders in CLIP and DINOv2 have brought promising performance, we fo…

2024

Phased Consistency Models

NeurIPS 2024poster

Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of…

2024

Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance

CVPR 2024poster

Classifier-Free Guidance (CFG) has been widely used in text-to-image diffusion models where the CFG scale is introduced to control the strength of text guidance on the whole image space. However we argue that a global CFG scale results in spatial inconsistency on varying semantic strengths and subop…

2024

Three Things We Need to Know About Transferring Stable Diffusion to Visual Dense Prediciton Tasks

ECCV 2024poster

"In this paper, we investigate how to conduct transfer learning to adapt Stable Diffusion to downstream visual dense prediction tasks such as semantic segmentation and depth estimation. We focus on fine-tuning the Stable Diffusion model, which has demonstrated impressive abilities in modeling image…

Cited by 3SourcePDFScholar
2024

Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

NeurIPS 2024spotlight

Multi-Modal Large Language Models (MLLMs) have demonstrated impressive performance in various VQA tasks. However, they often lack interpretability and struggle with complex visual inputs, especially when the resolution of the input image is high or when the interested region that could provide key i…

2024

ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model

ECCV 2024oral

"Although video generation has made great progress in capacity and controllability and is gaining increasing attention, currently available video generation models still make minimal progress in the video length they can generate. Due to the lack of well-annotated long video data, high training/infe…

2023

Decoupled DETR: Spatially Disentangling Localization and Classification for Improved End-to-End Object Detection

ICCV 2023poster

The introduction of DETR represents a new paradigm for object detection. However, its decoder conducts classification and box localization using shared queries and cross-attention layers, leading to suboptimal results. We observe that different regions of interest in the visual feature map are sui…

Cited by 23PDFScholar
2023

RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths

NeurIPS 2023poster

Text-to-image generation has recently witnessed remarkable achievements. We introduce a text-conditional image diffusion model, termed RAPHAEL, to generate highly artistic images, which accurately portray the text prompts, encompassing multiple nouns, adjectives, and verbs. This is achieved by stac…

2023

Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction

ICCV 2023poster

In this paper, we propose a new paradigm, named Historical Object Prediction (HoP) for multi-view 3D detection to leverage temporal information more effectively. The HoP approach is straightforward: given the current timestamp t, we generate a pseudo Bird's-Eye View (BEV) feature of timestamp t-k fr…

Cited by 37PDFcodeScholar
2023

UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object Detectors

ICCV 2023poster

Knowledge distillation (KD) has become a standard method to boost the performance of lightweight object detectors. Most previous works are feature-based, where students mimic the features of homogeneous teacher detectors. However, distilling the knowledge from the heterogeneous teacher fails in this…

Cited by 9PDFScholar
2022

"UniNet: Unified Architecture Search with Convolution, Transformer, and MLP"

ECCV 2022poster

"Recently, transformer and multi-layer perceptron (MLP) architectures have achieved impressive results on various vision tasks. However, how to effectively combine those operators to form high-performance hybrid visual architectures still remains a challenge. In this work, we study the learnable com…

2022

Large-batch Optimization for Dense Visual Predictions: Training Faster R-CNN in 4.2 Minutes

NeurIPS 2022accept

Training a large-scale deep neural network in a large-scale dataset is challenging and time-consuming. The recent breakthrough of large-batch optimization is a promising way to tackle this challenge. However, although the current advanced algorithms such as LARS and LAMB succeed in classification mo…

2022

Rethinking Robust Representation Learning under Fine-Grained Noisy Faces

ECCV 2022poster

"Learning robust feature representation from large-scale noisy faces stands out as one of the key challenges in high-performance face recognition. Recent attempts have been made to cope with this challenge by alleviating the intra-class conflict and inter-class conflict. However, the unconstrained n…

Cited by 1SourcePDFScholar
2022

Self-Slimmed Vision Transformer

ECCV 2022poster

"Vision transformers (ViTs) have become the popular structures and outperformed convolutional neural networks (CNNs) on various vision tasks. However, such powerful transformers bring a huge computation burden, because of the exhausting token-to-token comparison. The previous works focus on dropping…

2022

UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation Learning

ICLR 2022poster

It is a challenging task to learn rich and multi-scale spatiotemporal semantics from high-dimensional videos, due to large local redundancy and complex global dependency between video frames. The recent advances in this research have been mainly driven by 3D convolutional neural networks and vision…

2022

Unifying Visual Perception by Dispersible Points Learning

ECCV 2022poster

"We present a conceptually simple, flexible, and universal visual perception head for variant visual tasks, e.g., classification, object detection, instance segmentation and pose estimation, and different frameworks, such as one-stage or two-stage pipelines. Our approach effectively identifies an ob…

2021

Switchable K-Class Hyperplanes for Noise-Robust Representation Learning

ICCV 2021poster

Optimizing the K-class hyperplanes in the latent space has become the standard paradigm for efficient representation learning. However, it's almost impossible to find an optimal K-class hyperplane to accurately describe the latent space of massive noisy data. For this potential problem, we construct…

Cited by 7PDFcodeScholar
2018

Beyond Trade-Off: Accelerate FCN-Based Face Detector With Higher Accuracy

CVPR 2018poster

Fully convolutional neural network (FCN) has been dominating the game of face detection task for a few years with its congenital capability of sliding-window-searching with shared kernels, which boiled down all the redundant calculation, and most recent state-of-the-art methods such as Faster-RCNN,…

Cited by 40SourcePDFScholar
2018

Transductive Centroid Projection for Semi-supervised Large-scale Recognition

ECCV 2018poster

Conventional deep semi-supervised learning methods, such as recursive clustering and training process, suffer from cumulative error and high computational complexity when collaborating with Convolutional Neural Networks. To this end, we design a simple but effective learning mechanism that merely su…

Cited by 40SourcePDFScholar