← Search

Songcen Xu

37 accepted papers

2026

Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing

CVPR 2026

Always-on sensing is essential for next-generation edge/wearable AI systems, yet continuous high-fidelity RGB video capture remains prohibitively expensive for resource-constrained mobile and edge platforms. We present a new paradigm for efficient streaming video understanding: grayscale-always, col

Cited by 0SourcecodeScholar
2026

Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Features

CVPR 2026

Current diffusion-based makeup transfer methods commonly use the makeup information encoded by off-the-shelf foundation models (e.g., CLIP) as condition to preserve the makeup style of reference image in the generation. Although effective, these works mainly have two limitations: (1) foundation mode

Cited by 0SourcecodeScholar
2026

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

CVPR 2026

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are typically sparse and discrete. Conversely, synthetic data sc

Cited by 0SourcecodeScholar
2026

Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting

CVPR 2026

Feed-forward 3D Gaussian Splatting (3DGS) models enable real-time scene generation but are hindered by suboptimal pixel-aligned primitive placement, which relies on a dense, rigid grid that limits both quality and efficiency. We introduce a new feed-forward architecture that detects 3D Gaussian prim

Cited by 0SourceScholar
2024

Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images

AAAI 2024technical

Stable diffusion, a generative model used in text-to-image synthesis, frequently encounters resolution-induced composition problems when generating images of varying sizes. This issue primarily stems from the model being trained on pairs of single-scale images and their corresponding text descriptio…

2024

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

CVPR 2024poster

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks such as controllable image generation and image editing while downstream video synthesis tasks are less explore…

2024

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

CVPR 2024poster

Co-speech gestures if presented in the lively form of videos can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons resulting in the omission of appearance information we focus on the direct generation of audio-driven co-spee…

2024

DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior

CVPR 2024poster

3D generation has raised great attention in recent years. With the success of text-to-image diffusion models the 2D-lifting technique becomes a promising route to controllable 3D generation. However these methods tend to present inconsistent geometry which is also known as the Janus problem. We obse…

2024

EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head

ECCV 2024poster

"We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view consistency and a lack of emotional expressiveness. To address…

2024

GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction

ECCV 2024poster

"We present GSD, a diffusion model approach based on Gaussian Splatting (GS) representation for 3D object reconstruction from a single view. Prior works suffer from inconsistent 3D geometry or mediocre rendering quality due to improper representations. We take a step towards resolving these shortcom…

Cited by 7SourcePDFScholar
2024

JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation

ECCV 2024poster

"Score Distillation Sampling (SDS) by well-trained 2D diffusion models has shown great promise in text-to-3D generation. However, this paradigm distills view-agnostic 2D image distributions into the rendering distribution of 3D representation for each view independently, overlooking the coherence ac…

2024

LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model

ECCV 2024poster

"Despite the success of generating high-quality images given any text prompts by diffusion-based generative models, prior work directly generates the entire images, but cannot provide object-wise manipulation capability. To support wider real applications like professional graphic design and digital…

2024

MagicEraser: Erasing Any Objects via Semantics-Aware Control

ECCV 2024poster

"The traditional image inpainting task aims to restore corrupted regions by referencing surrounding background and foreground. However, the object erasure task, which is in increasing demand, aims to erase objects and generate harmonious background. Previous GAN-based inpainting methods struggle wit…

2024

MirrorGaussian: Reflecting 3D Gaussians for Reconstructing Mirror Reflections

ECCV 2024poster

"3D Gaussian Splatting showcases notable advancements in photo-realistic and real-time novel view synthesis. However, it faces challenges in modeling mirror reflections, which exhibit substantial appearance variations from different viewpoints. To tackle this problem, we present MirrorGaussian, the…

2024

PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion

ECCV 2024poster

"Current large-scale diffusion models represent a giant leap forward in conditional image synthesis, capable of interpreting diverse cues like text, human poses, and edges. However, their reliance on substantial computational resources and extensive data collection remains a bottleneck. On the other…

2024

Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution

CVPR 2024poster

Artifact-free super-resolution (SR) aims to translate low-resolution images into their high-resolution counterparts with a strict integrity of the original content eliminating any distortions or synthetic details. While traditional diffusion-based SR techniques have demonstrated remarkable abilities…

2024

Semantics-aware Motion Retargeting with Vision-Language Models

CVPR 2024poster

Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here we present a novel Semantics-aware Motion reTargeting (SMT) metho…

Cited by 5SourcePDFScholar
2024

TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling

ECCV 2024poster

"Given a 3D mesh, we aim to synthesize 3D textures that correspond to arbitrary textual descriptions. Current methods for generating and assembling textures from sampled views often result in prominent seams or excessive smoothing. To tackle these issues, we present TexGen, a novel multi-view sampli…

2024

TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields

ICLR 2024poster

Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tack…

Cited by 12SourcePDFScholar
2024

VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction

CVPR 2024poster

Existing NeRF-based methods for large scene reconstruction often have limitations in visual quality and rendering speed. While the recent 3D Gaussian Splatting works well on small-scale and object-centric scenes scaling it up to large scenes poses challenges due to limited video memory long optimiza…

Cited by 116SourcePDFScholar
2023

CLIPPING: Distilling CLIP-Based Models With a Student Base for Video-Language Retrieval

CVPR 2023poster

Pre-training a vison-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vison-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capabi…

Cited by 47SourcePDFScholar
2023

Co-Speech Gesture Synthesis by Reinforcement Learning With Contrastive Pre-Trained Rewards

CVPR 2023poster

There is a growing demand of automatically synthesizing co-speech gestures for virtual characters. However, it remains a challenge due to the complex relationship between input speeches and target gestures. Most existing works focus on predicting the next gesture that fits the data best, however, su…

2023

Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the Wild

NeurIPS 2023poster

This paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependen…

2023

Few-Shot Learning With Visual Distribution Calibration and Cross-Modal Distribution Alignment

CVPR 2023poster

Pre-trained vision-language models have inspired much research on few-shot learning. However, with only a few training images, there exist two crucial problems: (1) the visual feature distributions are easily distracted by class-irrelevant information in images, and (2) the alignment between the vis…

2023

HiVLP: Hierarchical Interactive Video-Language Pre-Training

ICCV 2023poster

Video-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and…

Cited by 6PDFScholar
2023

Low-Light Image Enhancement with Illumination-Aware Gamma Correction and Complete Image Modelling Network

ICCV 2023poster

This paper presents a novel network structure with illumination-aware gamma correction and complete image modelling to solve the low-light image enhancement problem. Low-light environments usually lead to less informative large-scale dark areas, directly learning deep representations from low-light…

Cited by 37PDFScholar
2023

PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval

ICCV 2023poster

Text-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video…

Cited by 19PDFScholar
2023

Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images

ICCV 2023poster

Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation. However, due to the limited text-3D face data pairs, text-driven 3D…

Cited by 18PDFcodeScholar
2022

CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation

ECCV 2022poster

"Top-down methods dominate the field of 3D human pose and shape estimation, because they are decoupled from human detection and allow researchers to focus on the core problem. However, cropping, their first step, discards the location information from the very beginning, which makes themselves unabl…

2021

Agreement-Discrepancy-Selection: Active Learning with Progressive Distribution Alignment

AAAI 2021technical

In active learning, the ignorance of aligning unlabeled samples' distribution with that of labeled samples hinders the model trained upon labeled samples from selecting informative unlabeled samples. In this paper, we propose an agreement-discrepancy-selection (ADS) approach, and target at unifying…

Cited by 12SourcePDFScholar
2021

DualPoseNet: Category-Level 6D Object Pose and Size Estimation Using Dual Pose Network With Refined Learning of Pose Consistency

ICCV 2021poster

Category-level 6D object pose and size estimation is to predict full pose configurations of rotation, translation, and size for object instances observed in single, arbitrary views of cluttered scenes. In this paper, we propose a new method of Dual Pose Network with refined learning of pose consiste…

Cited by 155PDFcodeScholar
2021

Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE

CVPR 2021poster

Given an incomplete image without additional constraint, image inpainting natively allows for multiple solutions as long as they appear plausible. Recently, multiple-solution inpainting methods have been proposed and shown the potential of generating diverse results. However, these methods have diff…

Cited by 274PDFcodeScholar
2021

Instance Segmentation in 3D Scenes Using Semantic Superpoint Tree Networks

ICCV 2021poster

Instance segmentation in 3D scenes is fundamental in many applications of scene understanding. It is yet challenging due to the compound factors of data irregularity and uncertainty in the numbers of instances. State-of-the-art methods largely rely on a general pipeline that first learns point-wise…

Cited by 143PDFcodeScholar
2021

Multiple Instance Active Learning for Object Detection

CVPR 2021poster

Despite the substantial progress of active learning for image recognition, there still lacks an instance-level active learning method specified for object detection. In this paper, we propose Multiple Instance Active Object Detection (MI-AOD), to select the most informative images for detector train…

Cited by 168PDFcodeScholar
2020

Renovating Parsing R-CNN for Accurate Multiple Human Parsing

ECCV 2020poster

Multiple human parsing aims to segment various human parts and associate each part with the corresponding instance simultaneously. This is a very challenging task due to the diverse human appearance, semantic ambiguity of different body parts and clothing, and complex background. Through analysis of…

2016

Adaptive distributed compressed estimation based on recursive least squares with sensing matrix design

ICASSP 2016accepted

In this paper, a distributed compressed estimation (DCE) scheme is presented based on a distributed recursive-least squares algorithm for sparse signals and systems along with a sensing matrix design procedure based on compressive sensing techniques. The D-CE scheme consists of compression and decom…

Cited by 0SourceScholar