← Search

Xiao Dong

7 accepted papers

2025

CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

ICLR 2025poster

Virtual try-on methods based on diffusion models achieve realistic effects but often require additional encoding modules, a large number of training parameters, and complex preprocessing, which increases the burden on training and inference. In this work, we re-evaluate the necessity of additional m…

2025

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

ICASSP 2025accepted

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods frequently integrate semantic information from images or simply concatenate images, which often leads to low fidelity and flickering in the generated videos. T…

Cited by 25SourceScholar
2025

EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation

ICASSP 2025accepted

The diffusion model is widely leveraged for either controllable video generation or video interpolation. As each field has its task-specific problems, it is difficult to merely develop a single model for completing both tasks simultaneously. Moreover, most existing works only support image condition…

Cited by 0SourceScholar
2025

PD-SDF: Dynamic Surface Reconstruction Based on Plane Decomposition for Single View RGB-D Videos

ICASSP 2025accepted

Surface reconstruction of dynamic scenes from single view videos is a challenging task due to the highly ill-posed and under-constrained nature. Existing single view reconstruction methods suffer from severe quality issues, such as surface distortion and mesh adherison. In this paper, we propose an…

Cited by 0SourceScholar
2024

DRSM: Efficient Neural 4D Decomposition for Dynamic Reconstruction in Stationary Monocular Cameras

ICASSP 2024accepted

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to scene reconstructions that exploit multi-view observations,…

Cited by 0SourceScholar
2022

M5Product: Self-Harmonized Contrastive Learning for E-Commercial Multi-Modal Pretraining

CVPR 2022poster

Despite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets. By leveraging the natural suitability of E-commerce, where different mod…

Cited by 44PDFcodeScholar
2021

Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal Pretraining

ICCV 2021poster

Nowadays, customer's demands for E-commerce are more diversified, which introduces more complications to the product retrieval industry. Previous methods are either subject to single-modal input or perform supervised image-level product retrieval, thus fail to accommodate real-life scenarios where e…

Cited by 75PDFcodeScholar