← Search

Weiwei Xu

26 accepted papers

2026

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

ICML 2026poster

Large Language Models(LLMs) have revolutionized text generation and multimodal perception, but their capabilities in 3D content generation remain underexplored. Existing methods compromise by producing either low-resolution meshes or coarse structural proxies, failing to capture fine-grained geometr…

Cited by 0SourceScholar
2026

FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization

AAAI 2026technical

Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment

Cited by 0SourcePDFScholar
2026

MatMart: Material Reconstruction of 3D Objects via Diffusion

CVPR 2026

Applying diffusion models to physically-based material estimation and generation has recently gained prominence. In this paper, we propose MatMart, a novel material reconstruction framework for 3D objects, offering the following advantages. First, MatMart adopts a two-stage reconstruction, starting

Cited by 0SourcecodeScholar
2026

Mimic-X: A Large-Scale Motion Dataset via Fast Physics-Based Controller Adaptation

AAAI 2026technical

Large and high-quality motion datasets are essential for advancing human motion modeling. However, limitations of existing motion datasets, such as insufficient scale or inadequate quality, significantly hinder the progress of this field. To address these limitations, we introduce Mimic-X, a large-s

Cited by 0SourcePDFScholar
2026

TurboGS: Accelerating 3D Gaussian Splatting via Error-Guided Sparse Pixel Sampling and Optimization

ICML 2026poster

Consumer-level applications require fast optimization of 3D Gaussian Splatting (3DGS) with high-fidelity novel view rendering. However, existing 3DGS acceleration approaches still incur substantial computation on redundant pixels while sacrificing fine details. In this paper, we present TurboGS, an …

Cited by 0SourceScholar
2025

AnimateAnything: Consistent and Controllable Animation for Video Generation

CVPR 2025poster

We propose a unified approach for video-controlled generation, enabling text-based guidance and manual annotations to control the generation of videos, similar to camera direction guidance. Specifically, we designed a two-stage algorithm. In the first stage, we convert all control information into f…

Cited by 10SourcePDFScholar
2025

DecoupledGaussian: Object-Scene Decoupling for Physics-Based Interaction

CVPR 2025poster

We present DecoupledGaussian, a novel system that decouples static objects from their contacted surfaces captured in-the-wild videos, a key prerequisite for realistic Newtonian-based physical simulations. Unlike prior methods focused on synthetic data or elastic jittering along the contact surface,…

2025

Detail-Preserving Latent Diffusion for Stable Shadow Removal

CVPR 2025poster

Achieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To addr…

Cited by 0SourcePDFScholar
2025

Intensity-Augmented LiDAR-Visual-Inertial Odometry and Meshing

IROS 2025

This paper presents a tightly-coupled LiDAR-Visual-Inertial Odometry (LIVO) system that integrates both LIO and VIO subsystems. The system jointly estimates the state by fusing LiDAR or visual data with Inertial Measurement Units (IMUs). It employs point-to-mesh tracking to optimize LiDAR poses and

Cited by 0SourceScholar
2025

Neural Shell Texture Splatting: More Details and Fewer Primitives

ICCV 2025poster

Gaussian splatting techniques have shown promising results in novel view synthesis, achieving high fidelity and efficiency. However, their high reconstruction quality comes at the cost of requiring a large number of primitives. We identify this issue as stemming from the entanglement of geometry and…

Cited by 0SourcePDFScholar
2025

OmniSR: Shadow Removal Under Direct and Indirect Lighting

AAAI 2025technical

Shadows can originate from occlusions in both direct and indirect illumination. Although most current shadow removal research focuses on shadows caused by direct illumination, shadows from indirect illumination are often just as pervasive, particularly in indoor scenes. A significant challenge in re…

2025

Real-Time Consistent Monocular Depth Recovery System for Dynamic Environments

IROS 2025

Monocular depth estimation is essential for applications such as autonomous navigation and 3D reconstruction. However, achieving accurate and temporally consistent depth estimation in dynamic environments remains challenging due to scale ambiguity, sensitivity to dynamic objects, and inconsistent de

Cited by 1SourceScholar
2025

SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations

ICCV 2025poster

Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-vie…

Cited by 0SourcePDFScholar
2025

UniTransfer: Video Concept Transfer via Progressive Spatio-Temporal Decomposition

NeurIPS 2025poster

Recent advancements in video generation models have enabled the creation of diverse and realistic videos, with promising applications in advertising and film production. However, as one of the essential tasks of video generation models, video concept transfer remains significantly challenging. Exist…

Cited by 0SourceScholar
2024

3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation

CVPR 2024poster

Text-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly attributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However these methods heavily rely on the outputs of existing models…

Cited by 8SourcePDFScholar
2024

Text-Guided 3D Face Synthesis - From Generation to Editing

CVPR 2024poster

Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However most existing works focus solely on the direct generation ignoring the editing restricting them from synthesizing customized 3D faces through iterative adjustments. In this paper…

2023

CF-Font: Content Fusion for Few-Shot Font Generation

CVPR 2023poster

Content and style disentanglement is an effective way to achieve few-shot font generation. It allows to transfer the style of the font image in a source domain to the style defined with a few reference images in a target domain. However, the content feature extracted using a representative font migh…

2023

Unsupervised Domain Adaption With Pixel-Level Discriminator for Image-Aware Layout Generation

CVPR 2023poster

Layout is essential for graphic design and poster generation. Recently, applying deep learning models to generate layouts has attracted increasing attention. This paper focuses on using the GAN-based model conditioned on image contents to generate advertising poster graphic layouts, which requires a…

Cited by 19SourcePDFScholar
2022

Active Boundary Loss for Semantic Segmentation

AAAI 2022technical

This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted b…

2022

Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs

IJCAI 2022poster

In this paper, we study the graphic layout generation problem of producing high-quality visual-textual presentation designs for given images. We note that image compositions, which contain not only global semantics but also spatial information, would largely affect layout results. Hence, we propose…

2022

Geometry-aware Two-scale PIFu Representation for Human Reconstruction

NeurIPS 2022accept

Although PIFu-based 3D human reconstruction methods are popular, the quality of recovered details is still unsatisfactory. In a sparse (e.g., 3 RGBD sensors) capture setting, the depth noise is typically amplified in the PIFu representation, resulting in flat facial surfaces and geometry-fallible bo…

Cited by 17SourcePDFScholar
2022

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

CVPR 2022poster

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene reconstructions and have difficulty scaling up to large scenes. These l…

Cited by 766PDFcodeScholar
2021

Location-Aware Single Image Reflection Removal

ICCV 2021poster

This paper proposes a novel location-aware deep-learning-based single image reflection removal method. Our network has a reflection detection module to regress a probabilistic reflection confidence map, taking multi-scale Laplacian features as inputs. This probabilistic map tells if a region is refl…

Cited by 110PDFcodeScholar