← Search

Xiaojie Xu

12 accepted papers

2026

PHOTONS: Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views

AAAI 2026technical

We present PHOTONS (Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views), a real-time framework for novel view synthesis without requiring camera calibration. Our method reconstructs consistent 3D Gaussian point clouds and synthesizes 2K photo-realistic novel vie

Cited by 0SourcePDFScholar
2026

SVBench: Evaluation of Video Generation Models on Social Reasoning

CVPR 2026

Recent text-to-video generation models have made remarkable progress in visual realism, motion fidelity, and text-video alignment, yet they still struggle to produce socially coherent behavior. Unlike humans, who readily infer intentions, beliefs, emotions, and social norms from brief visual cues, c

Cited by 0SourcecodeScholar
2026

Yume1.5: A Text-Controlled Interactive World Generation Model

CVPR 2026

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, whi

Cited by 0SourcecodeScholar
2025

Orchestrating Audio: Multi-Agent Framework for Long-Video Audio Synthesis

EMNLP 2025

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an unsolved challenge due to dynamic semantic shifts, audio diver

Cited by 0SourcePDFScholar
2025

POSTA: A Go-to Framework for Customized Artistic Poster Generation

CVPR 2025poster

Poster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appeal, limiting their applicability in artistic domains such as movies and exhibiti…

Cited by 4SourcePDFScholar
2025

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

EMNLP 2025

Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding, leading to mismatched visuals and text. These limitations r

Cited by 0SourcePDFScholar
2024

Every Dataset Counts: Scaling up Monocular 3D Object Detection with Joint Datasets Training

IROS 2024poster

Monocular 3D object detection is essential for autonomous driving. However, current monocular 3D detection algorithms rely on expensive 3D labels from LiDAR scans, making it difficult to use in new datasets and unfamiliar environments. This study explores training a monocular 3D object detection mod…

Cited by 4SourceScholar
2024

From Bird’s-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model

ICRA 2024poster

We explore Bird’s-Eye View (BEV) generation, converting a BEV map into its corresponding multi-view street images. Valued for its unified spatial representation aiding multi-sensor fusion, BEV is pivotal for various autonomous driving applications. Creating accurate street-view images from BEV maps…

Cited by 1SourceScholar
2024

HPFF: Hierarchical Locally Supervised Learning with Patch Feature Fusion

ECCV 2024poster

"Traditional deep learning relies on end-to-end backpropagation for training, but it suffers from drawbacks such as high memory consumption and not aligning with biological neural networks. Recent advancements have introduced locally supervised learning, which divides networks into modules with isol…

2024

Momentum Auxiliary Network for Supervised Local Learning

ECCV 2024oral

"Deep neural networks conventionally employ end-to-end backpropagation for their training process, which lacks biological credibility and triggers a locking dilemma during network parameter updates, leading to significant GPU memory use. Supervised local learning, which segments the network into mul…

2021

3DCaricShop: A Dataset and a Baseline Method for Single-View 3D Caricature Face Reconstruction

CVPR 2021poster

Caricature is an artistic representation that deliberately exaggerates the distinctive features of a human face to convey humor or sarcasm. However, reconstructing a 3D caricature from a 2D caricature image remains a challenging task, mostly due to the lack of data. We propose to fill this gap by in…

Cited by 27PDFScholar