← Search

Xiaogang Jin

18 accepted papers

2026

Unifying Precise Keyframes and Semantic Control via Multi-level Diffusion

CVPR 2026

Text-conditioned human motion in-betweening leverages keyframes for spatio-temporal control, with text providing high-level semantic guidance for the transitions. However, existing methods are unable to establish a coherent alignment between textual semantics and the spatio-temporal constraints prov

Cited by 0SourceScholar
2025

Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesis

CVPR 2025poster

Sewing patterns, the essential blueprints for fabric cutting and tailoring, act as a crucial bridge between design concepts and producible garments. However, existing uni-modal sewing pattern generation models struggle to effectively encode complex design concepts with a multi-modal nature and corre…

2025

POMP: Physics-consistent Motion Generative Model through Phase Manifolds

CVPR 2025poster

Numerous researches on real-time motion generation primarily focus on kinematic aspects, often resulting in physically implausible outcomes. In this paper, we present POMP ("\underline P hysics-c\underline O nsistent Human \underline M otion \underline P rior through Phase Manifolds"), a novel kinem…

Cited by 0SourcePDFScholar
2025

Towards Realistic Example-based Modeling via 3D Gaussian Stitching

CVPR 2025poster

Using parts of existing models to rebuild new models, commonly termed as example-based modeling, is a classical methodology in the realm of computer graphics. Previous works mostly focus on shape composition, making them very hard to use for realistic composition of 3D objects captured from real-wor…

Cited by 1SourcePDFScholar
2024

A General Implicit Framework for Fast NeRF Composition and Rendering

AAAI 2024technical

A variety of Neural Radiance Fields (NeRF) methods have recently achieved remarkable success in high render speed. However, current accelerating methods are specialized and incompatible with various implicit methods, preventing real-time composition over various types of NeRF works. Because NeRF rel…

Cited by 3SourcePDFScholar
2024

Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction

CVPR 2024poster

Implicit neural representation has paved the way for new approaches to dynamic scene reconstruction. Nonetheless cutting-edge dynamic neural rendering methods rely heavily on these implicit representations which frequently struggle to capture the intricate details of objects in the scene. Furthermor…

2024

MaskFactory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation

NeurIPS 2024poster

Dichotomous Image Segmentation (DIS) tasks require highly precise annotations, and traditional dataset creation methods are labor intensive, costly, and require extensive domain expertise. Although using synthetic data for DIS is a promising solution to these challenges, current generative models an…

2024

RobIR: Robust Inverse Rendering for High-Illumination Scenes

NeurIPS 2024poster

Implicit representation has opened up new possibilities for inverse rendering. However, existing implicit neural inverse rendering methods struggle to handle strongly illuminated scenes with significant shadows and slight reflections. The existence of shadows and reflections can lead to an inaccurat…

Cited by 0SourcePDFScholar
2024

SocialCVAE: Predicting Pedestrian Trajectory via Interaction Conditioned Latents

AAAI 2024technical

Pedestrian trajectory prediction is the key technology in many applications for providing insights into human behavior and anticipating human future motions. Most existing empirical models are explicitly formulated by observed human behaviors using explicable mathematical terms with deterministic na…

2024

Spec-Gaussian: Anisotropic View-Dependent Appearance for 3D Gaussian Splatting

NeurIPS 2024poster

The recent advancements in 3D Gaussian splatting (3D-GS) have not only facilitated real-time rendering through modern GPU rasterization pipelines but have also attained state-of-the-art rendering quality. Nevertheless, despite its exceptional rendering quality and performance on standard datasets, 3…

Cited by 45SourcePDFScholar
2022

BRNet: Exploring Comprehensive Features for Monocular Depth Estimation

ECCV 2022poster

"Self-supervised monocular depth estimation has achieved promising performance recently. A consensus is that high-resolution inputs often yield better results. However, we find that the performance gap between high and low resolutions lies in the inappropriate feature representation of the widely us…

2022

TraEDITS: Diversity and Irregularity-Aware Traffic Trajectory Editing

RA-L 2022

We present TraEDITS, a novel traffic trajectory editing framework for autonomous vehicle testing, which can generate new traffic behaviors by controlling each vehicle interactively to increase the diversity or irregularity of traffic testing data. Given a traffic flow with its original trajectories,

Cited by 4SourceScholar
2019

Force-based Heterogeneous Traffic Simulation for Autonomous Vehicle Testing

ICRA 2019poster

Recent failures in real-world self-driving tests have suggested a paradigm shift from directly learning in real-world roads to building a high-fidelity driving simulator as an alternative, effective, and safe tool to handle intricate traffic environments in urban areas. To date, traffic simulation c…

Cited by 44SourceScholar
2016

Steering micro-robotic swarm by dynamic actuating fields

ICRA 2016

We present a general solution for steering microrobotic swarm by dynamic actuating fields. In our approach, the motion of micro-robots is controlled by changing the actuating direction of a field applied to them. The time-series sequence of actuating field's directions can be computed automatically.

Cited by 7SourceScholar