← Search

Yuxuan Xue

10 accepted papers

2026

FastVMT: Eliminating Redundancy in Video Motion Transfer

ICLR 2026poster

Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent methods predominantly use the Diffusion Transformer (DiT) architecture. To achieve satisfactory runtime, several methods…

Cited by 0SourceScholar
2026

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

CVPR 2026

Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry du

Cited by 0SourceScholar
2026

Human3R: Everyone Everywhere All at Once

ICLR 2026poster

We present Human3R, a unified, feed-forward framework for online 4D human-scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact-aware refinement between humans and scenes, and heavy dependen…

Cited by 0SourcecodeScholar
2025

New Network Protocol for Supermedia-Enhanced Telerobotics

IROS 2025

The growing complexity of robotic teleoperation systems necessitates the integration of multiple feedback modalities, including video, audio, force, tactile, and temperature feedback. The concept of supermedia is utilized to describe the aggregation of these feedback streams. By integrating multiple

Cited by 0SourceScholar
2024

Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion Models

NeurIPS 2024poster

Creating realistic avatars from a single RGB image is an attractive yet challenging problem. To deal with challenging loose clothing or occlusion by interaction objects, we leverage powerful shape prior from 2D diffusion models pretrained on large datasets. Although 2D diffusion models demonstrate s…

Cited by 5SourcePDFScholar
2024

Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization

ICLR 2024poster

Large foundation models are becoming ubiquitous, but training them from scratch is prohibitively expensive. Thus, efficiently adapting these powerful models to downstream tasks is increasingly important. In this paper, we study a principled finetuning paradigm -- Orthogonal Finetuning (OFT) -- for d…

Cited by 57SourcePDFScholar
2023

Controlling Text-to-Image Diffusion by Orthogonal Finetuning

NeurIPS 2023poster

Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to perform different downstream tasks becomes an important open problem. To tackle this challenge, we introduce a princip…

Cited by 123SourcePDFScholar
2023

NSF: Neural Surface Fields for Human Modeling from Monocular Depth

ICCV 2023poster

Obtaining personalized 3D animatable avatars from a monocular camera has several real world applications in gaming, virtual try-on, animation, and VR/XR, etc. However, it is very challenging to model dynamic and fine-grained clothing deformations from such sparse data. Existing methods for modeling…

Cited by 15PDFScholar
2021

Robust Event Detection based on Spatio-Temporal Latent Action Unit using Skeletal Information

IROS 2021poster

This paper proposes a novel dictionary learning approach to detect event anomalities using skeletal information extracted from RGBD video. The event action is represented as several latent action atoms and composed of latent spatial and temporal attributes. We aim to construct a network able to lear…

Cited by 6SourceScholar
2020

Task Space Motion Control for AFM-Based Nanorobot Using Optimal and Ultralimit Archimedean Spiral Local Scan

RA-L 2020

Atomic force microscopy (AFM) based nanorobotic technology provides a unique manner for delicate operations at the nanoscale in various ambient, thanks to its ultrahigh spatial resolution, outstanding environmental adaptability, and numerous measurement approaches. However, one vital challenge behin

Cited by 9SourceScholar