← Search

Siyu Tang

69 accepted papers

2026

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

CVPR 2026

Emerging video diffusion models achieve high visual fidelity but fundamentally couple scene dynamics with camera motion, limiting their ability to provide precise spatial and temporal control. We introduce a 4D-controllable video diffusion framework that explicitly decouples scene dynamics from came

Cited by 0SourceScholar
2026

Feed-forward Gaussian Registration for Head Avatar Creation and Editing

CVPR 2026

We present MATCH (Multi-view Avatars from Topologically Corresponding Heads), a multi-view Gaussian registration method for high-quality head avatar creation and editing. State-of-the-art multi-view head avatars require time-consuming head tracking, which is followed by an expensive avatar optimizat

Cited by 0SourcecodeScholar
2026

TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrary-Shaped Modular Aerial Robot Systems

ICRA 2026poster

Modular Aerial Robot Systems (MARS) consist of multiple drone modules that are physically bound together to form a single structure for flight. Exploiting structural redundancy, MARS can be reconfigured into different formations to mitigate unit or rotor failures and maintain stable flight. Prior wo…

Cited by 0codeScholar
2025

DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control

ICLR 2025spotlight

Text-conditioned human motion generation, which allows for user interaction through natural language, has become increasingly popular. Existing methods typically generate short, isolated motions based on a single input sentence. However, human motions are continuous and can extend over long periods,…

2025

DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-free 3D Reconstruction

ICCV 2025poster

Reconstructing clean, distractor-free 3D scenes from real-world captures remains a significant challenge, particularly in highly dynamic and cluttered settings such as egocentric videos. To tackle this problem, we introduce DeGauss, a simple and robust self-supervised framework for dynamic scene rec…

Cited by 0SourcePDFScholar
2025

EgoM2P: Egocentric Multimodal Multitask Pretraining

ICCV 2025accepted

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the camera wearer's actions, intentions, and surrounding environ…

Cited by 0SourcePDFScholar
2025

Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting

NeurIPS 2025poster

Recent advances in feed-forward 3D Gaussian Splatting have led to rapid improvements in efficient scene reconstruction from sparse views. However, most existing approaches construct Gaussian primitives directly aligned with the pixels in one or more of the input images. This leads to redundancies in…

Cited by 0SourceScholar
2025

MARS-FTCP: Robust Fault-Tolerant Control and Agile Trajectory Planning for Modular Aerial Robot Systems

IROS 2025

Modular Aerial Robot Systems (MARS) consist of multiple drone units that can self-reconfigure to adapt to various mission requirements and fault conditions. However, existing fault-tolerant control methods exhibit significant oscillations during docking and separation, impacting system stability. To

Cited by 4SourcecodeScholar
2025

Multi-View 3D Point Tracking

ICCV 2025poster

We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle with depth ambiguities and occlusion, or prior multi-camera methods that require over 20 cameras and te…

2025

Robust Self-Reconfiguration for Fault-Tolerant Control of Modular Aerial Robot Systems

ICRA 2025

Modular Aerial Robotic Systems (MARS) consist of multiple drone units assembled into a single, integrated rigid flying platform. With inherent redundancy, MARS can self-reconfigure into different configurations to mitigate rotor or unit failures and maintain stable flight. However, existing works on

Cited by 9SourcecodeScholar
2025

SplatFormer: Point Transformer for Robust 3D Gaussian Splatting

ICLR 2025spotlight

3D Gaussian Splatting (3DGS) has recently transformed photorealistic reconstruction, achieving high visual fidelity and real-time performance. However, rendering quality significantly deteriorates when test views deviate from the camera angles used during training, posing a major challenge for appli…

2025

UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control

ICCV 2025poster

Generating natural and physically plausible character motion remains challenging, particularly for long-horizon control with diverse guidance signals. While prior work combines high-level diffusion-based motion planners with low-level physics controllers, these systems suffer from domain gaps that d…

Cited by 0SourcePDFScholar
2025

VolumetricSMPL: A Neural Volumetric Body Model for Efficient Interactions, Contacts, and Collisions

ICCV 2025poster

Parametric human body models play a crucial role in computer graphics and vision, enabling applications ranging from human motion analysis to understanding human-environment interactions. Traditionally, these models use surface meshes, which pose challenges in efficiently handling interactions with…

2024

3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting

CVPR 2024poster

We introduce an approach that creates animatable human avatars from monocular videos using 3D Gaussian Splatting (3DGS). Existing methods based on neural radiance fields (NeRFs) achieve high-quality novel-view/novel-pose image synthesis but often require days of training and are extremely slow at in…

Cited by 123SourcePDFScholar
2024

Degrees of Freedom Matter: Inferring Dynamics from Point Trajectories

CVPR 2024poster

Understanding the dynamics of generic 3D scenes is fundamentally challenging in computer vision essential in enhancing applications related to scene reconstruction motion tracking and avatar creation. In this work we address the task as the problem of inferring dense long-range motion of 3D points.…

2024

EgoGen: An Egocentric Synthetic Data Generator

CVPR 2024poster

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered third-person-view vision models but its application to embodied egocentr…

Cited by 17SourcePDFScholar
2024

Improving 2D Feature Representations by 3D-Aware Fine-Tuning

ECCV 2024poster

"Current visual foundation models are trained purely on unstructured 2D data, limiting their understanding of 3D structure of objects and scenes. In this work, we show that fine-tuning on 3D-aware data improves the quality of emerging semantic features. We design a method to lift semantic 2D feature…

2024

IntrinsicAvatar: Physically Based Inverse Rendering of Dynamic Humans from Monocular Videos via Explicit Ray Tracing

CVPR 2024poster

We present IntrinsicAvatar a novel approach to recovering the intrinsic properties of clothed human avatars including geometry albedo material and environment lighting from only monocular videos. Recent advancements in human-based neural rendering have enabled high-quality geometry and appearance re…

Cited by 12SourcePDFScholar
2024

LaRa: Efficient Large-Baseline Radiance Fields

ECCV 2024poster

"Radiance field methods have achieved photorealistic novel view synthesis and geometry reconstruction. But they are mostly applied in per-scene optimization or small-baseline settings. While several recent works investigate feed-forward reconstruction with large baselines by utilizing transformers,…

2024

Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar Creation

CVPR 2024poster

Recent advances in generative diffusion models have enabled the previously unfeasible capability of generating 3D assets from a single input image or a text prompt. In this work we aim to enhance the quality and functionality of these models for the task of creating controllable photorealistic human…

2024

Optimizing Diffusion Noise Can Serve As Universal Motion Priors

CVPR 2024poster

We propose Diffusion Noise Optimization (DNO) a new method that effectively leverages existing motion diffusion models as motion priors for a wide range of motion-related tasks. Instead of training a task-specific diffusion model for each new task DNO operates by optimizing the diffusion latent nois…

Cited by 41SourcePDFScholar
2024

ResFields: Residual Neural Fields for Spatiotemporal Signals

ICLR 2024spotlight

Neural fields, a category of neural networks trained to represent high-frequency signals, have gained significant attention in recent years due to their impressive performance in modeling complex 3D data, such as signed distance (SDFs) or radiance fields (NeRFs), via a single multi-layer perceptron…

2024

RoHM: Robust Human Motion Reconstruction via Diffusion

CVPR 2024poster

We propose RoHM an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn data-driven motion priors and combine them with optimization at…

2024

SplatFields: Neural Gaussian Splats for Sparse 3D and 4D Reconstruction

ECCV 2024poster

"Digitizing 3D static scenes and 4D dynamic events from multi-view images has long been a challenge in computer vision and graphics. Recently, 3D Gaussian Splatting (3DGS) has emerged as a practical and scalable reconstruction method, gaining popularity due to its impressive reconstruction quality,…

Cited by 16SourcePDFScholar
2023

3D Segmentation of Humans in Point Clouds with Synthetic Data

ICCV 2023poster

Segmenting humans in 3D indoor scenes has become increasingly important with the rise of human-centered robotics and AR/VR applications. To this end, we propose the task of joint 3D human semantic segmentation, instance segmentation and multi-human body-part segmentation. Few works have attempted to…

Cited by 29PDFScholar
2023

Aspect-to-Scope Oriented Multi-view Contrastive Learning for Aspect-based Sentiment Analysis

EMNLP 2023long findings

Aspect-based sentiment analysis (ABSA) aims to align aspects and corresponding sentiment expressions, so as to identify the sentiment polarities of specific aspects. Most existing ABSA methods focus on mining syntactic or semantic information, which still suffers from noisy interference introduced b…

Cited by 0SourceScholar
2023

Guided Motion Diffusion for Controllable Human Motion Synthesis

ICCV 2023poster

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge despite being essential for bridging the gap between isolat…

Cited by 126PDFScholar
2023

HARP: Personalized Hand Reconstruction From a Monocular RGB Video

CVPR 2023poster

We present HARP (HAnd Reconstruction and Personalization), a personalized hand avatar creation approach that takes a short monocular RGB video of a human hand as input and reconstructs a faithful hand avatar exhibiting a high-fidelity appearance and geometry. In contrast to the major trend of neural…

Cited by 31SourcePDFScholar
2023

Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

ICRA 2023poster

Modern 3D semantic instance segmentation approaches predominantly rely on specialized voting mechanisms followed by carefully designed geometric clustering techniques. Building on the successes of recent Transformer-based methods for object detection and image segmentation, we propose the first Tran…

Cited by 272SourcecodeScholar
2023

Probabilistic Human Mesh Recovery in 3D Scenes from Egocentric Views

ICCV 2023oral

Automatic perception of human behaviors during social interactions is crucial for AR/VR applications, and an essential component is estimation of plausible 3D human pose and shape of our social partners from the egocentric view. One of the biggest challenges of this task is severe body truncation du…

Cited by 30PDFcodeScholar
2023

Synthesizing Diverse Human Motions in 3D Indoor Scenes

ICCV 2023poster

We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on high-quality training sequences that contain captured human motions and the 3D scenes they interact with. Ho…

Cited by 66PDFcodeScholar
2022

ARAH: Animatable Volume Rendering of Articulated Human SDFs

ECCV 2022poster

"Combining human body models with differentiable rendering has recently enabled animatable avatars of clothed humans from sparse sets of multi-view RGB videos. While state-of-the-art approaches achieve a realistic appearance with neural radiance fields (NeRF), the inferred geometry often lacks detai…

Cited by 150SourcePDFScholar
2022

Accurate 3D Body Shape Regression Using Metric and Semantic Attributes

CVPR 2022oral

While methods that regress 3D human meshes from images have progressed rapidly, the estimated body shapes often do not capture the true human shape. This is problematic since, for many applications, accurate body shape is as important as pose. The key reason that body shape accuracy lags pose accura…

Cited by 72PDFcodeScholar
2022

Affective Knowledge Enhanced Multiple-Graph Fusion Networks for Aspect-based Sentiment Analysis

EMNLP 2022main

Aspect-based sentiment analysis aims to identify sentiment polarity of social media users toward different aspects. Most recent methods adopt the aspect-centric latent tree to connect aspects and their corresponding opinion words, thinking that would facilitate establishing the relationship between…

2022

COAP: Compositional Articulated Occupancy of People

CVPR 2022poster

We present a novel neural implicit representation for articulated human bodies. Compared to explicit template meshes, neural implicit body representations provide an efficient mechanism for modeling interactions with the environment, which is essential for human motion reconstruction and synthesis i…

Cited by 57PDFcodeScholar
2022

Compositional Human-Scene Interaction Synthesis with Semantic Control

ECCV 2022poster

"Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Recent methods mainly focus on modeling geometric relations between 3D environments and humans, where the high-level semantics of t…

2022

EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices

ECCV 2022poster

"Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner from the egocentric view. However, research in this area is se…

2022

Human-Aware Object Placement for Visual Environment Reconstruction

CVPR 2022poster

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate that these human-scene interactions (HSIs) can be leveraged to…

Cited by 70PDFcodeScholar
2022

Improving Anomaly Detection with a Self-Supervised Task Based on Generative Adversarial Network

ICASSP 2022accepted

Existing anomaly detection models show success in detecting abnormal images with generative adversarial networks on the insufficient annotation of anomalous samples. However, existing models cannot accurately identify the anomaly samples which are close to the normal samples. We assume that the main…

Cited by 0SourceScholar
2022

Improving Multi-task Stance Detection with Multi-task Interaction Network

EMNLP 2022main

Stance detection aims to identify people’s standpoints expressed in the text towards a target, which can provide powerful information for various downstream tasks.Recent studies have proposed multi-task learning models that introduce sentiment information to boost stance detection.However, they negl…

2022

KeypointNeRF: Generalizing Image-Based Volumetric Avatars Using Relative Spatial Encoding of Keypoints

ECCV 2022poster

"Image-based volumetric avatars using pixel-aligned features promise generalization to unseen poses and identities. Prior work leverages global spatial encodings and multi-view geometric consistency to reduce spatial ambiguity. However, global encodings often suffer from overfitting to the distribut…

2022

SAGA: Stochastic Whole-Body Grasping with Contact

ECCV 2022poster

"The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction for object grasping and manipulation, these typically only consider interacting hand alone. Our goal is to synthesize w…

2022

The Wanderings of Odysseus in 3D Scenes

CVPR 2022poster

Our goal is to populate digital environments, in which digital humans have diverse body shapes, move perpetually, and have plausible body-scene contact. The core challenge is to generate realistic, controllable, and infinitely long motions for diverse 3D bodies. To this end, we propose generative mo…

Cited by 53PDFScholar
2021

Learning Motion Priors for 4D Human Body Capture in 3D Scenes

ICCV 2021poster

Recovering high-quality 3D human motion in complex scenes from monocular videos is important for many applications, ranging from AR/VR to robotics. However, capturing realistic human-scene interactions, while dealing with occlusions and partial views, is challenging; current approaches are still far…

Cited by 116PDFcodeScholar
2021

MetaAvatar: Learning Animatable Clothed Human Models from Few Depth Images

NeurIPS 2021poster

In this paper, we aim to create generalizable and controllable neural signed distance fields (SDFs) that represent clothed humans from monocular depth observations. Recent advances in deep learning, especially neural implicit representations, have enabled human shape reconstruction and controllable…

2021

SCALE: Modeling Clothed Humans with a Surface Codec of Articulated Local Elements

CVPR 2021poster

Learning to model and reconstruct humans in clothing is challenging due to articulation, non-rigid deformation, and varying clothing types and topologies. To enable learning, the choice of representation is the key. Recent work uses neural networks to parameterize local surface elements. This approa…

Cited by 114PDFcodeScholar
2020

ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose and Gaze Variation

ECCV 2020poster

Gaze estimation is a fundamental task in many applications of computer vision, human computer interaction and robotics. Many state-of-the-art methods are trained and tested on custom datasets, making comparison across methods challenging. Furthermore, existing gaze estimation datasets have limited h…

2020

Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video

ECCV 2020poster

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods either ignore how the camera wearer interacts with objects, or simply considers body motion as a separate modality. In contrast, we observe that the intentional hand movement reveal…

2020

Learning to Dress 3D People in Generative Clothing

CVPR 2020poster

Three-dimensional human body models are widely used in the analysis of human pose and motion. Existing models, however, are learned from minimally-clothed 3D scans and thus do not generalize to the complexity of dressed people in common images and videos. Additionally, current models lack the expres…

Cited by 435PDFcodeScholar
2020

MATE: Plugging in Model Awareness to Task Embedding for Meta Learning

NeurIPS 2020poster

Meta-learning improves generalization of machine learning models when faced with previously unseen tasks by leveraging experiences from different, yet related prior tasks. To allow for better generalization, we propose a novel task representation called model-aware task embedding (MATE) that incorpo…

2019

Local Temporal Bilinear Pooling for Fine-Grained Action Parsing

CVPR 2019poster

Fine-grained temporal action parsing is important in many applications, such as daily activity understanding, human motion analysis, surgical robotics and others requiring subtle and precise operations over a long-term period. In this paper we propose a novel bilinear pooling operation, which is use…

Cited by 32PDFScholar
2018

Part-Aligned Bilinear Representations for Person Re-Identification

ECCV 2018poster

Comparing the appearance of corresponding body parts is essential for person re-identification. As body parts are frequently misaligned between the detected human boxes, an image representation that can handle this misalignment is required. In this paper, we propose a network that learns a part-alig…

Cited by 671SourcePDFScholar
2017

ArtTrack: Articulated Multi-Person Tracking in the Wild

CVPR 2017oral

In this paper we propose an approach for articulated tracking of multiple people in unconstrained videos. Our starting point is a model that resembles existing architectures for single-frame pose estimation but is substantially faster. We achieve this in two ways: (1) by simplifying and sparsifying…

Cited by 379PDFScholar
2017

Generating Descriptions With Grounded and Co-Referenced People

CVPR 2017poster

Learning how to generate descriptions of images or videos received major interest both in the Computer Vision and Natural Language Processing communities. While a few works have proposed to learn a grounding during the generation process in an unsupervised way (via an attention mechanism), it remain…

Cited by 76PDFScholar
2017

Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications

CVPR 2017poster

We state a combinatorial optimization problem whose feasible solutions define both a decomposition and a node labeling of a given graph. This problem offers a common mathematical abstraction of seemingly unrelated computer vision tasks, including instance-separating semantic segmentation, articulate…

Cited by 131PDFcodeScholar
2017

Multiple People Tracking by Lifted Multicut and Person Re-Identification

CVPR 2017poster

Tracking multiple persons in a monocular video of a crowded scene is a challenging task. Humans can master it even if they loose track of a person locally by re-identifying the same person based on their appearance. Care must be taken across long distances, as similar-looking persons need not be ide…

Cited by 703PDFScholar
2016

DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation

CVPR 2016spotlight

This paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts…

Cited by 1446PDFScholar