← Search

Sirui Xu

14 accepted papers

2026

BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning

ICLR 2026poster

Recent advances in Vision-Language Models (VLMs) have propelled embodied agents by enabling direct perception, reasoning, and planning task-oriented actions from visual inputs. However, such vision-driven embodied agents open a new attack surface: visual backdoor attacks, where the agent behaves no…

Cited by 0SourcecodeScholar
2026

HandX: Scaling Bimanual Motion and Interaction Generation

CVPR 2026

Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing, and inter-hand coordination, and existing resources lack hig

Cited by 0SourcecodeScholar
2026

InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

CVPR 2026

Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors

Cited by 0SourceScholar
2026

Unleashing Guidance Without Classifiers for Human-Object Interaction Animation

ICLR 2026poster

Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on handcrafted contact priors or human-imposed kinematic constraints to improve con…

Cited by 0SourceScholar
2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference Scoped Exploration

CoRL 2025poster

Hand–object motion-capture (MoCap) repositories provide abundant, contact-rich human demonstrations for scaling dexterous manipulation on robots. Yet demonstration inaccuracy and embodiment gaps between human and robot hands challenge direct policy learning. Existing pipelines adapt a three-stage wo…

Cited by 11SourceScholar
2025

InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

CVPR 2025poster

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack extensive, high-quality motion and annotation and exhibit artifacts s…

Cited by 2SourcePDFScholar
2025

InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions

CVPR 2025highlight

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate human-object coupling, variability in object geometries, and artif…

2024

InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction

NeurIPS 2024poster

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic human-object interaction (HOI) generation faces notable challenges, pr…

Cited by 23SourcePDFScholar
2023

InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion

ICCV 2023poster

This paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or static objects. Our task is significantly more challenging, as…

Cited by 116PDFcodeScholar
2022

Diverse Human Motion Prediction Guided by Multi-level Spatial-Temporal Anchors

ECCV 2022poster

"Predicting diverse human motions given a sequence of historical poses has received increasing attention. Despite rapid progress, existing work captures the multi-modal nature of human motions primarily through likelihood-based sampling, where the mode collapse has been widely observed. In this pape…

2019

Spatial and Channel Attention Based Convolutional Neural Networks for Modeling Noisy Speech

ICASSP 2019accepted

In recent years, Residual Networks (ResNets) have significantly increased the modeling power of convolutional neural networks (CNNs) by introducing residual connections. In this paper, we explore the incorporation of spatial and channel attention into the structure of ResNets for noisy speech recogn…

Cited by 0SourceScholar
2018

Application of Progressive Neural Networks for Multi-Stream Wfst Combination in One-Pass Decoding

ICASSP 2018accepted

Many state-of-the-art automatic speech recognition (ASR) systems adopt system combination techniques to improve recognition performance. In this paper, we investigate the possibility of transferring knowledge between models for different noisy speech domains and integrating these models via system c…

Cited by 0SourceScholar