← Search

Xu He

23 accepted papers

2026

3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation

CVPR 2026

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel-view synthesis. Explicit 3D models, though structurally i

Cited by 0SourcecodeScholar
2026

A Content-Preserving Secure Linguistic Steganography

AAAI 2026technical

Existing linguistic steganography methods primarily rely on content transformations to conceal secret messages. However, they often cause subtle yet looking-innocent deviations between normal and stego texts, posing potential security risks in real-world applications. To address this challenge, we p

Cited by 1SourcePDFScholar
2026

AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic

ICML 2026poster

Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-symbolic construction of circuit structure from images and correct multi-step symbolic derivation, both of which…

Cited by 0SourceScholar
2026

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

ICML 2026poster

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spa…

Cited by 0SourceScholar
2026

Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

AAAI 2026technical

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty

Cited by 0SourcePDFScholar
2026

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

CVPR 2026

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and limited modal diversity for comprehensive world knowledge represe

Cited by 0SourcecodeScholar
2025

Conditional-Balanced Adversarial Delta Tuning for Cross-Domain Implicit Discourse Relation Recognition

ICASSP 2025accepted

Implicit discourse relation recognition (IDRR) is faced with a domain dilemma. Recent studies have achieved breakthroughs in standard datasets, while they are not appropriate in domains with insufficient data, such as bio-medicine. In this paper, we treat this problem as a cross-domain IDRR task, wh…

Cited by 0SourceScholar
2025

Forgetting Transformer: Softmax Attention with a Forget Gate

ICLR 2025poster

An essential component of modern recurrent sequence models is the forget gate. While Transformers do not have an explicit recurrent form, we show that a forget gate can be naturally incorporated into Transformers by down-weighting the unnormalized attention scores in a data-dependent way. We name th…

2025

MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement

AAAI 2025technical

Existing works in single-image human reconstruction suffer from weak generalizability due to insufficient training data or 3D inconsistencies for a lack of comprehensive multi-view knowledge. In this paper, we introduce MagicMan, a human-specific multi-view diffusion model to generate high-quality n…

Cited by 8SourcePDFScholar
2025

Trajectory World Models for Heterogeneous Environments

ICML 2025poster

Heterogeneity in sensors and actuators across environments poses a significant challenge to building large-scale pre-trained world models on top of this low-dimensional sensor information. In this work, we explore pre-training world models for heterogeneous environments by addressing key transfer ba…

2024

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

CVPR 2024poster

Co-speech gestures if presented in the lively form of videos can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons resulting in the omission of appearance information we focus on the direct generation of audio-driven co-spee…

2024

Enhancing Expressiveness in Dance Generation Via Integrating Frequency and Music Style Information

ICASSP 2024accepted

Dance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre matching, beat alignment, and dance dynamics, from certain aspects. However, the enhancement is quite limited as they lack…

Cited by 0SourceScholar
2024

Improving Unsupervised Hierarchical Representation with Reinforcement Learning

CVPR 2024poster

Learning representations to capture the very fundamental understanding of the world is a key challenge in machine learning. The hierarchical structure of explanatory factors hidden in data is such a general representation and could be potentially achieved with a hierarchical VAE. However training a…

2024

Manifold Constraints for Imperceptible Adversarial Attacks on Point Clouds

AAAI 2024technical

Adversarial attacks on 3D point clouds often exhibit unsatisfactory imperceptibility, which primarily stems from the disregard for manifold-aware distortion, i.e., distortion of the underlying 2-manifold surfaces. In this paper, we develop novel manifold constraints to reduce such distortion, aiming…

Cited by 11SourcePDFScholar
2024

PoRank: A Practical Framework for Learning to Rank Policies

IJCAI 2024poster

In many real-world scenarios, we need to select from a set of candidate policies before online deployment. Although existing Off-policy evaluation (OPE) methods can be used to estimate the online performance, they suffer from high variance. Fortunately, we care only about the ranking of the candidat…

2024

iVideoGPT: Interactive VideoGPTs are Scalable World Models

NeurIPS 2024poster

World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in video generative models for developing world models at scal…

2021

Personalized Adaptive Meta Learning for Cold-start User Preference Prediction

AAAI 2021technical

A common challenge in personalized user preference prediction is the cold-start problem. Due to the lack of user-item interactions, directly learning from the new users' log data causes serious over-fitting problem. Recently, many existing studies regard the cold-start personalized preference predic…

Cited by 76SourcePDFScholar
2021

RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning Agents

NeurIPS 2021poster

Current value-based multi-agent reinforcement learning methods optimize individual Q values to guide individuals' behaviours via centralized training with decentralized execution (CTDE). However, such expected, i.e., risk-neutral, Q value is not sufficient even with CTDE due to the randomness of rew…

Cited by 59SourcePDFScholar
2020

Learning Efficient Multi-agent Communication: An Information Bottleneck Approach

ICML 2020poster

We consider the problem of the limited-bandwidth communication for multi-agent reinforcement learning, where agents cooperate with the assistance of a communication protocol and a scheduler. The protocol and scheduler jointly determine which agent is communicating what message and to whom. Under the…

2017

Ego-centric traffic behavior understanding through multi-level vehicle trajectory analysis

ICRA 2017poster

This study proposes a multi-level trajectory analysis method for modeling traffic behavior from an ego-centric view, where on-road vehicle trajectories are collected based on the authors' previous studies of an on-board system consisting of multiple 2D lidar sensors. From an input set of trajectorie…

Cited by 6SourceScholar