ICML 2026poster0 citations

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

Donghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li, Jingyu Lin, Xiaohu Huang, Yichen Liu, Xin Gao

Abstract

In this work, we study **Human-Object Interaction Video Generation (HOIVG)**, which aims to synthesize high-quality HOI videos via text, reference image, audio, and pose conditions. To address the challenges of harmonious multimodal injection and heterogeneous data utility, we present **OmniShow**, an end-to-end framework tailored for HOIVG. We introduce *Unified Channel-wise Conditioning* to efficiently inject image and pose cues, *Gated Local-Context Attention* to ensure precise audio-visual synchronization, and a *Decoupled-then-Joint Training strategy* to effectively harness heterogeneous data. Extensive experiments on the proposed *HOIVG-Bench* demonstrate that OmniShow achieves state-of-the-art performance.

TransformerVisionMultimodal
BibTeX
@inproceedings{
zhou2026omnishow,
title={OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation},
author={Donghao Zhou and Guisheng Liu and Hao Yang and Jiatong Li and Jingyu Lin and Xiaohu Huang and Yichen Liu and Xin Gao and Cunjian Chen and Shilei Wen and Chi-Wing Fu and Pheng-Ann Heng},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=uNWexq8gQB}
}