OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Zhen Wan, Jinchuan Tian, An-Chieh Cheng, Ligeng Zhu
Abstract
Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curation. For model architecture, we present three key innovations: (i) OmniAlignNet for strengthening alignment between vision and audio embeddings in a shared omni-modal latent space; (ii) Temporal Embedding Grouping for capturing relative temporal alignment between vision and audio signals; and (iii) Constrained Rotary Time Embedding for encoding absolute temporal information in omni-modal embeddings. We introduce a curation and synthesis pipeline that generates 24M single-modal and omni-modal conversations. We find that modalities reinforce one another in both perception and reasoning. Our model, OmniVinci, improves over Qwen2.5-Omni with +19.05 on DailyOmni (cross-modal understanding), +1.7 on MMAR (audio), and +3.9 on Video-MME (vision), while using just 0.2T training tokens — a 6× reduction compared to Qwen2.5-Omni’s 1.2T. We finally demonstrate omni-modal advantages in downstream applications spanning robotics, medical AI, and smart factory.
BibTeX
@inproceedings{
ye2026omnivinci,
title={OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding {LLM}},
author={Hanrong Ye and Chao-Han Huck Yang and Arushi Goel and Wei Huang and Zhen Wan and Jinchuan Tian and An-Chieh Cheng and Ligeng Zhu and Yuanhang Su and Yuming Lou and Yong-Xiang Lin and Dong Yang and Sreyan Ghosh and Zhijian Liu and Yukang Chen and Ehsan Jahangiri and Ambrish Dantrey and Daguang Xu and Ehsan Hosseini-Asl and Seyed Danial Mohseni Taheri and Vidya Nariyambut Murali and Sifei Liu and Yao Lu and Oluwatobi Olabiyi and Yu-Chiang Frank Wang and Rafael Valle and Bryan Catanzaro and Andrew Tao and Song Han and Jan Kautz and Hongxu Yin and Pavlo Molchanov},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=DZeic3NpHy}
}