ICML 2026poster0 citations

MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities

Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, Oğuzhan Fatih Kar, Jesse Allardice, Roman Bachmann, David Mizrahi

Abstract

Any-to-any modeling aims to flexibly relate arbitrary modalities within a single system, a requirement that arises across multimodal learning and scientific domains such as ecology and astronomy. However, existing any-to-any approaches are typically trained from scratch using encoder–decoder or diffusion architectures, limiting empirical performance and the use of pretrained models. We investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. As a consequence of this unified design, the resulting model MODUS naturally enables chained generation through intermediate modalities, cross-modal consistency verification, and analysis of visual representations by combining semantic and reconstruction features. Across a range of benchmarks, MODUS demonstrates strong out-of-the-box performance and flexible multimodal composition within a single model.

DiffusionVisionMultimodalBenchmark
BibTeX
@inproceedings{
ye2026modus,
title={{MODUS}: Decoder-only Any-to-Any Modeling of Diverse Modalities},
author={Mingqiao Ye and Zhaochong An and Zhitong Gao and Xian Liu and O{\u{g}}uzhan Fatih Kar and Jesse Allardice and Roman Bachmann and David Mizrahi and Fran{\c{c}}ois Fleuret and Chuan Li and Amir Zadeh and Serge Belongie and Afshin Dehghan and Amir Zamir},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=t7kTpjsRIw}
}