← Search

Eldad Klaiman

1 accepted papers

2023

Training Transitive and Commutative Multimodal Transformers with LoReTTa

NeurIPS 2023poster

Training multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images with audio or text with audio. Even rarer are datasets that align all three modalities at once. Critical domains such as h…

Cited by 3SourcePDFScholar