AAAI 2025technical0 citations

Open-World Multimodal Understanding and Generation with Efficiently Finetuned Foundation Models

Long Chen

Abstract

With the astonishing ability of different pretrained foundation models (e.g., large language models (LLMs), vision-language models, diffusion models), today’s AI research and development tendency has been revolutionized. In this talk, I will answer two questions: Q1: How can we efficiently train or fine-tune foundation models? Q2: How can we build strong open-world multimodal understanding and generation models with these pretrained foundation models?

BibTeX
@article{Chen_2025, title={Open-World Multimodal Understanding and Generation with Efficiently Finetuned Foundation Models}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/35101}, DOI={10.1609/aaai.v39i27.35101}, abstractNote={With the astonishing ability of different pretrained foundation models (e.g., large language models (LLMs), vision-language models, diffusion models), today’s AI research and development tendency has been revolutionized. In this talk, I will answer two questions: Q1: How can we efficiently train or fine-tune foundation models? Q2: How can we build strong open-world multimodal understanding and generation models with these pretrained foundation models?}, number={27}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Chen, Long}, year={2025}, month={Apr.}, pages={28706-28706} }
Open-World Multimodal Understanding and Generation with Efficiently Finetuned Foundation Models · AAAI 2025