Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
Yi Zhang, Chun-Wun Cheng, Junyi He, Zhihai He, Carola-Bibiane Schönlieb, Yuyan Chen, Angelica I Aviles-Rivero
Abstract
We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
BibTeX
@article{Zhang_Cheng_He_He_Schönlieb_Chen_Aviles-Rivero_2025, title={Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/33118}, DOI={10.1609/aaai.v39i10.33118}, abstractNote={We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model’s expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP’s robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.}, number={10}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Zhang, Yi and Cheng, Chun-Wun and He, Junyi and He, Zhihai and Schönlieb, Carola-Bibiane and Chen, Yuyan and Aviles-Rivero, Angelica I}, year={2025}, month={Apr.}, pages={10302-10310} }