AAAI 2026technical0 citations

Investigating Data Pruning for Pretraining Biological Foundation Models at Scale

Yifan Wu, Jiyue Jiang, Xichen Ye, Yiqi Wang, Chang Zhou, Yitao Xu, Jiayang Chen, He Hu

Abstract

Biological foundation models (BioFMs), pretrained on large-scale biological sequences, have recently shown strong potential in providing meaningful representations for diverse downstream bioinformatics tasks. However, such models often rely on millions to billions of training sequences and billions of parameters, resulting in prohibitive computational costs and significant barriers to reproducibility and accessibility—particularly for academic labs. To address these challenges, we investigate the feasibility of data pruning for BioFM pretraining and propose a post-hoc influence-guided data pruning framework tailored to biological domains. Our approach first introduces a subset-based self-influence formulation that enables efficient estimation of sample importance at low computational cost. Built upon this, we propose two simple yet effective selection strategies: Top-k Influence (Top I) and Coverage-Centric Influence (CCI). Then, we empirically validate our method on two representative BioFMs: RNA-FM and ESM-C. For RNA, our framework consistently outperforms random selection baselines under an extreme pruning rate of over 99%, which displays our framework

BibTeX
@inproceedings{aaai2026_investigatingdat,
  title = {Investigating Data Pruning for Pretraining Biological Foundation Models at Scale},
  author = {Yifan Wu and Jiyue Jiang and Xichen Ye and Yiqi Wang and Chang Zhou and Yitao Xu and Jiayang Chen and He Hu and Weizhong Zhang and Cheng Jin and Jiao Yuan and Yu Li},
  booktitle = {AAAI 2026},
  year = {2026}
}