← Search

Sergey Ovchinnikov

2 accepted papers

2025

The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling

ICLR 2025poster

Biological language model performance depends heavily on pretraining data quality, diversity, and size. While metagenomic datasets feature enormous biological diversity, their utilization as pretraining data has been limited due to challenges in data accessibility, quality filtering and deduplicatio…

Cited by 49SourcePDFScholar
2021

Transformer protein language models are unsupervised structure learners

ICLR 2021poster

Unsupervised contact prediction is central to uncovering physical, structural, and functional constraints for protein structure determination and design. For decades, the predominant approach has been to infer evolutionary constraints from a set of related sequences. In the past year, protein langua…

Cited by 380SourcePDFScholar