← Search

Praveenkumar Kanithi

4 accepted papers

2026

BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation Models

ICML 2026poster

Existing genomic foundation models (GFMs) typically treat DNA as raw nucleotide sequences, often overlooking the regulatory context required to interpret genetic variation accurately. We introduce BioToken, a tokenization framework that directly encodes variants and biological annotations into genom…

Cited by 0SourceScholar
2026

Genomic Foundationless Models: Pretraining Does Not Promise Performance

ICLR 2026poster

The success of Large Language Models has inspired the development of Genomic Foundation Models (GFMs) through similar pretraining techniques. However, the relationship between pretraining performance and effectiveness in downstream ge- nomic tasks remains unclear. Additionally, the high computationa…

Cited by 0SourcecodeScholar
2025

Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency

EMNLP 2025

Large language models offer transformative potential for healthcare, yet their responsible and equitable development depends critically on a deeper understanding of how training data characteristics influence model behavior, including the potential for bias. Current practices in dataset curation and

2024

Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs.

EMNLP 2024finding

Large Language Models (LLMs) have demonstrated significant potential in revolutionizing clinical applications. In this study, we investigate the efficacy of four techniques in adapting LLMs for clinical use-cases: continuous pretraining, instruct fine-tuning, NEFTune, and prompt engineering. We empl…

Cited by 0SourcePDFScholar