← Search

Anshul Kundaje

7 accepted papers

2024

DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

NeurIPS 2024poster

Recent advances in self-supervised models for natural language, vision, and protein sequences have inspired the development of large genomic DNA language models (DNALMs). These models aim to learn generalizable representations of diverse DNA elements, potentially enabling various genomic prediction,…

2023

Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design

NeurIPS 2023poster

The efficient exploration of chemical space to design molecules with intended properties enables the accelerated discovery of drugs, materials, and catalysts, and is one of the most important outstanding challenges in chemistry. Encouraged by the recent surge in computer power and artificial intelli…

2021

WILDS: A Benchmark of in-the-Wild Distribution Shifts

ICML 2021oral

Distribution shifts—where the training distribution differs from the test distribution—can substantially degrade the accuracy of machine learning (ML) systems deployed in the wild. Despite their ubiquity in the real-world deployments, these distribution shifts are under-represented in the datasets w…

2020

Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics

NeurIPS 2020poster

Deep learning models can accurately map genomic DNA sequences to associated functional molecular readouts such as protein-DNA binding data. Base-resolution importance (i.e. "attribution") scores inferred from these models can highlight predictive sequence motifs and syntax. Unfortunately, these mode…

Cited by 39SourcePDFScholar
2020

Maximum Likelihood with Bias-Corrected Calibration is Hard-To-Beat at Label Shift Adaptation

ICML 2020poster

Label shift refers to the phenomenon where the prior class probability p(y) changes between the training and test distributions, while the conditional probability p(x|y) stays fixed. Label shift arises in settings like medical diagnosis, where a classifier trained to predict disease given symptoms m…

2017

Learning Important Features Through Propagating Activation Differences

ICML 2017poster

The purported “black box” nature of neural networks is a barrier to adoption in applications where interpretability is essential. Here we present DeepLIFT (Deep Learning Important FeaTures), a method for decomposing the output prediction of a neural network on a specific input by backpropagating the…

Cited by 5417SourcePDFScholar
2016

Unsupervised Learning from Noisy Networks with Applications to Hi-C Data

NeurIPS 2016poster

Complex networks play an important role in a plethora of disciplines in natural sciences. Cleaning up noisy observed networks, poses an important challenge in network analysis Existing methods utilize labeled data to alleviate the noise effect in the network. However, labeled data is usually expens…

Cited by 7SourcePDFScholar