← Search

Muhammed Razzak

3 accepted papers

2026

MADE: Benchmark Environments for Closed-Loop Materials Discovery

ICML 2026poster

Existing benchmarks for computational materials discovery primarily evaluate static predictive tasks or isolated computational sub-tasks. While valuable, these evaluations neglect the inherently iterative and adaptive nature of scientific discovery. We introduce MAterials Discovery Environments (MAD…

Cited by 0SourceScholar
2025

Scaling Up Active Testing to Large Language Models

NeurIPS 2025poster

Active testing enables label-efficient evaluation of predictive models through careful data acquisition, but it can pose a significant computational cost. We identify cost-saving measures that enable active testing to be scaled up to large language models (LLMs). In particular we show that the surro…

Cited by 0SourceScholar
2025

Simple Factuality Probes Detect Hallucinations in Long-Form Natural Language Generation

EMNLP 2025

Large language models (LLMs) often mislead users with confident hallucinations. Current approaches to detect hallucination require many samples from the LLM generator, which is computationally infeasible as frontier model sizes and generation lengths continue to grow. We present a remarkably simple