← Search

Shayan Shabihi

2 accepted papers

2026

PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach

ICLR 2026poster

Recent advances in Large Language Models (LLMs) have sparked concerns over their potential to acquire and misuse dangerous capabilities, posing frontier risks to society. Current safety evaluations primarily test for what a model *can* do---its capabilities---without assessing what it *would* do if…

Cited by 0SourcecodeScholar
2026

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

ICML 2026poster

Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scientific knowledge and reasoning, their ability to predict experimental outcomes---…

Cited by 0SourceScholar