← Search

Leon Bergen

7 accepted papers

2025

Adapting While Learning: Grounding LLMs for Scientific Problems with Tool Usage Adaptation

ICML 2025poster

Large Language Models (LLMs) demonstrate promising capabilities in solving scientific problems but often suffer from the issue of hallucination. While integrating LLMs with tools can mitigate this issue, models fine-tuned on tool usage become overreliant on them and incur unnecessary costs. Insp…

2025

ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

ICLR 2025poster

The use of Large Language Models (LLMs) in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation framework capable of assessing the quality and scientific validity of model outputs. To address this issue, we develop *Clim…

2025

Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark

EMNLP 2025

Systems that answer questions by reviewing the scientific literature are becoming increasingly feasible. To draw reliable conclusions, these systems should take into account the quality of available evidence from different studies, placing more weight on studies that use a valid methodology. We pres

2024

IR2: Information Regularization for Information Retrieval

COLING 2024main

Effective information retrieval (IR) in settings with limited training data, particularly for complex queries, remains a challenging task. This paper introduces IR2, Information Regularization for Information Retrieval, a technique for reducing overfitting during synthetic data generation. This appr…

2023

Scientific Document Retrieval using Multi-level Aspect-based Queries

NeurIPS 2023poster

In scientific research, the ability to effectively retrieve relevant documents based on complex, multifaceted queries is critical. Existing evaluation datasets for this task are limited, primarily due to the high costs and effort required to annotate resources that effectively represent complex quer…