← Search

Michael Krumdick

7 accepted papers

2025

BLEUBERI: BLEU is a surprisingly effective reward for instruction following

NeurIPS 2025poster

Reward models are central to aligning LLMs with human preferences, but they are costly to train, requiring large-scale human-labeled preference data and powerful pretrained LLM backbones. Meanwhile, the increasing availability of high-quality synthetic instruction-following datasets raises the quest…

Cited by 0SourcecodeScholar
2025

Complexity Scaling Laws for Neural Models using Combinatorial Optimization

NeurIPS 2025poster

Recent work on neural scaling laws demonstrates that model performance scales predictably with compute budget, model size, and dataset size. In this work, we develop scaling laws based on problem complexity. We analyze two fundamental complexity measures: solution space size and representation space…

Cited by 0SourceScholar
2025

Language Model Probabilities are Not Calibrated in Numeric Contexts

ACL 2025long

Some statements have one well-defined continuation (e.g., “the Eiffel Tower is in [Paris]"), whereas others have a natural distribution over multiple options (e.g., “the weighted coin flip was [Heads/Tails].") We argue that language model (LM) outputs should capture these natural distributions. Our…

Cited by 0SourcePDFScholar
2024

An Analysis of Multilingual FActScore

EMNLP 2024main

FActScore has gained popularity as a metric to estimate the factuality of long-form texts generated by Large Language Models (LLMs) in English. However, there has not been any work in studying the behavior of FActScore in other languages. This paper studies the limitations of each component in the f…

Cited by 1SourcePDFScholar
2024

BizBench: A Quantitative Reasoning Benchmark for Business and Finance

ACL 2024long

Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. Together, these requirements make this domain difficult for large language models (LLMs). We introduce BizBench, a benchmark for evaluating models’ ability to reason about realis…

Cited by 13SourcePDFScholar
2024

DocFinQA: A Long-Context Financial Reasoning Dataset

ACL 2024short

For large language models (LLMs) to be effective in the financial domain – where each decision can have a significant impact – it is necessary to investigate realistic tasks and data. Financial professionals often interact with documents spanning hundreds of pages, but most financial research datase…

Cited by 17SourcePDFScholar
2020

APRICOT: A Dataset of Physical Adversarial Attacks on Object Detection

ECCV 2020poster

Physical adversarial attacks threaten to fool object detection systems, but reproducible research on the real-world effectiveness of physical patches and how to defend against them requires a publicly available benchmark dataset. We present APRICOT, a collection of over 1,000 annotated photographs o…

Cited by 59SourcePDFScholar