← Search

Deyuan Li

2 accepted papers

2024

M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models

EMNLP 2024finding

Existing evaluation benchmarks for foundation models in understanding scientific literature predominantly focus on single-document, text-only tasks. Such benchmarks often do not adequately represent the complexity of research workflows, which typically also involve interpreting non-textual data, suc…

2022

Bridging the Gap: Unifying the Training and Evaluation of Neural Network Binary Classifiers

NeurIPS 2022accept

While neural network binary classifiers are often evaluated on metrics such as Accuracy and $F_1$-Score, they are commonly trained with a cross-entropy objective. How can this training-evaluation gap be addressed? While specific techniques have been adopted to optimize certain confusion matrix based…

Cited by 1SourcePDFScholar