← Search

Mengxue Zhang

3 accepted papers

2025

CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs

NeurIPS 2025poster

Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded ha…

Cited by 0SourceScholar
2023

Interpretable Math Word Problem Solution Generation via Step-by-step Planning

ACL 2023long

Solutions to math word problems (MWPs) with step-by-step explanations are valuable, especially in education, to help students better comprehend problem-solving strategies. Most existing approaches only focus on obtaining the final correct answer. A few recent approaches leverage intermediate solutio…

Cited by 12SourcePDFScholar
2020

Evaluating the Performance of Reinforcement Learning Algorithms

ICML 2020poster

Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results are often inconsistent and difficult to replicate. In this work, we argue that the inconsistency of performance stems from…