← Search

Weijing Huang

2 accepted papers

2025

Training Medical QA Models Based on Mixed Rewards from Multiple-Choice and Open-Ended Questions

EMNLP 2025

Reinforcement learning (RL) for large language models (LLMs) typically requires clear reward signals, which are often unavailable for open-ended (OE) questions where answer evaluation is ambiguous without scalable expert labeling. We investigate whether LLMs benefit from training on mixed data with

Cited by 0SourcePDFScholar
2020

Generating Reasonable Legal Text through the Combination of Language Modeling and Question Answering

IJCAI 2020poster

Due to the improvement of Language Modeling, the emerging NLP assistant tools aiming for text generation greatly reduce the human workload on writing documents. However, the generation of legal text faces greater challenges than ordinary texts because of its high requirement for keeping logic reason…