← Search

Jiarui Yu

3 accepted papers

2025

ReDit: Reward Dithering for Improved LLM Policy Optimization

NeurIPS 2025poster

DeepSeek-R1 has successfully enhanced Large Language Model (LLM) reasoning capabilities through its rule-based reward system. While it's a ''perfect'' reward system that effectively mitigates reward hacking, such reward functions are often discrete. Our experimental observations suggest that discret…

Cited by 0SourceScholar
2024

Event Detection from Social Media for Epidemic Prediction

NAACL 2024long

Social media is an easy-to-access platform providing timely updates about societal trends and events. Discussions regarding epidemic-related events such as infections, symptoms, and social interactions can be crucial for informing policymaking during epidemic outbreaks. In our work, we pioneer explo…

2024

SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness

EMNLP 2024main

Social media is often the first place where communities discuss the latest societal trends. Prior works have utilized this platform to extract epidemic-related information (e.g. infections, preventive measures) to provide early warnings for epidemic prediction. However, these works only focused on E…

Cited by 1SourcePDFScholar