← Search

Haoke Zhang

4 accepted papers

2025

Generative Reward Modeling via Synthetic Criteria Preference Learning

ACL 2025long

Generative Reward Models (GenRMs) leverage synthesized Chains of Thought (CoT) to reduce the need for massive labeled data, but this approach introduces risks of overoptimization due to the inability to guarantee the correctness of the CoTs. Identifying and optimizing unexpected behaviors within the…

2025

Unlocking Recursive Thinking of LLMs: Alignment via Refinement

ACL 2025finding

The OpenAI o1-series models have demonstrated that leveraging long-form Chain of Thought (CoT) can substantially enhance performance. However, the recursive thinking capabilities of Large Language Models (LLMs) remain limited, particularly in the absence of expert-curated data for distillation. In t…

2025

𝒜3: Automatic Alignment Framework for Attributed Text Generation

ACL 2025long

Attributed text generation aims to enhance the reliability of content generated from large language models by providing citations for each claim, which thereby enables users to easily verify the correctness of the responses.However, the scarcity of high-quality training samples presents a significan…

2023

G-SPEED: General SParse Efficient Editing MoDel

EMNLP 2023long findings

Large Language Models (LLMs) have demonstrated incredible capabilities in understanding, generating, and manipulating languages. Through human-model interactions, LLMs can automatically understand human-issued instructions and output the expected contents, which can significantly increase working ef…

Cited by 0SourcecodeScholar