← Search

Yixiao Zeng

2 accepted papers

2026

HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithimic Coding

ICLR 2026poster

Verifiers provide important reward signals for reinforcement learning of large language models (LLMs). However, it is challenging to develop or create reliable verifiers, especially for code generation tasks. A well-disguised wrong solution program may only be detected by carefully human-written edg…

Cited by 0SourcecodeScholar
2026

Prompt-MII: Meta-Learning Instruction Induction for LLMs

ICLR 2026poster

A popular method to adapt large language models (LLMs) to new tasks is in-context learning (ICL), which is effective but incurs high inference costs as context length grows. In this paper we propose a method to perform instruction induction, where we take training examples and reduce them to a compa…

Cited by 0SourcecodeScholar