← Search

Yanfang Zhang

2 accepted papers

2026

Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning

AAAI 2026technical

Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capability of Large Language Models (LLMs). Current RLVR approaches typically conduct training across all generated tokens, but neglect to explore which tokens (e.g., prefix tokens) actually contribute to reas

Cited by 0SourcePDFScholar
2025

Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning

COLING 2025main

Recently, increasing attention has been focused on improving the ability of Large Language Models (LLMs) to perform complex reasoning. Advanced methods, such as Chain-of-Thought (CoT) and its variants, are found to enhance their reasoning skills by designing suitable prompts or breaking down complex…

Cited by 2SourcePDFScholar