← Search

Fangzhi Zhu

2 accepted papers

2026

Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations

ICLR 2026poster

Large Language Models (LLMs) have achieved impressive performance on complex tasks by generating human-like, step-by-step rationales, referred to as \textit{reasoning trajectory}, before arriving at final answers. However, the length of these reasoning trajectories often far exceeds that of the fina…

Cited by 0SourcecodeScholar
2025

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples

COLING 2025main

Aligning Large Language Models (LLMs) with human feedback is crucial for their development. Existing preference optimization methods such as DPO and KTO, while improved based on Reinforcement Learning from Human Feedback (RLHF), are inherently derived from PPO, requiring a reference model that adds…

Cited by 0SourcePDFScholar