← Search

Chunshan Li

3 accepted papers

2025

Maximizing Intermediate Checkpoint Value in LLM Pretraining with Bayesian Optimization

ICML 2025poster

The rapid proliferation of large language models (LLMs), such as GPT-4 and Gemini, underscores the intense demand for resources during their training processes, posing significant challenges due to substantial computational and environmental costs. In this paper, we introduce a novel checkpoint merg…

Cited by 0SourcePDFScholar
2025

VPO: Reasoning Preferences Optimization Based on $\mathcal{V}$-Usable Information

NeurIPS 2025spotlight

Direct Preference Optimization (DPO) is a widely used preference optimization algorithm in large language model (LLM) alignment, which reparameterizes the reward function in reinforcement learning with human feedback (RLHF) without requiring a separate reward model. However, during the DPO training…

Cited by 0SourceScholar
2024

Analyzing Chain-of-thought Prompting in Black-Box Large Language Models via Estimated V-information

COLING 2024main

Chain-of-Thought (CoT) prompting combined with large language models (LLM) has shown great potential in improving performance on challenging reasoning tasks. While understanding why CoT prompting is effective is crucial for the application and improvement of CoT prompting, few studies have addressed…

Cited by 1SourcePDFScholar