← Search

Zhenqing Ling

5 accepted papers

2026

BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning

ICLR 2026poster

Reinforcement finetuning (RFT) is a key technique for aligning Large Language Models (LLMs) with human preferences and enhancing reasoning, yet its effectiveness is highly sensitive to which tasks are explored during training. Uniform task sampling is inefficient, wasting computation on tasks that a…

Cited by 0SourceScholar
2026

GeoAlign: Geometric Rollout Curation for Robust LLM Reinforcement Learning

ICML 2026poster

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollouts induces represen…

Cited by 0SourceScholar
2025

Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data

NeurIPS 2025poster

Fine-tuning large language models (LLMs) using diverse datasets is crucial for enhancing their overall performance across various domains. In practical scenarios, existing methods based on modeling the mixture proportions of data composition often struggle with data whose domain labels are missing,…

Cited by 0SourcecodeScholar
2025

Enhancing Factual Consistency in Text Summarization via Counterfactual Debiasing

COLING 2025main

Despite significant progress in abstractive text summarization aimed at generating fluent and informative outputs, how to ensure the factual consistency of generated summaries remains a crucial and challenging issue. In this study, drawing inspiration from advancements in causal inference, we constr…

Cited by 1SourcePDFScholar
2025

MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?

NeurIPS 2025poster

Large foundation models face challenges in acquiring transferable, structured thinking abilities, especially when supervised with rigid templates or crowd-annotated instruction datasets. Unlike prior approaches, we focus on a thinking-centric data synthesis paradigm that enables models to evolve thr…

Cited by 0SourceScholar