← Search

Sheng Jia

3 accepted papers

2026

Training Large Language Models To Reason In Parallel With Global Forking Tokens

ICLR 2026poster

Although LLMs have demonstrated improved performance by scaling parallel test-time compute, doing so relies on generating reasoning paths that are both diverse and accurate. For challenging problems, the forking tokens that trigger diverse yet correct reasoning modes are typically deep in the sampli…

Cited by 0SourcecodeScholar
2021

Efficient Statistical Tests: A Neural Tangent Kernel Approach

ICML 2021spotlight

For machine learning models to make reliable predictions in deployment, one needs to ensure the previously unknown test samples need to be sufficiently similar to the training data. The commonly used shift-invariant kernels do not have the compositionality and fail to capture invariances in high-dim…