← Search

Bochen Lyu

8 accepted papers

2026

MoDr: Mixture-of-Depth-Recurrent Transformers for Test-Time Reasoning

ICLR 2026poster

Large Language Models have demonstrated superior reasoning capabilities by generating step-by-step reasoning in natural language before deriving the final answer. Recently, Geiping et al. introduced 3.5B-Huginn as an alternative to this paradigm, a depth-recurrent Transformer that increases computat…

Cited by 0SourceScholar
2026

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently

ICML 2026poster

Transformers can acquire Chain-of-Thought (CoT) capabilities to solve complex reasoning tasks through fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we examine them specifically for learning k-sparse Boolean functions w…

Cited by 0SourceScholar
2025

Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear Networks

AAAI 2025technical

This paper targets on the regularization effect of momentum-based methods in regression settings and analyzes the popular diagonal linear networks to precisely characterize the implicit bias of continuous versions of heavy-ball (HB) and Nesterov's method of accelerated gradients (NAG). We show that,…

Cited by 0SourcePDFScholar
2025

Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis

NeurIPS 2025poster

This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization error. Investigating continuous differential equations has been a promising approach for studying the discrete optimiza…

Cited by 0SourceScholar