← Search

Shi Fu

8 accepted papers

2026

The State of Reinforcement Finetuning for Transformer-based Generative Agents

ICLR 2026poster

Reinforcement finetuning (RFT) has garnered significant attention in recent years, particularly for enhancing large reasoning models such as OpenAI o1 and Deepseek R1. The appeal of RFT largely stems from its ability to refine model knowledge, better align outputs with user intent, and address chall…

Cited by 0SourceScholar
2026

Towards a Theoretical Understanding of In-context Learning: Stability and Non-I.I.D Generalisation

ICLR 2026poster

In-context learning (ICL) has demonstrated significant performance improvements in transformer-based large models. This study identifies two key factors influencing ICL generalisation under complex non-i.i.d. scenario: algorithmic stability and distributional discrepancy. First, we establish a stabi…

Cited by 0SourceScholar
2025

A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops

ICLR 2025poster

High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical r…

Cited by 0SourcePDFScholar
2025

Self-Verification Provably Prevents Model Collapse in Recursive Synthetic Training

NeurIPS 2025poster

Large generative models are increasingly trained on synthetic data from earlier generations, raising concerns about *model collapse*, a progressive performance decline consistently observed in empirical studies. However, theoretical understanding of recursive training dynamics and their failure mode…

Cited by 0SourceScholar
2024

Adaptive Time-Stepping Schedules for Diffusion Models

UAI 2024poster

This paper studies how to tune the stepping schedule in diffusion models, which is mostly fixed in current practice, lacking theoretical foundations and assurance of optimal performance at the chosen discretization points. In this paper, we advocate the use of adaptive time-stepping schedules and de…

2024

Towards Theoretical Understandings of Self-Consuming Generative Models

ICML 2024poster

This paper tackles the emerging challenge of training generative models within a self-consuming loop, wherein successive generations of models are recursively trained on mixtures of real and synthetic data from previous generations. We construct a theoretical framework to rigorously evaluate how thi…

Cited by 8SourcePDFScholar
2023

Sharper Bounds for Uniformly Stable Algorithms with Stationary Mixing Process

ICLR 2023poster

Generalization analysis of learning algorithms often builds on a critical assumption that training examples are independently and identically distributed, which is often violated in practical problems such as time series prediction. In this paper, we use algorithmic stability to study the generaliza…

Cited by 5SourcePDFScholar