← Search

Liangze Jiang

8 accepted papers

2026

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

ICLR 2026poster

Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. **Method.** We present a method to optimize a transformer ar…

Cited by 0SourceScholar
2026

Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers

CVPR 2026

Transformers are remarkably versatile, suggesting the existence of generic inductive biases beneficial across modalities. In this work, we explore a new way to instil such biases in vision transformers (ViTs) through pretraining on procedurally generated data devoid of visual or semantic content. We

Cited by 0SourceScholar
2026

Meta-RL Induces Exploration in Language Agents

ICLR 2026poster

Reinforcement learning (RL) has enabled the training of Large Language Model (LLM) agents to interact with the environment and to solve multi-turn longhorizon tasks. However, the RL-trained agents often struggle in tasks that require active exploration and fail to efficiently adapt from trial-and-er…

Cited by 0SourcecodeScholar
2026

Procedural Pretraining: Warming Up Language Models with Abstract Data

ICML 2026oral

Pretraining directly on web-scale corpora is the de facto paradigm for building language models. We study an alternative setting where the model is initially exposed to abstract structured data, as a means to ease the subsequent acquisition of rich semantic knowledge, much like humans learn simple l…

Cited by 0SourceScholar
2025

Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild

CVPR 2025poster

Common choices of architecture give neural networks a preference for fitting data with simple functions. This simplicity bias is known as key to their success. This paper explores the limits of this assumption. Building on recent work that showed that activation functions are the origin of the simpl…

2024

Unraveling the Key Components of OOD Generalization via Diversification

ICLR 2024poster

Supervised learning datasets may contain multiple cues that explain the training set equally well, i.e., learning any of them would lead to the correct predictions on the training data. However, many of them can be spurious, i.e., lose their predictive power under a distribution shift and consequent…

Cited by 2SourcePDFScholar