← Search

Zhixuan Lin

6 accepted papers

2026

MSP: Probabilistically Consistent Multi-Scale Action Generation

ICML 2026spotlight

In robotic imitation learning, accurately modeling the multimodality and temporal correlations of long-horizon action sequences remains challenging. Long-horizon tasks require preserving global task intent while executing precise low-level control; otherwise, local errors can accumulate and lead to …

Cited by 0SourceScholar
2025

Forgetting Transformer: Softmax Attention with a Forget Gate

ICLR 2025poster

An essential component of modern recurrent sequence models is the forget gate. While Transformers do not have an explicit recurrent form, we show that a forget gate can be naturally incorporated into Transformers by down-weighting the unnormalized attention scores in a data-dependent way. We name th…

2024

The Curse of Diversity in Ensemble-Based Exploration

ICLR 2024poster

We uncover a surprising phenomenon in deep reinforcement learning: training a diverse ensemble of data-sharing agents -- a well-established exploration strategy -- can significantly impair the performance of the individual ensemble members when compared to standard single-agent training. Through car…

Cited by 3SourcePDFScholar
2020

Improving Generative Imagination in Object-Centric World Models

ICML 2020poster

The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these ingredients can be integrated into a unified framework. Second,…

Cited by 88SourcePDFScholar
2020

SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition

ICLR 2020poster

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either based on spatial-attention or scene-mixture approaches and li…

Cited by 269SourceScholar
2019

GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs

NeurIPS 2019poster

Finding local correspondences between images with different viewpoints requires local descriptors that are robust against geometric transformations. An approach for transformation invariance is to integrate out the transformations by pooling the features extracted from transformed versions of an ima…

Cited by 109SourcePDFScholar