← Search

Mateusz Ostaszewski

10 accepted papers

2026

Huxley-G\"odel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

ICLR 2026oral

Recent studies operationalize self-improvement through coding agents that edit their own codebases, grow a tree of self-modifications through expansion strategies that favor higher software engineering benchmark performance, considering that this implies more promising subsequent self-modifications…

Cited by 0SourcecodeScholar
2025

Learning Continually by Spectral Regularization

ICLR 2025poster

Loss of plasticity is a phenomenon where neural networks can become more difficult to train over the course of learning. Continual learning algorithms seek to mitigate this effect by sustaining good performance while maintaining network trainability. We develop a new technique for improving continua…

Cited by 4SourcePDFScholar
2025

PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors

NeurIPS 2025poster

Evaluating the scientific discovery capabilities of large language model based agents, particularly how they cope with varying environmental complexity and utilize prior knowledge, requires specialized benchmarks currently lacking in the landscape. To address this gap, we introduce PhysGym, a novel…

Cited by 0SourceScholar
2024

Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous control

NeurIPS 2024spotlight

Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial improvements. We conduct a thorough investigation into the interplay of scaling model capacity and domain-specific RL en…

2024

Curriculum reinforcement learning for quantum architecture search under hardware errors

ICLR 2024poster

The key challenge in the noisy intermediate-scale quantum era is finding useful circuits compatible with current device limitations. Variational quantum algorithms (VQAs) offer a potential solution by fixing the circuit architecture and optimizing individual gate parameters in an external loop. Howe…

Cited by 23SourcePDFScholar
2024

Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem

ICML 2024spotlight

Fine-tuning is a widespread technique that allows practitioners to transfer pre-trained capabilities, as recently showcased by the successful applications of foundation models. However, fine-tuning reinforcement learning (RL) models remains a challenge. This work conceptualizes one specific cause of…

2024

Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

ICML 2024poster

Recent advancements in off-policy Reinforcement Learning (RL) have significantly improved sample efficiency, primarily due to the incorporation of various forms of regularization that enable more gradient update steps than traditional agents. However, many of these techniques have been tested in lim…

Cited by 22SourcePDFScholar
2023

The Tunnel Effect: Building Data Representations in Deep Neural Networks

NeurIPS 2023poster

Deep neural networks are widely known for their remarkable effectiveness across various tasks, with the consensus that deeper networks implicitly learn more complex data representations. This paper shows that sufficiently deep networks trained for supervised image classification split into two disti…

Cited by 18SourcePDFScholar
2021

Reinforcement learning for optimization of variational quantum circuit architectures

NeurIPS 2021poster

The study of Variational Quantum Eigensolvers (VQEs) has been in the spotlight in recent times as they may lead to real-world applications of near-term quantum devices. However, their performance depends on the structure of the used variational ansatz, which requires balancing the depth and expressi…

Cited by 169SourcePDFScholar