← Search

Jenny Zhang

6 accepted papers

2026

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

ICLR 2026poster

Most of today's AI systems are constrained by human-designed, fixed architectures and cannot autonomously and continuously improve themselves. The scientific method, on the other hand, is a cumulative and open-ended system, where each innovation builds upon previous artifacts, enabling future discov…

Cited by 0SourcecodeScholar
2025

OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

ICLR 2025poster

Open-ended and AI-generating algorithms aim to continuously generate and solve increasingly complex tasks indefinitely, offering a promising path toward more general intelligence. To accomplish this grand vision, learning must occur within a vast array of potential tasks. Existing approaches to auto…

2024

Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback

ICRA 2024poster

We present a minimal phase oscillator model for learning quadrupedal locomotion. Each of the four oscillators is coupled only to itself and its corresponding leg through local feedback of the ground reaction force, which can be interpreted as an observer feedback gain. We interpret the oscillator it…

Cited by 3SourcecodeScholar
2024

OMNI: Open-endedness via Models of human Notions of Interestingness

ICLR 2024poster

Open-ended algorithms aim to learn new, interesting behaviors forever. That requires a vast environment search space, but there are thus infinitely many possible tasks. Even after filtering for tasks the current agent can learn (i.e., learning progress), countless learnable yet uninteresting tasks r…

2024

Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization

ICML 2024poster

Reinforcement Learning from Human Feedback (RLHF) has shown potential in qualitative tasks where easily defined performance measures are lacking. However, there are drawbacks when RLHF is commonly used to optimize for average human preferences, especially in generative tasks that demand diverse mode…

2024

Quality-Diversity through AI Feedback

ICLR 2024poster

In many text-generation problems, users may prefer not only a single response, but a diverse range of high-quality outputs from which to choose. Quality-diversity (QD) search algorithms aim at such outcomes, by continually improving and diversifying a population of candidates. However, the applicabi…

Cited by 50SourcePDFScholar