← Search

Onur Celik

14 accepted papers

2026

PAWS: Preference Learning with Advantage-Weighted Segments

ICML 2026poster

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates…

Cited by 0SourceScholar
2026

Trust-Region Diffusion Policies for Massively Parallel On-Policy RL

ICML 2026poster

Reinforcement learning with massively parallel simulations has become an emerging trend; however, most existing approaches still rely on simple Gaussian policy parameterizations. Diffusion models provide a more expressive policy class and have shown strong performance on challenging control problems…

Cited by 0SourceScholar
2025

DIME: Diffusion-Based Maximum Entropy Reinforcement Learning

ICML 2025poster

Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized using Gaussian distributions, which significantly limits their representational capacity. Diffusion-based policies offer a…

Cited by 0SourcePDFScholar
2025

Scaffolding Dexterous Manipulation with Vision-Language Models

NeurIPS 2025poster

Dexterous robotic hands are essential for performing complex manipulation tasks, yet remain difficult to train due to the challenges of demonstration collection and high-dimensional control. While reinforcement learning (RL) can alleviate the data bottleneck by generating experience in simulation, i…

Cited by 0SourceScholar
2024

Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of Experts

ICML 2024poster

Reinforcement learning (RL) is a powerful approach for acquiring a good-performing policy. However, learning diverse skills is challenging in RL due to the commonly used Gaussian policy parameterization. We propose Diverse Skill Learning (Di-SkilL), an RL method for learning diverse skills using Mix…

Cited by 7SourcePDFScholar
2024

MaIL: Improving Imitation Learning with Selective State Space Models

CoRL 2024poster

This work introduces Mamba Imitation Learning (MaIL), a novel imitation learning (IL) architecture that offers a computationally efficient alternative to state-of-the-art (SoTA) Transformer policies. Transformer-based policies have achieved remarkable results due to their ability in handling human-r…

Cited by 7SourceScholar
2024

MuTT: A Multimodal Trajectory Transformer for Robot Skills

IROS 2024poster

High-level robot skills represent an increasingly popular paradigm in robot programming. However, configuring the skills’ parameters for a specific task remains a manual and time-consuming endeavor. Existing approaches for learning or optimizing these parameters often require numerous real-world exe…

Cited by 2SourceScholar
2024

Variational Distillation of Diffusion Policies into Mixture of Experts

NeurIPS 2024poster

This work introduces Variational Diffusion Distillation (VDD), a novel method that distills denoising diffusion policies into Mixtures of Experts (MoE) through variational inference. Diffusion Models are the current state-of-the-art in generative modeling due to their exceptional ability to accurate…

2023

Curriculum-Based Imitation of Versatile Skills

ICRA 2023poster

Learning skills by imitation is a promising concept for the intuitive teaching of robots. A common way to learn such skills is to learn a parametric model by maximizing the likelihood given the demonstrations. Yet, human demonstrations are often multi-modal, i.e., the same task is solved in multiple…

Cited by 4SourcecodeScholar
2023

Information Maximizing Curriculum: A Curriculum-Based Approach for Learning Versatile Skills

NeurIPS 2023poster

Imitation learning uses data for training policies to solve complex tasks. However, when the training data is collected from human demonstrators, it often leads to multimodal distributions because of the variability in human actions. Most imitation learning methods rely on a maximum likelihood (ML)…

Cited by 15SourcePDFScholar
2022

Deep Black-Box Reinforcement Learning with Movement Primitives

CoRL 2022poster

Episode-based reinforcement learning (ERL) algorithms treat reinforcement learning (RL) as a black-box optimization problem where we learn to select a parameter vector of a controller, often represented as a movement primitive, for a given task descriptor called a context. ERL offers several distinc…

Cited by 29SourcecodeScholar
2021

Specializing Versatile Skill Libraries using Local Mixture of Experts

CoRL 2021poster

A long-cherished vision in robotics is to equip robots with skills that match the versatility and precision of humans. For example, when playing table tennis, a robot should be capable of returning the ball in various ways while precisely placing it at the desired location. A common approach to mod…

Cited by 41SourcecodeScholar
2019

Chance-Constrained Trajectory Optimization for Non-linear Systems with Unknown Stochastic Dynamics

IROS 2019poster

Iterative trajectory optimization techniques for non-linear dynamical systems are among the most powerful and sample-efficient methods of model-based reinforcement learning and approximate optimal control. By leveraging time-variant local linear-quadratic approximations of system dynamics and reward…

Cited by 10SourceScholar