← Search

Alper Kamil Bozkurt

5 accepted papers

2026

Accelerated Learning with Linear Temporal Logic using Differentiable Simulation

ICLR 2026poster

Ensuring that reinforcement learning (RL) controllers satisfy safety and reliability constraints in real-world settings remains challenging: state-avoidance and constrained Markov decision processes often fail to capture trajectory-level requirements or induce overly conservative behavior. Formal sp…

Cited by 0SourceScholar
2024

Steering Decision Transformers via Temporal Difference Learning

IROS 2024poster

Decision Transformers (DTs) have been highly effective for offline reinforcement learning (RL) tasks, successfully modeling the sequences of actions in a given set of demonstrations. However, DTs may perform poorly in stochastic environments, which are prevalent in robotics scenarios. In this paper,…

Cited by 0SourceScholar
2021

Model-Free Reinforcement Learning for Stochastic Games with Linear Temporal Logic Objectives

ICRA 2021poster

We study synthesis of control strategies from linear temporal logic (LTL) objectives in unknown environments. We model this problem as a turn-based zero-sum stochastic game between the controller and the environment, where the transition probabilities and the model topology are fully unknown. The wi…

Cited by 20SourceScholar
2021

Secure Planning Against Stealthy Attacks via Model-Free Reinforcement Learning

ICRA 2021poster

We consider the problem of security-aware planning in an unknown stochastic environment, in the presence of attacks on control signals (i.e., actuators) of the robot. We model the attacker as an agent who has the full knowledge of the controller as well as the employed intrusion-detection system and…

Cited by 21SourceScholar
2020

Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning

ICRA 2020poster

We present a reinforcement learning (RL) frame-work to synthesize a control policy from a given linear temporal logic (LTL) specification in an unknown stochastic environment that can be modeled as a Markov Decision Process (MDP). Specifically, we learn a policy that maximizes the probability of sat…

Cited by 168SourceScholar