← Search

Bilgehan Sel

9 accepted papers

2026

Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning

ICML 2026spotlight

Fine-tuning APIs offered by major AI providers create new attack surfaces where adversaries can bypass safety measures through targeted fine-tuning. We introduce **Trojan-Speak**, an adversarial fine-tuning method that bypasses Anthropic's Constitutional Classifiers. Our approach uses curriculum lea…

Cited by 0SourceScholar
2025

Reinforcement Learning with Backtracking Feedback

NeurIPS 2025poster

Addressing the critical need for robust safety in Large Language Models (LLMs), particularly against adversarial attacks and in-distribution errors, we introduce Reinforcement Learning with Backtracking Feedback (RLBF). This framework advances upon prior methods, such as BSAFE, by primarily leveragi…

Cited by 0SourceScholar
2024

Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models

ICML 2024poster

Current literature, aiming to surpass the "Chain-of-Thought" approach, often resorts to external modi operandi involving halting, modifying, and then resuming the generation process to boost Large Language Models' (LLMs) reasoning capacities. Due to their *myopic perspective*, they escalate the numb…

2024

Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation

AAAI 2024technical

Ensuring the safety of Reinforcement Learning (RL) is crucial for its deployment in real-world applications. Nevertheless, managing the trade-off between reward and safety during exploration presents a significant challenge. Improving reward performance through policy adjustments may adversely affec…

2024

Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs

ACL 2024long

Large Language Models (LLMs) have shown remarkable capabilities in tasks such as summarization, arithmetic reasoning, and question answering. However, they encounter significant challenges in the domain of moral reasoning and ethical decision-making, especially in complex scenarios with multiple sta…

Cited by 6SourcePDFScholar
2023

A CMDP-within-online framework for Meta-Safe Reinforcement Learning

ICLR 2023top-25%

Meta-reinforcement learning has widely been used as a learning-to-learn framework to solve unseen tasks with limited experience. However, the aspect of constraint violations has not been adequately addressed in the existing works, making their application restricted in real-world settings. In this p…

Cited by 22SourcePDFScholar
2023

On Solution Functions of Optimization: Universal Approximation and Covering Number Bounds

AAAI 2023technical

We study the expressibility and learnability of solution functions of convex optimization and their multi-layer architectural extension. The main results are: (1) the class of solution functions of linear programming (LP) and quadratic programming (QP) is a universal approximant for the smooth model…

Cited by 8SourcePDFScholar