← Search

Takumi Tanabe

4 accepted papers

2026

Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment

AAAI 2026technical

Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO alignment have been studied empirically, their theoretical foundations remain unclear. We investigate the minimum-cost po

Cited by 0SourcePDFScholar
2025

A Provable Approach for End-to-End Safe Reinforcement Learning

NeurIPS 2025poster

A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing safe RL paradigms inherently struggle to achieve this objective. We propose a method, called Provably Lifetime Safe RL (PL…

Cited by 0SourceScholar
2024

Stepwise Alignment for Constrained Language Model Policy Optimization

NeurIPS 2024poster

Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment as an optimization problem of the language model policy to maximize reward under a safety constraint, and then proposes…

2022

Max-Min Off-Policy Actor-Critic Method Focusing on Worst-Case Robustness to Model Misspecification

NeurIPS 2022accept

In the field of reinforcement learning, because of the high cost and risk of policy training in the real world, policies are trained in a simulation environment and transferred to the corresponding real-world environment. However, the simulation environment does not perfectly mimic the real-world en…