← Search

Keita Saito

1 accepted papers

2026

Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment

AAAI 2026technical

Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO alignment have been studied empirically, their theoretical foundations remain unclear. We investigate the minimum-cost po

Cited by 0SourcePDFScholar