← Search

Rohan Bindu

1 accepted papers

2024

Removing RLHF Protections in GPT-4 via Fine-Tuning

NAACL 2024short

As large language models (LLMs) have increased in their capabilities, so doestheir potential for dual use. To reduce harmful outputs, produces and vendors ofLLMs have used reinforcement learning with human feedback (RLHF). In tandem,LLM vendors have been increasingly enabling fine-tuning of their mo…

Cited by 104SourcePDFScholar