2025
GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models
EMNLP 2025
In natural language processing (NLP) tasks, pure reinforcement learning fine-tuning methods often suffer from inefficient exploration and slow convergence; while supervised fine-tuning (SFT) methods, although efficient in training, have limited performance ceiling and less solid theoretical foundati