Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
Reinforcement Learning (RL) traditionally relies on scalar reward signals, limiting its ability to leverage the rich semantic knowledge often available in real-world tasks. In contrast, humans learn efficiently by combining numerical feedback with language, prior knowledge, and common sense. We intr…