From Dice to Agents: Reinforcement Learning in One Set of Symbols
Most introductions to reinforcement learning for language models start in the middle, with PPO or GRPO already on the table. This post starts from a dice game instead and adds...
Read post