PPO, One Question at a Time: From Rewards to Policy Updates
The questions behind my PPO notes: why rewards cannot simply backpropagate through sampled tokens, what advantage and GAE estimate, why on-policy rollouts still need probability ratios, and what clipping actually changes. Includes numerical examples, runnable gradient checks, and a focused reading path.