Reading guide
This nine-chapter series develops reinforcement learning from first principles. Begin with the learning problem and Markov models, use returns and value functions to derive the Bellman equations, and turn those equations into planning algorithms. Policies as probability distributions then lead to policy gradients, PPO, GRPO, and DPO.
Definitions, proofs, and worked explanations develop the links between chapters. Previous/next navigation follows the same order as the path below. Later extensions may explore other preference objectives and decision processes in diffusion models. The conceptual starting point is Datawhale Easy RL, with independently developed derivations and additional material.