Reading guide

These notes explain the problems and design choices behind learning systems. The current series develops reinforcement learning from interaction and Markov models through policy gradients, PPO, GRPO, and DPO. Follow its reading path for the detailed progression. Representation learning, sequence models, and generative models are future directions for this collection.