Skip to the content

What This Unit Is About

Reinforcement learning replaces the dataset with an interaction loop. An agent in state \(s\) takes action \(a\), receives reward \(r\), and lands in state \(s'\). Nobody says which action was correct — only how much reward followed — and the consequences may arrive many steps later. The whole subject is built on one recursive identity, the Bellman equation, and on whether you have a model of the environment or must learn without one.

Learning Outcomes

By the end of this unit you should be able to:

  1. Write down the Bellman optimality equation and explain why it makes value iteration converge.
  2. Explain what makes Q-learning off-policy, and what the \(\varepsilon\)-greedy policy is trading off.
  3. Derive the REINFORCE update from the policy gradient theorem, and say why a baseline reduces its variance.
  4. Explain how Dyna-Q buys sample efficiency with planning — and what a wrong model costs.
  5. Choose between model-free and model-based methods from the cost of collecting real experience.

Topics in This Unit

4.1

🎲 Model-Free Methods

Learn directly from experience, without ever building a model of how the environment behaves.

2 algorithms • Q-Learning, Policy Gradient (REINFORCE)
4.2

🗺️ Model-Based Methods

Learn or be given the environment’s dynamics, then plan against them.

2 algorithms • Dyna-Q, Value Iteration (Dynamic Programming)

Every Algorithm in Reinforcement Learning

No.AlgorithmCharacterTopic
4.1Q-LearningOff-Policy • Temporal Difference4.1 Model-Free Methods
4.2Policy Gradient (REINFORCE)On-Policy • Continuous Actions4.1 Model-Free Methods
4.3Dyna-QModel-Based • Planning4.2 Model-Based Methods
4.4Value Iteration (Dynamic Programming)Model-Based • Exact4.2 Model-Based Methods