Reinforcement learning replaces the dataset with an interaction loop. An agent in state \(s\) takes action \(a\), receives reward \(r\), and lands in state \(s'\). Nobody says which action was correct — only how much reward followed — and the consequences may arrive many steps later. The whole subject is built on one recursive identity, the Bellman equation, and on whether you have a model of the environment or must learn without one.
By the end of this unit you should be able to:
Learn directly from experience, without ever building a model of how the environment behaves.
2 algorithms • Q-Learning, Policy Gradient (REINFORCE) 4.2Learn or be given the environment’s dynamics, then plan against them.
2 algorithms • Dyna-Q, Value Iteration (Dynamic Programming)| No. | Algorithm | Character | Topic |
|---|---|---|---|
| 4.1 | Q-Learning | Off-Policy • Temporal Difference | 4.1 Model-Free Methods |
| 4.2 | Policy Gradient (REINFORCE) | On-Policy • Continuous Actions | 4.1 Model-Free Methods |
| 4.3 | Dyna-Q | Model-Based • Planning | 4.2 Model-Based Methods |
| 4.4 | Value Iteration (Dynamic Programming) | Model-Based • Exact | 4.2 Model-Based Methods |