強化學習原理
Reinforcement Learning
| 節 | 週二 | 週五 |
|---|---|---|
3 10:10–11:00 | 強化學習原理 EC115 2 節連堂 | |
4 11:10–12:00 | ||
7 15:30–16:20 | 強化學習原理 EC115 |
* 根據陽明交大上課時間表所列
- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team project
- Some math maturity: Familiarity with calculus and probability (basic understanding of optimization would help) - Programming language: Python (familiarity with Tensorflow/Pytorch would help)
**Important** The first lecture on 2/14 (Tuesday) will be delivered via Webex at https://nycu.webex.com/nycu/j.php?MTID=m4c8d72b8f0c204e24785b52c2d9511cf
Homework: 30% Theory Project: 30% Team Implementation Project: 40% (including 10% for presentation)
| 週次 | 主題 |
|---|---|
| 第 1 週 | We will try our best to discuss all of the following topics:1. Markov decision process (MDP) and planning in MDPs2. Distributional Perspective of MDPs3. Policy Optimization and Gradient Descent4. Stochastic Policy Gradient Methods (REINFORCE, A2C, NPG)5. Variance Reduction and Model-Free Prediction6. Global Convergence of Policy Gradient7. Value Function Approximation8. Trust Region Policy Optimization: TRPO, PPO, and CPO9. Deterministic Policy Gradient for Continuous Control10. Off-Policy Learning via Deterministic and Stochastic Policy Gradients11. Value-Based Methods and Stochastic Approximation (Expected Sarsa, Q-Learning, and Double Q-Learning)12. Distributional RL (C51, QR-DQN, and IQN)13. Soft Actor Critic14. Imitation Learning and Inverse Reinforcement Learning |
| 第 2 週 | |
| 第 3 週 | |
| 第 4 週 | |
| 第 5 週 | |
| 第 6 週 | |
| 第 7 週 | |
| 第 8 週 | |
| 第 9 週 | |
| 第 10 週 | |
| 第 11 週 | |
| 第 12 週 | |
| 第 13 週 | |
| 第 14 週 | |
| 第 15 週 | |
| 第 16 週 | |
| 第 17 週 | |
| 第 18 週 |
Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2020 Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019
- 地點
- EC418
- 時間
- 4:30pm-5pm on Tuesdays
- 聯絡方式
- By email: pinghsieh@nycu.edu.tw