強化學習原理(英文授課)
Reinforcement Learning
| 節 | 週二 | 週五 |
|---|---|---|
3 10:10–11:00 | 強化學習原理(英文授課) EDB27 2 節連堂 | |
4 11:10–12:00 | ||
7 15:30–16:20 | 強化學習原理(英文授課) EDB27 |
* 根據陽明交大上課時間表所列
- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team project
- Some math maturity: Familiarity with calculus and probability (basic understanding of optimization would help) - Programming language: Python (familiarity with Tensorflow/Pytorch would help)
*Important Announcements* 1. The lectures for the first week (i.e., 2/15 and 2/18) will be delivered via Webex: https://nycu.webex.com/nycu/j.php?MTID=me3ff04a5c477f67f4ee5a42ccc23f8f0 2. For those who would like to ask for manual registration, please fill out the following form by 9pm, 2/15 (Tuesday). https://forms.gle/7BmqWfejy6oWct5w9
Homework: 30% Theory Project: 30% Team Implementation Project: 40% (including 10% for presentation)
| 週次 | 主題 |
|---|---|
| 第 1 週 | We will try our best to discuss all of the following topics: 1. Markov decision process (MDP) and planning in MDPs 2. Distributional Perspective of MDPs 3. Policy Optimization and Gradient Descent 4. Stochastic Policy Gradient Methods (REINFORCE, A2C, NPG) 5. Variance Reduction and Model-Free Prediction 6. Global Convergence of Policy Gradient 7. Value Function Approximation 8. Trust Region Policy Optimization: TRPO, PPO, and CPO 9. Deterministic Policy Gradient for Continuous Control 10. Off-Policy Learning via Deterministic and Stochastic Policy Gradients 11. Value-Based Methods and Stochastic Approximation (Expected Sarsa, Q-Learning, and Double Q-Learning) 12. Distributional RL (C51, QR-DQN, and IQN) 13. Soft Actor Critic 14. Imitation Learning and Inverse Reinforcement Learning |
Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2020 Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019
- 地點
- EC713
- 時間
- 4:30pm-5:20pm, Tuesdays (starting from 2/22)
- 聯絡方式
- By email: pinghsieh@nctu.edu.tw