強化學習原理(英文授課)
Reinforcement Learning
| 節 | 週二 | 週五 |
|---|---|---|
2 09:00–09:50 | 強化學習原理(英文授課) EC115 | |
5 13:20–14:10 | 強化學習原理(英文授課) EC115 2 節連堂 | |
6 14:20–15:10 |
* 根據陽明交大上課時間表所列
- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team project
- Some math maturity: Familiarity with calculus and probability (basic understanding of optimization would help ) - Programming language: python (familiarity with tensorflow/pytorch would help)
Homework: 30% Theory Project: 30% Team Implementation Project: 40% (including 10% for presentation)
| 週次 | 主題 |
|---|---|
| 第 1 週 | We will try our best to discuss all of the following topics: 1. Markov decision process (MDP) and planning in MDPs 2. Introduction to numerical optimization: line search and trust region methods 3. Model-free policy evaluation 4. Model-free control and stochastic approximation 5. Policy-gradient methods (REINFORCE, A2C, DPG) 6. Function approximation 7. Trust-region policy optimization: TRPO, PPO, and CPO 8. Deep value-based approach 9. Off-policy evaluation 10. Distributional RL 11. Multi-armed Bandits |
Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2019 Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019
- 地點
- EC713
- 時間
- By appointment
- 聯絡方式
- By email: pinghsieh@nctu.edu.tw