校際選修

115-1 選課時程

進行中

  • 初選第一階段 6/15/2026
  • 初選第二階段 6/22/2026
  • 校際選修 8/24/2026
  • 初選第三階段 8/31/2026
  • 開學後加退選 9/7/2026
  • 逾期加退選 9/21/2026
選課資源

強化學習原理(英文授課)

Reinforcement Learning

學期
110-2
學分
3 學分
當期課號
5259
永久課號
IOC5208
開課單位
資訊科學與工程研究所
授課教師
謝秉均
校區
光復
類別
選修
上課時間表
週二
週五
3
10:10–11:00
強化學習原理(英文授課)
EDB27
2 節連堂
4
11:10–12:00
7
15:30–16:20
強化學習原理(英文授課)
EDB27

* 根據陽明交大上課時間表所列

概述

- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team project

先修科目

- Some math maturity: Familiarity with calculus and probability (basic understanding of optimization would help) - Programming language: Python (familiarity with Tensorflow/Pytorch would help)

教學方式

*Important Announcements* 1. The lectures for the first week (i.e., 2/15 and 2/18) will be delivered via Webex: https://nycu.webex.com/nycu/j.php?MTID=me3ff04a5c477f67f4ee5a42ccc23f8f0 2. For those who would like to ask for manual registration, please fill out the following form by 9pm, 2/15 (Tuesday). https://forms.gle/7BmqWfejy6oWct5w9

評分方式

Homework: 30% Theory Project: 30% Team Implementation Project: 40% (including 10% for presentation)

週次計畫
週次主題
第 1 週We will try our best to discuss all of the following topics: 1. Markov decision process (MDP) and planning in MDPs 2. Distributional Perspective of MDPs 3. Policy Optimization and Gradient Descent 4. Stochastic Policy Gradient Methods (REINFORCE, A2C, NPG) 5. Variance Reduction and Model-Free Prediction 6. Global Convergence of Policy Gradient 7. Value Function Approximation 8. Trust Region Policy Optimization: TRPO, PPO, and CPO 9. Deterministic Policy Gradient for Continuous Control 10. Off-Policy Learning via Deterministic and Stochastic Policy Gradients 11. Value-Based Methods and Stochastic Approximation (Expected Sarsa, Q-Learning, and Double Q-Learning) 12. Distributional RL (C51, QR-DQN, and IQN) 13. Soft Actor Critic 14. Imitation Learning and Inverse Reinforcement Learning
教科書

Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2020 Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019

Office Hours
地點
EC713
時間
4:30pm-5:20pm, Tuesdays (starting from 2/22)
聯絡方式
By email: pinghsieh@nctu.edu.tw