校際選修

115-1 選課時程

進行中

  • 初選第一階段 6/15/2026
  • 初選第二階段 6/22/2026
  • 校際選修 8/24/2026
  • 初選第三階段 8/31/2026
  • 開學後加退選 9/7/2026
  • 逾期加退選 9/21/2026
選課資源

強化學習原理

Reinforcement Learning

學期
112-2
學分
3 學分
當期課號
535514
永久課號
CSIC30046
開課單位
資訊科學與工程研究所
授課教師
謝秉均
校區
光復
類別
選修
上課時間表
週一
週四
3
10:10–11:00
強化學習原理
ED202
2 節連堂
4
11:10–12:00
7
15:30–16:20
強化學習原理
ED202

* 根據陽明交大上課時間表所列

概述

- Learn how to model tasks as RL problems. - Understand RL from a theoretical viewpoint - Learn how to systematically solve RL problems by using various RL algorithms and perform analysis of these algorithms - Learn how to implement deep RL algorithms using software packages (e.g. Tensorflow and Pytorch) through team projects

先修科目

- Some math maturity: Familiarity with calculus and probability (basic understanding of numerical optimization would help) - Programming language: Python (familiarity with Tensorflow/Pytorch would help)

評分方式

Homework: 35% Theory Project: 30% Team Implementation Project: 35% (including 10% for presentation)

週次計畫
週次主題
第 1 週- Course Logistics - Markov Decision Process (MDP)
第 2 週- Planning in MDPs - Bellman Equations - Value Iteration - Policy Iteration - Regularized MDPs
第 3 週- Policy Optimization - Introduction to Optimization (Convexity, Smoothness, Gradient Descent, and Mirror Descent) - Policy Gradient (PG)
第 4 週- Stochastic PG (REINFORCE, A2C, and Natural PG) - Variance Reduction
第 5 週- Model-Free Prediction - Generalized Advantage Estimation
第 6 週- Global Convergence of Policy Gradient - Global Convergence of Natural PG
第 7 週- Value Function Approximation
第 8 週- Deterministic PG, DDPG, TD3 - Off-Policy Learning via Deterministic and Stochastic Policy Gradients
第 9 週- Trust Region Policy Optimization (TRPO) - Global Convergence of TRPO - Proximal Policy Optimization (PPO)
第 10 週- Value-Based Methods and Stochastic Approximation - Sarsa, Expected Sarsa, Q-Learning, and Double Q-Learning
第 11 週- Distributional Perspective of MDPs - Distributional RL (C51, QR-DQN, and IQN)
第 12 週- Entropy-Regularized RL - Soft Q-learning - Soft Actor-Critic
第 13 週- Reinforcement Learning from Human Feedback (RLHF) - Recent Theoretical Results on RLHF - Dueling Bandits
第 14 週- Imitation Learning - Inverse Reinforcement Learning (GAIL, WAIL, AIL, and IQ-Learn)
第 15 週- Upside-Down RL - Sequence-to-Sequence Modeling for RL
第 16 週- No class (exam week)
第 17 週- Final Presentations
第 18 週
教科書

- Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, MIT Press, 2nd edition, 2018 - Alekh Agarwal, Nan Jiang, and Sham M. Kakade, Reinforcement Learning: Theory and Algorithms, 2020 - Nocedal, Jorge, and Stephen Wright. Numerical optimization. Springer Science & Business Media, 2006 - Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning. arXiv 2016 - Tor Lattimore and Csaba Szepesvari, Bandit Algorithms. 2019

Office Hours
地點
EC418
時間
1pm-1:30pm on Mondays
聯絡方式
By email: pinghsieh@nycu.edu.tw