校際選修

115-1 選課時程

進行中

  • 初選第一階段 6/15/2026
  • 初選第二階段 6/22/2026
  • 校際選修 8/24/2026
  • 初選第三階段 8/31/2026
  • 開學後加退選 9/7/2026
  • 逾期加退選 9/21/2026
選課資源

強化學習專論

Selected Topics in Reinforcement Learning

學期
113-1
學分
3 學分
當期課號
535518
永久課號
CSIC30163
開課單位
資訊科學與工程研究所
授課教師
吳毅成
校區
光復
類別
選修
上課時間表
週二
A
18:30–19:20
強化學習專論
EC114
3 節連堂
B
19:30–20:20
C
20:30–21:20

* 根據陽明交大上課時間表所列

概述

This course is designed to teach students how to develop reinforcement learning (RL) algorithms for a wide range of applications, such as computer games, video games, intelligent traffic management, manufacturing scheduling, autonomous driving/racing, and robotics. The objectives are summarized as follows. (1) To understand the basic core concepts of reinforcement learning (RL) (2) To understand many latest RL techniques for applications (3) To familiarize with tools for developing RL, such as PyTorch, Gazebo, etc. (4) To develop practical working systems via projects such as DeepRacer.

先修科目

Machine Learning/Deep Learning (suggested)

評分方式

Projects (done individually) 50% Paper presentation (done in groups of 2 members) 20% Final exam 30%

課程大綱
  • Core of RL
  • Advanced Topics of RL
  • Presentation
  • Introduction to RL
週次計畫
週次主題
第 1 週Introduction to Reinforcement Learning
第 2 週Case studies of lightweight model applications: 2048 and Go
第 3 週Fundamentals: Markov Decision Process (MDP), Dynamic Programming (Tabular RL), Q-Learning, Function Approximation
第 4 週Value-Based Reinforcement Learning: DQN, DDQN (Double DQN), Dueling Network (with Advantage), Distributional DQN
第 5 週Policy-based Reinforcement Learning: Policy Gradient, Actor-Critic (Discrete actions), A2C and A3C (Asynchronous Advantage Actor-Critic)
第 6 週Policy-based Reinforcement Learning: TRPO &amp PPO, GAE, DDPG, TD3 (Continuous Actions), SAC (Soft Actor-Critic)
第 7 週Applications: DeepRacer: Augmentation, RL-cycleGAN, DrQ
第 8 週Applications: Solving Rubik Cube, RL for optimization (JSP/TSP)
第 9 週Exploration vs. Exploitation: Multi-Arm Bandits, UCB, Sequential Halving
第 10 週Planning: Dyna, Monte-Carlo Tree Search (MCTS), AlphaGo, AlphaZero, MuZero, Path Consistency, Abstraction
第 11 週Advanced Exploration: ICM, RND Experience Replay: PER, Ape-X
第 12 週Model-based RL: DQfD, R2D3 Multi-Agents RL (MARL) Q-mix, COMA
第 13 週Presentation
第 14 週Presentation
第 15 週Presentation
第 16 週Final exam
第 17 週Final competition for DeepRacer
第 18 週
教科書

1. R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, Nov. 2017 2. David Silver, Online Course for Deep Reinforcement Learning. http://www.cs.ucl.ac.uk/staff/D.Silver/web/Teaching.html 3. Papers and slides.

Office Hours
地點
TBA
時間
TBA
聯絡方式
TBA