遞迴神經網路與變形器
Recurrent Neural Network and Transformer
| 節 | 週二 |
|---|---|
5 13:20–14:10 | 遞迴神經網路與變形器 CM212 3 節連堂 |
6 14:20–15:10 | |
7 15:30–16:20 |
* 根據陽明交大上課時間表所列
This course provides an in-depth introduction to two significant AI models: Recurrent Neural Networks (RNNs) and Transformers. These models play a critical role in handling sequential data and natural language processing tasks. The course begins with an overview of the basic structure and operational principles of RNNs, focusing on how their recurrent architecture learns temporal dependencies. It also delves into how Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRUs) address the long-term dependency issues inherent in RNNs, highlighting their respective advantages. Next, we will explore the fundamental architecture and principles of Transformers, including their core mechanisms such as self-attention and multi-head attention. The course will also introduce the development and applications of Bidirectional Encoder Representations from Transformers (BERT) and related models. The applications of Transformers extend beyond natural language processing to areas like image processing and computer vision. Furthermore, we will examine the development of the latest AI models based on Transformers, with particular focus on their applications in large language models (LLMs), such as the GPT series. Topics include continual domain-adaptive pretraining, low-rank fine-tuning, retrieval-augmented generation, and other techniques. Through practical examples and programming assignments, this course aims to help students understand how to apply these network models to real-world problems. Objectives: Objective 1: Develop students' understanding and ability to apply Recurrent Neural Networks (RNNs, LSTMs, GRUs) for handling time-series data and natural language texts. Objective 2: Equip students with a mastery of Transformers and their applications in natural language processing, image processing, and large language models, as well as familiarity with the latest related technologies. Objective 3: Introduce the applications of Transformers in image processing and computer vision, such as Vision Transformer, SWIN, DETR, MAE, and BEiT.
None
Instruction and project-based learning. Course material is available on E3 learning platform.
Programming homework.: 60% (4x15%) Paper presentation: 10% Final project: 30%
| 週次 | 主題 |
|---|---|
| 第 1 週 | Course Introduction RNN 1: Network Architectures |
| 第 2 週 | RNN 2: Learning Processes RNN 3: Recurrent Neural Networks |
| 第 3 週 | RNN 4: LSTM |
| 第 4 週 | RNN 5 : LSTM, Seq-to-seq model, LSTM with Attention |
| 第 5 週 | RNN 6: Gated Recurrent Unit (GRU), Minimal Gated Unit (MGU) |
| 第 6 週 | Transformer |
| 第 7 週 | BERT |
| 第 8 週 | Pretraining a RoBERTa Model from Scratch |
| 第 9 週 | Downstream NLP Tasks with Transformers |
| 第 10 週 | Text Generation with OpenAI GPT models |
| 第 11 週 | Recent Development of Large Language Models |
| 第 12 週 | Recent Development of Computer Vision using Transformers Vision Transformer (ViT), BERT Pre-Training of Image Transformers (BEiT), End-to-End Object Detection with Transformers (DETR), CF-DERT, etc. |
| 第 13 週 | Continue Learning, Fine tune, Retrieval Augmented Generation of LLM |
| 第 14 週 | Paper Presentation |
| 第 15 週 | Paper presentation |
| 第 16 週 | Final project demo |
| 第 17 週 | |
| 第 18 週 |
1. Fathi M. Salem, Recurrent Neural Networks, Springer, 2022. ISBN 978-3-030-89928-8 2. Denis Rothman, Transformers for Natural Language Processing, Packt Publishing, Jan. 2021. ISBN: 9781800565791
- 地點
- ChiMei 303
- 時間
- Tuesday 10:00-12:00AM
- 聯絡方式
- 55729