機器學習晶片架構設計
Accelerator Architectures for Machine Learning
| 節 | 週四 |
|---|---|
3 10:10–11:00 | 機器學習晶片架構設計 ED302 2 節連堂 |
4 11:10–12:00 |
* 根據陽明交大上課時間表所列
Machine learning has captured tremendous successes to solve difficult learning problems. Hardware accelerators pursue continued performance and energy-efficient gains to meet the intensive computation in machine learning applications. This course explores leading approaches that tackle machine learning computational challenges and have been emerged in industrial and academic research. This course aims to build up students a foundation to understand the programming and accelerator architectural functions. This course begins with the fundamental basis of deep neural networks (DNN). The second potion of this course provides students accelerator hardware architectures specified for machine learning workloads. This course will address the graphic processing units (GPUs) that are widely used for the training of the neural networks and specialized machine learning accelerators such as tensor processor units (TPUs). The final portion of this course discusses challenges in designing accelerator architectures for machine learning applications and introduces emerging accelerator architectures. This course includes the programming assignments to use the computer architecture simulator, research paper reading and a class project to reflect ideas that improve accelerator architecture designs.
Computer architecture and digital logic circuit design
class website:https://people.cs.nycu.edu.tw/~ttyeh/course/2025_Fall/IOC5009/outline.html
10 % paper reading 40 % homework and lab assignments 20% midterm exam 30 % class project
- DNN Models
- GPU
- DNN accelerators
| 週次 | 主題 |
|---|---|
| 第 1 週 | Class Organization & amp
 amp Foundations of Deep Learning |
| 第 2 週 | DNN Methods and Models |
| 第 3 週 | DNN Kernel Computation |
| 第 4 週 | DNN Data Type Quantization |
| 第 5 週 | DNN Sparsity |
| 第 6 週 | Sparse DNN Accelerators |
| 第 7 週 | GPU Programming Model and Instruction Set Architecture |
| 第 8 週 | GPU SIMT Core architecture |
| 第 9 週 | GPU Memory System |
| 第 10 週 | Introduction to GPGPU-Sim Simulator |
| 第 11 週 | Machine Learning GPU Kernel Optimization |
| 第 12 週 | DNN Dataflow Accelerators Part I |
| 第 13 週 | DNN Dataflow Accelerators Part II |
| 第 14 週 | DNN Benchmarking (MLPerf) |
| 第 15 週 | DNN HW-SW Co-design (Model Pruning) |
| 第 16 週 | DNN Near/In Memory Processing |
| 第 17 週 | Advanced Technology for Accelerated ML |
| 第 18 週 | Conclusion |
1. Efficient Processing of Deep Neural Networks, Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, Joel S. Emer, Synthesis Lectures on Computer Architecture, Morgan & Claypool, 2020 2. Deep Learning for Computer Architects, Brandon Reagen, Robert Adolf, Paul Whatmough, Gu-Yeon Wei, and David Brooks, Synthesis Lectures on Comput-er Architecture, Morgan & Claypool, 2017 3. General-Purpose Graphics Processor Architectures, Tor M. Aamodt, Wilson Wai Lun Fung, and Timothy G. Rogers, Synthesis Lectures on Computer Archi-tecture, Morgan & Claypool, 2018 4. Programming Massively Parallel Processors: A Hands-on Approach, Kirk, D.B., & Hwu, W.M.W., 3rd Edition, Elsevier, Inc., 2016.
- 地點
- TBA
- 時間
- TBA
- 聯絡方式
- TBA