基於深度學習之視覺辨識專論
Selected Topics in Visual Recognition using Deep Learning
| 節 | 週二 |
|---|---|
5 13:20–14:10 | 基於深度學習之視覺辨識專論 EC114 3 節連堂 |
6 14:20–15:10 | |
7 15:30–16:20 |
* 根據陽明交大上課時間表所列
Computer vision aims to empower computers with the ability to "see" – to perceive, understand, and interpret the visual world much like humans do. Deep learning has emerged as the driving force behind the current computer vision revolution. The availability of massive, annotated datasets, coupled with the accessibility of powerful GPUs, has enabled the training of complex deep learning models. These models, consisting of hundreds of layers and millions of parameters, have significantly advanced the performance of numerous computer vision applications. In this course, we will begin by exploring key deep learning architectures crucial for computer vision research. This will include a deep dive into foundational concepts such as deep neural networks, convolutional neural networks (CNNs), Transformers, and denoising diffusion models. Subsequently, we will delve into several important computer vision applications, including object recognition, detection, segmentation, low-level vision, and 3D vision. For each application, we will examine the state-of-the-art deep learning algorithms that are driving progress in that area.
1. Foundational mathematical skills, including linear algebra and calculus 2. Programming experience with Python and common libraries 3. Deep learning programming skills with frameworks like PyTorch
Four homework assignments 64% (=16% x 4) Final project 36%
| 週次 | 主題 |
|---|---|
| 第 1 週 | Introduction |
| 第 2 週 | Deep Neural Networks |
| 第 3 週 | Convolutional Neural Networks |
| 第 4 週 | Transformers |
| 第 5 週 | Object Detection I |
| 第 6 週 | Object Detection II |
| 第 7 週 | Object Segmentation I |
| 第 8 週 | Object Segmentation II |
| 第 9 週 | Denoising Diffusion Models |
| 第 10 週 | Low-level Vision |
| 第 11 週 | Mamba |
| 第 12 週 | 3D Point Clouds, Neural Radiance Fields (NeRF), and 3D Gaussian Splatting (3DGS) |
| 第 13 週 | 3D Vision |
| 第 14 週 | Guest Lecture (Date subject to change based on speakers' schedule) |
| 第 15 週 | Final Project Presentation I |
| 第 16 週 | Final Project Presentation II |
1. Ian Goodfellow, Yoshua Bengio, and Aaron Courville, Deep Learning, MIT Press, 2016 2. Richard Szeliski, Computer Vision: Algorithms and Applications, Springer, 2022
- 地點
- EC706 (Instructor) EC234-C or EC701 (TAs) Please send us an email in advance to make an appointment, and we will inform you where to have a discussion.
- 時間
- Tuesday 4:20 pm ~ 5:20 pm
- 聯絡方式
- Instructor: Yen-Yu Lin (林彥宇) Email: lin@cs.nycu.edu.tw TAs: Tsung-Lin Tsai (蔡宗霖) Email: sean19990323123@gmail.com Jian-Zhe Wang (王健哲) Email: jzwang.cs13@nycu.edu.tw Yi-Jen Tsai (蔡宜蓁) Email: tsai.cs14@nycu.edu.tw Nai-Yun Hsiao (蕭乃云) Email: alllllvin21292@gmail.com