數位語音訊號處理
Digital Speech Processing
| 節 | 週二 |
|---|---|
5 13:20–14:10 | 數位語音訊號處理 A301 3 節連堂 |
6 14:20–15:10 | |
7 15:30–16:20 |
* 根據陽明交大上課時間表所列
因為人類最主要的溝通工具就是語音,所以語音訊號處理一直是人工智慧研究的主要議題,其最終目標是建構具語音人機介面,可以跟人直接對答的智慧機器人。因此,本課程將著重在介紹語音訊號處理的各種基礎知識與技術,並實際利用機器學習/深度學習技術進行實現。課程內容包括如何建立語音訊號處理系統(例如語音辨認,語者辨認,語言辨認,語音合成,語音轉換等),並將相關知識延伸至音訊訊號處理(例如語音增強,麥克風陣列,音源定位與分離,聲音事件偵測等)。最後並將實際操作相關語音訊號處理工具程式庫,進行系統實作,使學生能建立語音信號處理技術之知識及能力。 Since the main communication tool of human beings is speech, digital speech processing (DSP) has always been the top research issue of modern artificial intelligence. Therefore, this course will focus on fundamental theories and practices of DSP technology. The course content includes first how to build a DSP system (such as speech recognition, speaker recognition, language recognition, speech synthesis, speech conversion, etc.), and then extends to audio signal processing (ASP, such as speech enhancement, microphone array, sound source localization and separation, sound event detection, etc.). To this end, several popular DSP toolkits and, especially, related machine learning/deep learning techniques will be introduced to help students build the knowledge and competencies in DSP technology.
具備程式能力 熟習Linux作業系統與與shell script使用方式 Programming linux system and shell script
Toolkits: 1. Kaldi: https://github.com/kaldi-asr/kaldi 2. ESPnet: https://github.com/espnet/espnet 3. SpeechBrain: https://github.com/speechbrain/speechbrain
Online InClass Kaggle Competitions (https://www.kaggle.com/) with corresponding Github projects (programs), reports and oral presentations, at least three times, 100%
| 週次 | 主題 |
|---|---|
| 第 1 週 | Introduction to Speech Signal Processing 1. Introduction to Speech Recognition, Synthesis, Voice Conversion, Speaker and Language Recognition |
| 第 2 週 | Introduction to Speech Signal Processing 2. Speech Production, Hearing, and Understanding 3, Audio and Acoustic Environment |
| 第 3 週 | Speech Recognition 1. Introduction, Problem, and Challenge 2. Speech Recognition - Feature Extraction, Pattern Recognition, VQ, * Homework #1 - Feature Extraction with Python |
| 第 4 週 | Speech Recognition 1. Statistical Model for Speech Recognition (1/4), Probability Model |
| 第 5 週 | Speech Recognition 2. Statistical Model for Speech Recognition (2/4), GMMs * Homework #2 - InClass Kaggle Competition II - Speech Recognition with ESPnet |
| 第 6 週 | Speech Recognition 3. Statistical Model for Speech Recognition (3/4), HMMs |
| 第 7 週 | Speech Recognition 4. Statistical Model for Speech Recognition (4/4), Decoding |
| 第 8 週 | 期中考週,繼續上課,趕進度 |
| 第 9 週 | Speech Recognition 5. Deep Learning for Speech Recognition (1/2) * Homework #3 - InClass Kaggle Competition II - Speech Recognition with ESPnet + S3PRL + Whisper |
| 第 10 週 | Speech Recognition 6. Deep Learning for Speech Recognition (2/2) |
| 第 11 週 | Speech Synthesis 1. Statistic Model-based Speech Synthesis |
| 第 12 週 | Speech Synthesis 2. Deep Learning for Speech Synthesis * Homework #4 - Speech Synthesis with Vits2 |
| 第 13 週 | Speech Synthesis 3. Deep Learning for Speech Synthesis |
| 第 14 週 | Voice Conversion |
| 第 15 週 | Speaker and Langauge Recognition 1. Triple Loss |
| 第 16 週 | 期末考週(暫停上課) |
自編講義 Lecture Notes available
- 地點
- EF375
- 時間
- Wednesday 9:00-12:00 AM
- 聯絡方式
- Line Group and Email