數位語音訊號處理
Digital Speech Processing
| 節 | 週二 |
|---|---|
5 13:20–14:10 | 數位語音訊號處理 EDB06 3 節連堂 |
6 14:20–15:10 | |
7 15:30–16:20 |
* 根據陽明交大上課時間表所列
因為人類最主要的溝通工具就是語音,所以語音訊號處理一直是人工智慧研究的主要議題,其最終目標是建構具語音人機介面,可以跟人直接對答的智慧機器人。因此,本課程將著重在介紹語音訊號處理的各種基礎知識與技術,並實際利用機器學習/深度學習技術進行實現。課程內容包括如何建立語音訊號處理系統(例如語音辨認,語者辨認,語言辨認,語音合成,語音轉換等),並將相關知識延伸至音訊訊號處理(例如語音增強,麥克風陣列,音源定位與分離,聲音事件偵測等)。最後並將實際操作相關語音訊號處理工具程式庫,進行系統實作,使學生能建立語音信號處理技術之知識及能力。 Since the main communication tool of human beings is speech, digital speech processing (DSP) has always been the top research issue of modern artificial intelligence. Therefore, this course will focus on fundamental theories and practices of DSP technology. The course content includes first how to build a DSP system (such as speech recognition, speaker recognition, language recognition, speech synthesis, speech conversion, etc.), and then extends to audio signal processing (ASP, such as speech enhancement, microphone array, sound source localization and separation, sound event detection, etc.). To this end, several popular DSP toolkits and, especially, related machine learning/deep learning techniques will be introduced to help students build the knowledge and competencies in DSP technology.
具備程式能力 熟習Linux作業系統與與shell script使用方式 Programming linux system and shell script
Toolkits: 1. Kaldi: https://github.com/kaldi-asr/kaldi 2. ESPnet: https://github.com/espnet/espnet 3. SpeechBrain: https://github.com/speechbrain/speechbrain
Online InClass Kaggle Competitions (https://www.kaggle.com/) with corresponding Github projects (programs), reports and oral presentations, at least three times, 100%
| 週次 | 主題 |
|---|---|
| 第 1 週 | Introduction to Speech Signal Processing 1. Introduction to Speech Recognition, Synthesis, Voice Conversion, Speaker and Language Recognition 2. Speech Production, Hearing, and Understanding 3, Audio and Acoustic Environment |
| 第 2 週 | 1. Speech Recognition - Introduction, Problem, and Challenge 2. Speech Recognition - Feature Extraction, Pattern Recognition, VQ, * Homework #1 - Feature Extraction with Python |
| 第 3 週 | 參加 25th NRC-NSTC Anniversary Celebrations Workshop on AI and 3D Technologies (NRC-NSTC),暫停一次(期中考週補課)。 |
| 第 4 週 | Speech Recognition - Statistical Model for Speech Recognition (1/4), Probability Model |
| 第 5 週 | 國慶日 |
| 第 6 週 | Speech Recognition - Statistical Model for Speech Recognition (2/4), GMMs |
| 第 7 週 | Speech Recognition - Statistical Model for Speech Recognition (3/4), HMMs |
| 第 8 週 | Speech Recognition - Statistical Model for Speech Recognition (4/4), Decoding |
| 第 9 週 | Speech Recognition - Deep Learning for Speech Recognition (1/2) * Homework #3 - InClass Kaggle Competition II - Speech Recognition with ESPnet |
| 第 10 週 | Speech Recognition - Deep Learning for Speech Recognition (2/2) * Homework #3 補充 - Transformer |
| 第 11 週 | Speech Synthesis - Statistic Model-based Speech Synthesis * Homework #3 補充 - Wav2Vec, WaveLM |
| 第 12 週 | Speech Synthesis - Deep Learning for Speech Synthesis * Homework #4 - Speech Synthesis with Vits2 |
| 第 13 週 | 出國參加O-COCOSDA 2023,暫停一次(期末考週補課)。 |
| 第 14 週 | 老師回診,暫停一次(期末考週補課) |
| 第 15 週 | 參加 ASRU 2023, 暫停一次(第17週補課) |
| 第 16 週 | Voice Conversion |
| 第 17 週 | Speaker and Langauge Recognition |
自編講義 Lecture Notes available
- 地點
- EF375
- 時間
- Wednesday 9:00-12:00 AM
- 聯絡方式
- Line Group and Email