校際選修

115-1 選課時程

進行中

  • 初選第一階段 6/15/2026
  • 初選第二階段 6/22/2026
  • 校際選修 8/24/2026
  • 初選第三階段 8/31/2026
  • 開學後加退選 9/7/2026
  • 逾期加退選 9/21/2026
選課資源

數位語音訊號處理

Digital Speech Processing

學期
113-1
學分
3 學分
當期課號
539101
永久課號
IIAI30003
開課單位
智能系統研究所
授課教師
廖元甫
校區
光復
類別
選修
上課時間表
週二
5
13:20–14:10
數位語音訊號處理
A301
3 節連堂
6
14:20–15:10
7
15:30–16:20

* 根據陽明交大上課時間表所列

概述

因為人類最主要的溝通工具就是語音,所以語音訊號處理一直是人工智慧研究的主要議題,其最終目標是建構具語音人機介面,可以跟人直接對答的智慧機器人。因此,本課程將著重在介紹語音訊號處理的各種基礎知識與技術,並實際利用機器學習/深度學習技術進行實現。課程內容包括如何建立語音訊號處理系統(例如語音辨認,語者辨認,語言辨認,語音合成,語音轉換等),並將相關知識延伸至音訊訊號處理(例如語音增強,麥克風陣列,音源定位與分離,聲音事件偵測等)。最後並將實際操作相關語音訊號處理工具程式庫,進行系統實作,使學生能建立語音信號處理技術之知識及能力。 Since the main communication tool of human beings is speech, digital speech processing (DSP) has always been the top research issue of modern artificial intelligence. Therefore, this course will focus on fundamental theories and practices of DSP technology. The course content includes first how to build a DSP system (such as speech recognition, speaker recognition, language recognition, speech synthesis, speech conversion, etc.), and then extends to audio signal processing (ASP, such as speech enhancement, microphone array, sound source localization and separation, sound event detection, etc.). To this end, several popular DSP toolkits and, especially, related machine learning/deep learning techniques will be introduced to help students build the knowledge and competencies in DSP technology.

先修科目

具備程式能力 熟習Linux作業系統與與shell script使用方式 Programming linux system and shell script

教學方式

Toolkits: 1. Kaldi: https://github.com/kaldi-asr/kaldi 2. ESPnet: https://github.com/espnet/espnet 3. SpeechBrain: https://github.com/speechbrain/speechbrain

評分方式

Online InClass Kaggle Competitions (https://www.kaggle.com/) with corresponding Github projects (programs), reports and oral presentations, at least three times, 100%

週次計畫
週次主題
第 1 週Introduction to Speech Signal Processing 1. Introduction to Speech Recognition, Synthesis, Voice Conversion, Speaker and Language Recognition
第 2 週Introduction to Speech Signal Processing 2. Speech Production, Hearing, and Understanding 3, Audio and Acoustic Environment
第 3 週Speech Recognition 1. Introduction, Problem, and Challenge 2. Speech Recognition - Feature Extraction, Pattern Recognition, VQ, * Homework #1 - Feature Extraction with Python
第 4 週Speech Recognition 1. Statistical Model for Speech Recognition (1/4), Probability Model
第 5 週Speech Recognition 2. Statistical Model for Speech Recognition (2/4), GMMs * Homework #2 - InClass Kaggle Competition II - Speech Recognition with ESPnet
第 6 週Speech Recognition 3. Statistical Model for Speech Recognition (3/4), HMMs
第 7 週Speech Recognition 4. Statistical Model for Speech Recognition (4/4), Decoding
第 8 週期中考週,繼續上課,趕進度
第 9 週Speech Recognition 5. Deep Learning for Speech Recognition (1/2) * Homework #3 - InClass Kaggle Competition II - Speech Recognition with ESPnet + S3PRL + Whisper
第 10 週Speech Recognition 6. Deep Learning for Speech Recognition (2/2)
第 11 週Speech Synthesis 1. Statistic Model-based Speech Synthesis
第 12 週Speech Synthesis 2. Deep Learning for Speech Synthesis * Homework #4 - Speech Synthesis with Vits2
第 13 週Speech Synthesis 3. Deep Learning for Speech Synthesis
第 14 週Voice Conversion
第 15 週Speaker and Langauge Recognition 1. Triple Loss
第 16 週期末考週(暫停上課)
教科書

自編講義 Lecture Notes available

Office Hours
地點
EF375
時間
Wednesday 9:00-12:00 AM
聯絡方式
Line Group and Email