資料科學與決策支援
Data Science and Decision Support
| 節 | 週三 |
|---|---|
2 09:00–09:50 | 資料科學與決策支援 MB110 3 節連堂 |
3 10:10–11:00 | |
4 11:10–12:00 |
* 根據陽明交大上課時間表所列
The primary objective of this course will guide students to follow a PDCA (plan-do-check-action) loop in data science to solve real problems: defining your problem, selecting appropriate methods, evaluating the performance, and modifying the constructed models. In addition, the main objectives of this course are summarized as follows: 1. Applying statistical skills to real problems (quality control), 2. Applying clustering skills to real problems (target marketing), 3. Applying classification skills to real problems (bankruptcy prediction), 4. Applying regression skills to real problems (demand forecasting) 5. Applying dimension-reduction skills to real problems (business intelligence).
The prerequisites for this course are statistics and basic programming. Students are required to take laptops in the classroom for practicing coding skills and handling real datasets. This course is expected to be applied to two major areas: machine intelligence and business analytics. Course loading is heavy, totally different from the style of case oriented in-class discussing. Students are expected to employ the skills learned in class to conduct data-driven decision making & support.
*The course schedule may be subject to change.
Assessment Take-home assignment (4 times) using R package 60 % Midterm exam 20% Final exam 20% Total 100 % *Details will be announced in the first class.
| 週次 | 主題 |
|---|---|
| 第 1 週 | Introduction to data science and the top 10 algorithms |
| 第 2 週 | Overview of statistics and R programming |
| 第 3 週 | Data processing (outlier detection, Chi-square test, proportion test) |
| 第 4 週 | Statistical analysis (one-tail/two-tail T-test, ANOVA, regression) |
| 第 5 週 | Overview of data mining and typical applications/ HW1 due |
| 第 6 週 | Clustering (K-means, K-medoids, C-means) |
| 第 7 週 | Clustering (Gaussian mixture modeling, hierarchical clustering, DBSCAN) |
| 第 8 週 | Association (Apriori algorithm) |
| 第 9 週 | Association (sequential rule mining)/ HW2 due |
| 第 10 週 | Basic classifiers (Naive Bayes, KNN, Logit/Probit regression) |
| 第 11 週 | Decision tree (C4.5, CART) |
| 第 12 週 | Ensemble learning (random forest, bagging, boosting ) /HW 3 due |
| 第 13 週 | Advanced classifiers (support vector machine, artificial neural network) |
| 第 14 週 | Statistical regression (MLR, MARS, PLS) |
| 第 15 週 | Machine-learning based regression (support vector machine , neural network, random forest) |
| 第 16 週 | Special regression (Ridge & amp Lasso)/ HW4 due |
1. Introduction to data mining (textbook), Tan et al., Pearson. 2. Data mining and business analytics: Concepts, Techniques, and Applications in R, Shmueli et al., Wiley. 3. Personal handouts for R coding in data science. 4. Published academic papers and industrial news/reports.
- 地點
- MB411
- 時間
- Professor's office hour
- 聯絡方式
- Available online (e3 campus)