# data process
# 需要在本地準備好原始資料,需修改main中的target_row範圍,決定處理的月份
# 需要新增名為dataset的資料夾以及名為output的資料夾
uv run main.py # 開始進行資料處理
uv run combine.py # 將同年份資料合併
# deep learning
uv run main.py # Start training and output test result at pres.csv
uv run bench.py # Do the benchmark for the model從kaggle取得youbike官方從2020四月到2025三月的使用紀錄,資料來源:Taipei YouBike 2.0 Rental Records
利用==main.py==輸入要下載的月份範圍,接著從第一個月開始迴圏
main傳送某月的位置給==download.py==執行
- 從Taipei YouBike 2.0 rental records.csv找到該月,抓取下載連結
- 下載到本地解壓縮
- 回傳單月份資料檔到main
- 刪除zip檔
main將收到的月份資料檔輸入給==catch.py==執行
- 統計當月每日使用量
- 將結果寫入「單月份.csv」檔,格式為
| 年月日 | 使用量 |
|---|
- 刪除月份原資料檔
- 回傳「單月份.csv」檔給main
main輸入「單月份.csv」檔給==change.py==執行
- 將
| 年月日 | 使用量 |
|---|
- 改為
| 年 | 月 | 日 | 使用量 |
|---|
- 將結果覆寫到原本的「單月份.csv」檔
處理 202004-202503 的資料後,使用==combine.py==將同一年的資料合併,完成資料處理
We divide the ubike.train.csv into training and validation set with a 9:1 ratio.
To prevent data leakage, we calculated the mean and standard deviation of features solely on the training set and applied these parameters to normalize all datasets.
Additionally, a logarithmic transformation was applied to the target variable to stabilize the training process; the results were transformed back to the original scale during testing for evaluation.
The model is a fully-connected neural network(MLP) consisting of 2 hidden layers with 32 and 16 neurons, respectively. Each layer sequentially perform a linear operation, followed by Batch Normalization, and a ReLU activation function.
We employed Mean Squared Error (MSE) as the loss function and used the Adam optimizer. To improve generalization and prevent overfitting, we implemented mini-batch training and an early stopping mechanism.
- True values mean: 189043.0000
- RMSE: 57228.4521
- MAPE: 34.6975 %
在這次期末專題中,我們在測試集上跑出了RMSE=57228.4521, MAPE=34.6975%的成績,這代表我們的預測人數與真實人數的平均誤差約為57000人,平均約偏離真實值35%。
以第一次處理這種大規模的迴歸預測任務且目前只透過日期進行預測來說,我們認為這已經算是不錯的成績了。
未來如果要繼續改進此模型,我們可能會選擇新增更多特徵並嘗試對時間特徵進行編碼。
Source of main.py: Heng-Jui Chang @ NTUEE (https://github.com/ga642381/ML2021-Spring/blob/main/HW01/HW01.ipynb)