Skip to content

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ubike-usage-ml-prediction

Usage

# data process
# 需要在本地準備好原始資料,需修改main中的target_row範圍,決定處理的月份
# 需要新增名為dataset的資料夾以及名為output的資料夾
uv run main.py # 開始進行資料處理
uv run combine.py # 將同年份資料合併

# deep learning
uv run main.py # Start training and output test result at pres.csv
uv run bench.py # Do the benchmark for the model

DataProcess

資料來源

從kaggle取得youbike官方從2020四月到2025三月的使用紀錄,資料來源:Taipei YouBike 2.0 Rental Records


程式碼運作

利用==main.py==輸入要下載的月份範圍,接著從第一個月開始迴圏

download

main傳送某月的位置給==download.py==執行

  1. 從Taipei YouBike 2.0 rental records.csv找到該月,抓取下載連結
  2. 下載到本地解壓縮
  3. 回傳單月份資料檔到main
  4. 刪除zip檔

catch

main將收到的月份資料檔輸入給==catch.py==執行

  1. 統計當月每日使用量
  2. 將結果寫入「單月份.csv」檔,格式為
年月日 使用量
  1. 刪除月份原資料檔
  2. 回傳「單月份.csv」檔給main

change

main輸入「單月份.csv」檔給==change.py==執行

  • 將
年月日 使用量
  • 改為
年 月 日 使用量
  • 將結果覆寫到原本的「單月份.csv」檔

combine

處理 202004-202503 的資料後,使用==combine.py==將同一年的資料合併,完成資料處理

Deep Learning Implementation

Data Pre-processing

We divide the ubike.train.csv into training and validation set with a 9:1 ratio.

To prevent data leakage, we calculated the mean and standard deviation of features solely on the training set and applied these parameters to normalize all datasets.

Additionally, a logarithmic transformation was applied to the target variable to stabilize the training process; the results were transformed back to the original scale during testing for evaluation.

Neural Network Architecture

The model is a fully-connected neural network(MLP) consisting of 2 hidden layers with 32 and 16 neurons, respectively. Each layer sequentially perform a linear operation, followed by Batch Normalization, and a ReLU activation function.

Training Strategy

We employed Mean Squared Error (MSE) as the loss function and used the Adam optimizer. To improve generalization and prevent overfitting, we implemented mini-batch training and an early stopping mechanism.

Benchmark

  • True values mean: 189043.0000
  • RMSE: 57228.4521
  • MAPE: 34.6975 %

Summary

在這次期末專題中,我們在測試集上跑出了RMSE=57228.4521, MAPE=34.6975%的成績,這代表我們的預測人數與真實人數的平均誤差約為57000人,平均約偏離真實值35%。

以第一次處理這種大規模的迴歸預測任務且目前只透過日期進行預測來說,我們認為這已經算是不錯的成績了。

未來如果要繼續改進此模型,我們可能會選擇新增更多特徵並嘗試對時間特徵進行編碼。

Reference

Source of main.py: Heng-Jui Chang @ NTUEE (https://github.com/ga642381/ML2021-Spring/blob/main/HW01/HW01.ipynb)

About

資料科學與問題解決-期末專題

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages