NYC Mobility Forecasting - ML-Driven Taxi/Rideshare Demand Prediction System
I built an end-to-end ML system forecasting NYC taxi/rideshare demand across 263 zones, cutting prediction error (RMSE) by 43% over a baseline Random Forest and serving live predictions through a REST API and dashboard.
01
TL;DR
- I built an end-to-end ML system forecasting NYC taxi/rideshare demand across 263 zones, cutting prediction error (RMSE) by 43% over a baseline Random Forest and serving live predictions through a REST API and dashboard.
- Best published result: Weighted MAPE improved 18.77% -> 16.52%
02
Problem
- Forecast NYC taxi/FHV demand at zone/hour granularity to support supply planning and surge-risk awareness Rideshare/taxi operations planning and demand-forecasting use cases. Accurate zone/hour demand forecasts reduce both undersupply and oversupply risk in a highly variable transportation market.
03
My Role
- personal build with end-to-end ownership
- Problem framing: Forecast NYC taxi/FHV demand at zone/hour granularity to support supply planning and surge-risk awareness
- Architecture: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
- Implementation: PySpark, Delta Lake, PostgreSQL, Redis caching
- Evaluation: Weighted MAPE improved 18.77% -> 16.52%; RMSE cut by 43%, MAE cut by 53%; 263 zones, 2.3M-row hourly demand dataset
- Before: Accurate zone/hour demand forecasts reduce both undersupply and oversupply risk in a highly variable transportation market
- Personally designed: Used a two-stage (classification + regression) Poisson hurdle model instead of a single regression model because Better captures the zero-inflated nature of hourly zone-level demand (many zone/hours have zero trips) than a plain regression baseline; Built a Spark-based medallion (bronze/silver/gold) pipeline because Processes trip and weather data into a clean 2.3M-row hourly demand dataset at scale with reproducible stages; Targeted free-tier Oracle Cloud for deployment because Keeps the project runnable end-to-end without ongoing cloud spend
- Others owned: No separate collaborator-owned subsystem is published in the source data.
04
Constraints
- Built Jul 2026 (personal). Large-scale data processing (2.3M-row hourly dataset); free-tier cloud deployment budget.
05
Architecture
- Input: NYC TLC trip data, weather data, geospatial zone data
- Backend: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
- Data & storage: PySpark, Delta Lake, PostgreSQL, Redis caching
- External APIs: Weather data feed
- Output: React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions
Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model; PySpark, Delta Lake, PostgreSQL, Redis caching; React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions
- input 01Input
NYC TLC trip data, weather data, geospatial zone data
- process 02Backendinput ->
Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
- storage 03Data / storagebackend ->
PySpark, Delta Lake, PostgreSQL, Redis caching
- external 04External APIsbackend ->
Weather data feed
- output 05Outputstorage ->external ->
React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions
Routes
- Input -> Backend
- Backend -> Data / storage
- Backend -> External APIs
- Data / storage -> Output
- External APIs -> Output
06
Key Technical Decisions
- Used a two-stage (classification + regression) Poisson hurdle model instead of a single regression model
- Built a Spark-based medallion (bronze/silver/gold) pipeline
- Targeted free-tier Oracle Cloud for deployment
07
Implementation
- Input layer: NYC TLC trip data, weather data, geospatial zone data
- Core system: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
- Data layer: PySpark, Delta Lake, PostgreSQL, Redis caching
- External boundary: Weather data feed
- User output: React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions
08
What Broke / What Didn't Work
- Rejected: Single-stage Random Forest regression (baseline). Chosen path: Used a two-stage (classification + regression) Poisson hurdle model instead of a single regression model.
- Rejected: Single-pass pandas ETL script. Chosen path: Built a Spark-based medallion (bronze/silver/gold) pipeline.
- Rejected: AWS/GCP paid tiers. Chosen path: Targeted free-tier Oracle Cloud for deployment.
- Two-stage model adds pipeline complexity over a single regressor but meaningfully improves accuracy on sparse-demand zones
- Free-tier cloud deployment constrains available compute/memory for serving
09
Results
- Weighted MAPE improved 18.77% -> 16.52% - Forecast accuracy vs baseline Random Forest - Forecast accuracy vs baseline Random Forest - master-resume
- RMSE cut by 43%, MAE cut by 53% - Prediction error vs baseline Random Forest - Prediction error vs baseline Random Forest - master-resume
- 263 zones, 2.3M-row hourly demand dataset - Scale of forecasting system and processed dataset - Scale of forecasting system and processed dataset - master-resume
10
What I'd Change Now
- Add real-time streaming ingestion instead of batch pipeline
- Expand to additional cities/regions
- Incorporate event-based demand signals (concerts, holidays)
11
Stack
- Python
- PySpark
- Apache Spark
- Delta Lake
- scikit-learn
- FastAPI
- SQL
- PostgreSQL
- Redis
- Docker
- Oracle Cloud
- React
- REST API
- CI/CD
12
Links
- Source docs: 2-projects.json
Ask me about the trade-offs.
- Why this architecture boundary exists: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
- How I evaluated Forecast accuracy vs baseline Random Forest
- The hardest tradeoff: Two-stage model adds pipeline complexity over a single regressor but meaningfully improves accuracy on sparse-demand zones
- What I would change next: Add real-time streaming ingestion instead of batch pipeline