TBSystemTanmay
Bhuskute
Level 1 Case StudyData & ML Systems / Jul 2026

NYC Mobility Forecasting - ML-Driven Taxi/Rideshare Demand Prediction System

I built an end-to-end ML system forecasting NYC taxi/rideshare demand across 263 zones, cutting prediction error (RMSE) by 43% over a baseline Random Forest and serving live predictions through a REST API and dashboard.

Weighted MAPE improved 18.77% -> 16.52%RMSE cut by 43%, MAE cut by 53%263 zones, 2.3M-row hourly demand dataset

01

TL;DR

  • I built an end-to-end ML system forecasting NYC taxi/rideshare demand across 263 zones, cutting prediction error (RMSE) by 43% over a baseline Random Forest and serving live predictions through a REST API and dashboard.
  • Best published result: Weighted MAPE improved 18.77% -> 16.52%

02

Problem

  • Forecast NYC taxi/FHV demand at zone/hour granularity to support supply planning and surge-risk awareness Rideshare/taxi operations planning and demand-forecasting use cases. Accurate zone/hour demand forecasts reduce both undersupply and oversupply risk in a highly variable transportation market.

03

My Role

  • personal build with end-to-end ownership
  • Problem framing: Forecast NYC taxi/FHV demand at zone/hour granularity to support supply planning and surge-risk awareness
  • Architecture: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
  • Implementation: PySpark, Delta Lake, PostgreSQL, Redis caching
  • Evaluation: Weighted MAPE improved 18.77% -> 16.52%; RMSE cut by 43%, MAE cut by 53%; 263 zones, 2.3M-row hourly demand dataset
  • Before: Accurate zone/hour demand forecasts reduce both undersupply and oversupply risk in a highly variable transportation market
  • Personally designed: Used a two-stage (classification + regression) Poisson hurdle model instead of a single regression model because Better captures the zero-inflated nature of hourly zone-level demand (many zone/hours have zero trips) than a plain regression baseline; Built a Spark-based medallion (bronze/silver/gold) pipeline because Processes trip and weather data into a clean 2.3M-row hourly demand dataset at scale with reproducible stages; Targeted free-tier Oracle Cloud for deployment because Keeps the project runnable end-to-end without ongoing cloud spend
  • Others owned: No separate collaborator-owned subsystem is published in the source data.

04

Constraints

  • Built Jul 2026 (personal). Large-scale data processing (2.3M-row hourly dataset); free-tier cloud deployment budget.

05

Architecture

  • Input: NYC TLC trip data, weather data, geospatial zone data
  • Backend: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
  • Data & storage: PySpark, Delta Lake, PostgreSQL, Redis caching
  • External APIs: Weather data feed
  • Output: React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions
FlowNYC Mobility Forecasting system flow

Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model; PySpark, Delta Lake, PostgreSQL, Redis caching; React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions

  1. input 01Input

    NYC TLC trip data, weather data, geospatial zone data

  2. process 02Backend
    input ->

    Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model

  3. storage 03Data / storage
    backend ->

    PySpark, Delta Lake, PostgreSQL, Redis caching

  4. external 04External APIs
    backend ->

    Weather data feed

  5. output 05Output
    storage ->external ->

    React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions

Routes

  • Input -> Backend
  • Backend -> Data / storage
  • Backend -> External APIs
  • Data / storage -> Output
  • External APIs -> Output

06

Key Technical Decisions

  • Used a two-stage (classification + regression) Poisson hurdle model instead of a single regression model
  • Built a Spark-based medallion (bronze/silver/gold) pipeline
  • Targeted free-tier Oracle Cloud for deployment

07

Implementation

  • Input layer: NYC TLC trip data, weather data, geospatial zone data
  • Core system: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
  • Data layer: PySpark, Delta Lake, PostgreSQL, Redis caching
  • External boundary: Weather data feed
  • User output: React/map-based dashboard consuming a FastAPI REST API for real-time demand predictions

08

What Broke / What Didn't Work

  • Rejected: Single-stage Random Forest regression (baseline). Chosen path: Used a two-stage (classification + regression) Poisson hurdle model instead of a single regression model.
  • Rejected: Single-pass pandas ETL script. Chosen path: Built a Spark-based medallion (bronze/silver/gold) pipeline.
  • Rejected: AWS/GCP paid tiers. Chosen path: Targeted free-tier Oracle Cloud for deployment.
  • Two-stage model adds pipeline complexity over a single regressor but meaningfully improves accuracy on sparse-demand zones
  • Free-tier cloud deployment constrains available compute/memory for serving

09

Results

  • Weighted MAPE improved 18.77% -> 16.52% - Forecast accuracy vs baseline Random Forest - Forecast accuracy vs baseline Random Forest - master-resume
  • RMSE cut by 43%, MAE cut by 53% - Prediction error vs baseline Random Forest - Prediction error vs baseline Random Forest - master-resume
  • 263 zones, 2.3M-row hourly demand dataset - Scale of forecasting system and processed dataset - Scale of forecasting system and processed dataset - master-resume

10

What I'd Change Now

  • Add real-time streaming ingestion instead of batch pipeline
  • Expand to additional cities/regions
  • Incorporate event-based demand signals (concerts, holidays)

11

Stack

  • Python
  • PySpark
  • Apache Spark
  • Delta Lake
  • scikit-learn
  • FastAPI
  • SQL
  • PostgreSQL
  • Redis
  • Docker
  • Oracle Cloud
  • React
  • REST API
  • CI/CD

12

Links

  • Source docs: 2-projects.json
Deep dive prompts

Ask me about the trade-offs.

  • Why this architecture boundary exists: Spark-based medallion (bronze/silver/gold) pipeline; two-stage classification + regression Poisson hurdle model
  • How I evaluated Forecast accuracy vs baseline Random Forest
  • The hardest tradeoff: Two-stage model adds pipeline complexity over a single regressor but meaningfully improves accuracy on sparse-demand zones
  • What I would change next: Add real-time streaming ingestion instead of batch pipeline