Skip to content
All projects
ML Data PlatformFeaturedAirtel X Labs

Data Drift Monitoring Framework for ML Pipelines

A reusable PySpark framework that compares feature distributions between time windows and raises early alerts, across multiple spam ML pipelines.

records per run
600M+
features tested for drift
~70

Context

Multiple spam-detection ML pipelines depend on features whose distributions can shift as network traffic and user behaviour change.

Problem

Without systematic monitoring, a change in upstream data can quietly degrade model inputs before anyone notices. Each pipeline needed the same checks, and running them at this data volume had to be efficient.

My role

I designed and built the framework and integrated it across multiple ML pipelines.

Approach & architecture

  • Compares two date-partitioned feature snapshots and runs statistical tests per feature: KL divergence, PSI, KS test and null-percentage shift.
  • Runs over ~600M records and ~70 features per run.
  • Per-feature tests are parallelised and results are written to a drift results table that drives alerting.
  • Optimised with a deliberate caching strategy, partition tuning, broadcast variables, adaptive query execution and single-pass aggregations.
Feature snapshot T0date-partitionedFeature snapshot T1date-partitionedPer-feature testsKL · PSI · KS · null %Drift resultsresults tableAlertingearly warning
Two date-partitioned feature snapshots feed a parallelised per-feature test engine running KL divergence, PSI, KS test and null-percentage shift. Results go to a drift results table, which drives alerting.

Results

  • One reusable framework adopted across multiple ML pipelines, instead of per-pipeline checks.
  • Early alerts when feature distributions change.

Tech stack

  • PySpark
  • Spark on YARN
  • Statistics
  • MLOps
  • Parquet
  • HDFS

What I'd do next

  • Add per-feature baselines and seasonality-aware thresholds to reduce alert noise.
  • Link drift results to model performance metrics so alerts can be prioritised by impact.