ML Data PlatformFeaturedAirtel X Labs
Data Drift Monitoring Framework for ML Pipelines
A reusable PySpark framework that compares feature distributions between time windows and raises early alerts, across multiple spam ML pipelines.
- records per run
- 600M+
- features tested for drift
- ~70
Context
Multiple spam-detection ML pipelines depend on features whose distributions can shift as network traffic and user behaviour change.
Problem
Without systematic monitoring, a change in upstream data can quietly degrade model inputs before anyone notices. Each pipeline needed the same checks, and running them at this data volume had to be efficient.
My role
I designed and built the framework and integrated it across multiple ML pipelines.
Approach & architecture
- Compares two date-partitioned feature snapshots and runs statistical tests per feature: KL divergence, PSI, KS test and null-percentage shift.
- Runs over ~600M records and ~70 features per run.
- Per-feature tests are parallelised and results are written to a drift results table that drives alerting.
- Optimised with a deliberate caching strategy, partition tuning, broadcast variables, adaptive query execution and single-pass aggregations.
Results
- One reusable framework adopted across multiple ML pipelines, instead of per-pipeline checks.
- Early alerts when feature distributions change.
Tech stack
- PySpark
- Spark on YARN
- Statistics
- MLOps
- Parquet
- HDFS
What I'd do next
- Add per-feature baselines and seasonality-aware thresholds to reduce alert noise.
- Link drift results to model performance metrics so alerts can be prioritised by impact.