Automation & Reliability
Self-Healing Metadata Watchdog for Hive/Alluxio
Detects stale snapshot tables, auto-remediates Hive/HDFS metadata drift, and alerts only when upstream data is genuinely missing.
Context
A critical Hive snapshot table sometimes went stale, which needed manual ops intervention.
Problem
Staleness had two different causes that needed different responses: Hive metadata out of sync with HDFS, which is safe to fix automatically, or an upstream loader that produced no data, which needs a person.
My role
I built the watchdog.
Approach & architecture
- Detects when the snapshot table is stale.
- Distinguishes metadata drift from missing upstream data.
- For metadata drift, auto-remediates with an Alluxio remount and partition repair.
- For missing upstream data, alerts only.
Results
- Replaced manual ops intervention for the safe-to-fix case.
Tech stack
- Bash
- Hive
- Alluxio
- HDFS
- Monitoring
What I'd do next
- Record every remediation so recurring drift can be traced to its root cause.