Skip to content
All projects
Automation & Reliability

Self-Healing Metadata Watchdog for Hive/Alluxio

Detects stale snapshot tables, auto-remediates Hive/HDFS metadata drift, and alerts only when upstream data is genuinely missing.

Context

A critical Hive snapshot table sometimes went stale, which needed manual ops intervention.

Problem

Staleness had two different causes that needed different responses: Hive metadata out of sync with HDFS, which is safe to fix automatically, or an upstream loader that produced no data, which needs a person.

My role

I built the watchdog.

Approach & architecture

  • Detects when the snapshot table is stale.
  • Distinguishes metadata drift from missing upstream data.
  • For metadata drift, auto-remediates with an Alluxio remount and partition repair.
  • For missing upstream data, alerts only.

Results

  • Replaced manual ops intervention for the safe-to-fix case.

Tech stack

  • Bash
  • Hive
  • Alluxio
  • HDFS
  • Monitoring

What I'd do next

  • Record every remediation so recurring drift can be traced to its root cause.