Skip to content
All projects
Data Pipelines & Migration

Oracle Hierarchical Logic → PySpark

Re-implemented an Oracle PL/SQL account-hierarchy function (CONNECT BY) as a distributed, iterative PySpark job with checkpointing.

Context

Account-hierarchy logic lived in an Oracle PL/SQL function built on CONNECT BY.

Problem

Row-by-row hierarchical logic does not carry over to a distributed engine as written.

My role

I re-implemented the function in PySpark.

Approach & architecture

  • Rewrote the hierarchy traversal as an iterative PySpark job.
  • Replaced row-by-row logic with set-based processing.
  • Used checkpointing to keep the lineage of each iteration manageable.

Results

  • Hierarchical account logic now runs as a distributed job at scale.

Tech stack

  • PySpark
  • Oracle PL/SQL
  • Spark SQL

What I'd do next

  • Add cycle detection and depth limits as explicit, tested safeguards.