Data Pipelines & Migration
Oracle Hierarchical Logic → PySpark
Re-implemented an Oracle PL/SQL account-hierarchy function (CONNECT BY) as a distributed, iterative PySpark job with checkpointing.
Context
Account-hierarchy logic lived in an Oracle PL/SQL function built on CONNECT BY.
Problem
Row-by-row hierarchical logic does not carry over to a distributed engine as written.
My role
I re-implemented the function in PySpark.
Approach & architecture
- Rewrote the hierarchy traversal as an iterative PySpark job.
- Replaced row-by-row logic with set-based processing.
- Used checkpointing to keep the lineage of each iteration manageable.
Results
- Hierarchical account logic now runs as a distributed job at scale.
Tech stack
- PySpark
- Oracle PL/SQL
- Spark SQL
What I'd do next
- Add cycle detection and depth limits as explicit, tested safeguards.