Data Pipelines & MigrationCapgemini
Oracle Data Lake Migration
Led the migration of Oracle source systems into a Hadoop data lake, with generated DMLs, test data and table-by-table validation.
Context
An enterprise moving data from Oracle source systems into a Hadoop-based data lake.
Problem
Many tables had to be migrated, and every one needed a matching DML, test data and proof that it arrived intact. Doing that by hand was slow and error-prone.
My role
I led the migration and built its supporting automation.
Approach & architecture
- Extracted data from Oracle source systems into Hadoop-based data lakes.
- Built utilities to auto-generate DMLs and test data.
- Designed batch validation that analysed every migrated table and produced reconciliation reports.
Results
- Less manual effort and fewer errors in DML creation and test-data generation.
- Every migrated table covered by a validation report.
Tech stack
- Ab Initio
- Oracle
- Hadoop
- Hive
- Unix shell
What I'd do next
- Package the DML and validation generators as a reusable migration toolkit.