Skip to content
All projects
Data Pipelines & MigrationCapgemini

Oracle Data Lake Migration

Led the migration of Oracle source systems into a Hadoop data lake, with generated DMLs, test data and table-by-table validation.

Context

An enterprise moving data from Oracle source systems into a Hadoop-based data lake.

Problem

Many tables had to be migrated, and every one needed a matching DML, test data and proof that it arrived intact. Doing that by hand was slow and error-prone.

My role

I led the migration and built its supporting automation.

Approach & architecture

  • Extracted data from Oracle source systems into Hadoop-based data lakes.
  • Built utilities to auto-generate DMLs and test data.
  • Designed batch validation that analysed every migrated table and produced reconciliation reports.

Results

  • Less manual effort and fewer errors in DML creation and test-data generation.
  • Every migrated table covered by a validation report.

Tech stack

  • Ab Initio
  • Oracle
  • Hadoop
  • Hive
  • Unix shell

What I'd do next

  • Package the DML and validation generators as a reusable migration toolkit.