Senior Cloud Data Engineer and Lead Data Engineer with 14+ years of experience designing and optimizing modern cloud data platforms and enterprise data warehouses across banking, insurance, healthcare, government, and retail. Deep hands-on expertise in Data Warehousing, Data Lakes, Delta Live Tables (DLT), ETL/ELT, CDC, and large-scale migrations from Teradata/Oracle/SQL Server to Snowflake, AWS Redshift, Azure, and Databricks.

PRABAKAR MUNUSAMY

Senior Cloud Data Engineer and Lead Data Engineer with 14+ years of experience designing and optimizing modern cloud data platforms and enterprise data warehouses across banking, insurance, healthcare, government, and retail. Deep hands-on expertise in Data Warehousing, Data Lakes, Delta Live Tables (DLT), ETL/ELT, CDC, and large-scale migrations from Teradata/Oracle/SQL Server to Snowflake, AWS Redshift, Azure, and Databricks.

Available to hire

Senior Cloud Data Engineer and Lead Data Engineer with 14+ years of experience designing and optimizing modern cloud data platforms and enterprise data warehouses across banking, insurance, healthcare, government, and retail. Deep hands-on expertise in Data Warehousing, Data Lakes, Delta Live Tables (DLT), ETL/ELT, CDC, and large-scale migrations from Teradata/Oracle/SQL Server to Snowflake, AWS Redshift, Azure, and Databricks.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more

Language

Work Experience

Lead Data Engineer & Technical Data Lead at NCS AUSTRALIA
June 1, 2026 - Present
Led end-to-end enterprise data warehouse, data lake, and Delta Live Tables (DLT) implementation for National Australia Bank (NAB) ADA projects. Built Databricks medallion-architecture pipelines (Bronze/Stage/Silver/Gold/Serve) using Python and PySpark, and implemented Databricks SQL transformations across layers. Developed DLT pipelines for reliable ingestion, transformations, data quality enforcement, and movement across layers. Created Serve layer datasets and generated outbound .dat/.csv files for downstream consumers. Drove performance tuning, operational stability, troubleshooting, and production readiness for high-volume banking data from multiple NAB source systems, collaborating with stakeholders and technical teams to translate regulatory and business requirements into reusable, production-grade solutions.
Lead Data Engineer & Technical Data Lead at CAPGEMINI AUSTRALIA
June 1, 2023 - June 1, 2026
Architected and delivered modern cloud data platforms for the Department of Justice and Community Safety using Airflow ingestion into Databricks transforms, dbt models, and deployment to Snowflake and AWS Redshift, as well as ADF ingestion into Databricks with dbt targeting Azure Synapse. Built Airflow DAGs and ADF pipelines in Python for automated ingestion from government sources into Databricks, including scheduling, dependencies, retries, error handling, and SLA monitoring. Implemented Databricks (PySpark/SQL) transformation jobs for cleansing, standardization, and SCD logic. Developed dbt staging/intermediate/mart layers with automated tests and documentation to reduce transformation errors. Optimized warehouse performance (Snowflake distribution/sort keys, Redshift distribution/sort keys) and implemented CI/CD automation for DDL/DML, Airflow DAGs, Databricks notebooks, and dbt projects across SIT/PREPROD/PROD, improving deployment speed and operational consistency while managing
Senior Cloud Data Engineer & Technical Data Lead at CAPGEMINI AUSTRALIA
June 1, 2023 - Present
Architected and delivered modern cloud data platform pipelines for Department of Justice and Community Safety using orchestration and transformation chain patterns: Airflow ingestion → Databricks transforms → dbt models targeting Snowflake/Redshift, and ADF ingestion → Databricks transforms → dbt models targeting Azure Synapse. Built production-grade Airflow DAGs and ADF pipelines in Python for automated ingestion, scheduling, retries, dependency management, SLA monitoring, and robust error handling. Developed Databricks PySpark/SQL transformations for cleansing, standardization, SCD handling, and business rule implementation. Implemented dbt staging/intermediate/mart layers with automated tests and lineage. Designed Delta Lake incremental processing for ACID-compliant synchronization and built performance-optimized SQL (cluster/distribution/sort strategies), improving query performance. Established CI/CD using Jenkins across SIT/PREPROD/PROD for deploying DDL/DML, Airflow DAG
Lead Data Engineer & Technical Data Lead at COGNIZANT AUSTRALIA
May 1, 2022 - June 1, 2023
Delivered a modern cloud data platform enabling a 360-degree customer view for National Australia Bank Customer Master project using Airflow → Databricks → dbt → Snowflake/Redshift and ADF → Databricks → dbt → Azure Synapse. Designed and implemented Airflow DAGs and ADF pipelines for ingestion from on-premises and cloud sources into Databricks with scheduling, retries, dependencies, and SLA monitoring. Built Databricks transformations for conformance and business rule implementation prior to dbt modeling. Developed dbt staging/intermediate/mart layers with automated tests, documentation, and lineage; strengthened data quality via profiling and cleansing in Databricks and dbt. Migrated 5+ years of customer history to Snowflake/Redshift with zero data loss. Tuned Snowflake clustering, Redshift distribution/sort keys, and Synapse indexing to improve query performance. Established CI/CD using Git and Jenkins for deployments across DEV/SIT/PROD and received project excellence re
Senior Cloud Data Engineer & Technical Data Lead at COGNIZANT AUSTRALIA
May 1, 2022 - June 1, 2023
Delivered a modern cloud data platform for National Australia Bank (Customer Master) using Airflow → Databricks → dbt → Snowflake/Redshift and ADF → Databricks → dbt → Azure Synapse patterns to enable 360-degree customer views. Architected Airflow DAGs and ADF pipelines for ingestion from on-prem/cloud sources into Databricks with scheduling, dependencies, retries, and SLA monitoring. Developed Databricks transformations (PySpark/SQL) for cleansing, standardization, and conformance before dbt modeling. Built dbt models (staging/intermediate/mart) with tests, documentation, and data lineage across multiple warehouses. Designed dimensional schemas (Star/Snowflake) and implemented data profiling and data quality processes. Established historical migration strategy for 5+ years of customer data to Snowflake and Redshift with zero data loss. Optimized clustering keys, distribution/sort keys, and indexing to improve query performance. Implemented Git/Jenkins CI/CD for deploying
Senior Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
July 1, 2020 - May 1, 2022
Led Oracle and Netezza to Snowflake and AWS Redshift migration for Suncorp Claim Centre Upgrade using Airflow → Python → dbt → Snowflake/Redshift architecture. Migrated 500+ database objects with zero data loss, converting Oracle/Netezza SQL to Snowflake/Redshift-compatible SQL with comprehensive testing and data validation. Used bulk loading and AWS DMS to move data efficiently with minimal downtime. Designed ETL/ELT strategies for insurance claims and policy/customer data from multiple source systems. Built Airflow DAGs for full/incremental loading from sources and S3 into Snowflake/Redshift with SLA monitoring and robust dependency/retry/error handling. Implemented Python transformation steps in Airflow, then produced dbt staging/intermediate/mart models with automated tests and lineage. Developed EMR scripts for processing CSV/JSON/Parquet from S3 with validation and exception handling. Set up Git repositories and Jenkins CI/CD to reduce manual deployments by 70%.
Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
December 1, 2019 - July 1, 2020
Executed Teradata to Snowflake cloud migration for Discover Financial Services investment banking data warehouse using Teradata → Airflow ingestion → Python → Snowflake pattern. Migrated 300+ tables and 1000+ stored procedures with SQL conversion, performance optimization, and end-to-end functional validation. Developed Airflow DAGs for full/incremental loads from Teradata and S3 into Snowflake with workflow dependencies, retries, SLA monitoring, and rollback procedures via CI/CD. Implemented CDC within Airflow-based Python pipelines to enable incremental synchronization. Designed Snowflake fact/dimension models with optimized clustering keys, micro-partitions, and compression. Performed data profiling/cleansing/mining in transformation layers and set up Git/Jenkins CI/CD for SIT/PROD deployment of Airflow DAGs, Python scripts, and Snowflake DDLs. Received Customer Focus Award for migration excellence.
Wellmark Blue Cross Blue Shield Projects — Lead Data Engineer / Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
October 1, 2017 - December 1, 2019
Led phased Teradata to AWS Redshift migration for healthcare insurance, dental coverage, and pharmaceutical data warehouse using Teradata → Airflow ingestion → Python → Redshift. Replaced legacy Teradata ETL and Informatica PowerCenter with cloud-native Airflow-orchestrated Python pipelines. Built Airflow DAGs for full and incremental extraction from Teradata into Redshift staging and dimensional layers. Migrated PowerCenter ETL sessions/mappings/parameterization into Python-based Airflow pipelines while preserving integration logic. Implemented CDC within Airflow pipelines for incremental ETL across healthcare source systems to maintain data freshness and accuracy. Converted Teradata bulk load utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT) to Redshift COPY commands for improved bulk load performance. Developed optimized Redshift SQL queries/stored procedures using distribution and sort keys.
Wellmark Blue Cross Blue Shield Projects (Cloud Data Engineer) at COGNIZANT TECHNOLOGY SOLUTIONS
October 1, 2017 - December 1, 2019
Led phased migration from Teradata to AWS Redshift for healthcare insurance, dental, and pharmaceutical data warehouse using Teradata → Airflow ingestion → Python → Redshift. Replaced legacy Teradata ETL and Informatica PowerCenter with cloud-native Python pipelines orchestrated by Airflow. Built Airflow DAGs for full and incremental extractions including dependencies, retries, and error handling. Migrated Informatica PowerCenter ETL workflows (sessions/mappings/parameter files) into Python-based Airflow pipelines preserving integration logic. Implemented CDC processes for incremental ETL and data freshness. Converted Teradata bulk loading utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT) into Python and Redshift COPY-based loading for performance. Tuned Redshift SQL using distribution keys, sort keys, and vacuum/analyze strategies, with production job monitoring and alerting.
Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
October 1, 2017 - December 1, 2019
Led phased migration from Teradata to AWS Redshift for Wellmark Blue Cross Blue Shield healthcare insurance and related data using Teradata → Airflow ingestion → Python → Redshift pattern. Replaced legacy Teradata ETL and Informatica PowerCenter with cloud-native Airflow-orchestrated Python pipelines. Built Airflow DAGs for full and incremental extractions into Redshift staging and dimensional layers with scheduling, dependencies, retries, and error handling. Migrated Informatica PowerCenter ETL workflows (sessions/mappings/parameter files) into Python-based Airflow pipelines, preserving integration logic. Implemented CDC in Airflow Python pipelines for incremental ETL and data freshness. Converted Teradata loading utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT) into Python and Redshift COPY-based bulk loading. Developed complex Redshift SQL and stored procedures using distribution/sort keys and vacuum/analyze tuning. Implemented automated monitoring, error handling, a
Senior Data Engineer at INFOSYS LIMITED
March 1, 2015 - October 1, 2017
Built and supported Informatica PowerCenter → SQL → Teradata architecture for RBS FATCA regulatory reporting (IRS/HMRC). Designed and implemented Informatica PowerCenter workflows, sessions, mappings, and transformations to populate on-prem Teradata with complex business rules. Developed complex Teradata SQL queries, stored procedures, macros, and views as the SQL transformation layer. Loaded data into Teradata staging and tactical tables using Teradata utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT). Implemented CDC processes in Informatica for incremental ETL and audit trail consistency. Created Unix/Windows scripts for automated file transfers, batch scheduling, and FTP/SFTP operations. Provided production support with automated error handling and notifications, maintaining 99.9% availability.
Data Engineer at TATA CONSULTANCY SERVICES LIMITED
December 1, 2011 - March 1, 2015
Developed end-to-end ETL strategies using Informatica PowerCenter → SQL → Teradata for a Teradata warehouse supporting investment deposit, funding accounts, and corporate services. Designed Informatica PowerCenter workflows, sessions, mappings, transformations, and parameter files across Teradata and Oracle source systems. Implemented Teradata SQL (queries, stored procedures, macros, views) for business logic and performance optimization. Loaded data into Teradata staging/tactical tables using high-volume utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT). Implemented CDC in Informatica for incremental ETL and consistency across Teradata systems. Automated file transfers, batch scheduling, and FTP/SFTP operations using Unix/Windows shell scripts. Provided production support with automated error handling and notification systems, reducing production incidents by 45%.

Education

Post Graduate Diploma in Data Mining at Annamalai University
January 1, 2019 - January 1, 2019
Bachelor of Engineering in Electronics and Communication Engineering at Alagappa Chettiar College of Engineering and Technology (Affiliated to Anna University, Tamil Nadu)
January 1, 2011 - January 1, 2011
Post Graduate Diploma in Data Mining at Annamalai University
January 1, 2019 - January 1, 2019
Bachelor of Engineering in Electronics and Communication Engineering at Alagappa Chettiar College of Engineering and Technology (Affiliated to Anna University, Tamil Nadu, India)
January 1, 2011 - January 1, 2011
Post Graduate Diploma in Data Mining at Annamalai University
January 1, 2019 - January 1, 2019
Bachelor of Engineering in Electronics and Communication Engineering at Alagappa Chettiar College of Engineering and Technology (Affiliated to Anna University, Tamil Nadu, India)
January 1, 2011 - January 1, 2011

Qualifications

Databricks Data Engineer Professional
July 1, 2023 - July 13, 2026
Databricks Accredited Lakehouse Fundamentals
August 1, 2022 - July 13, 2026
Apache Spark 3 Databricks Certified Associate Developer
August 1, 2022 - July 13, 2026
Mastering Amazon Redshift 2021 Development and Administration
December 1, 2021 - July 13, 2026
Oracle Database 11g Administrator Certified Professional
January 1, 2015 - July 13, 2026
Oracle Database 11g Administrator Certified Associate
October 1, 2014 - July 13, 2026
Oracle Database 11g Performance Tuning Certified Expert
February 1, 2014 - July 13, 2026
Oracle Database SQL Certified Expert
November 1, 2013 - July 13, 2026
Oracle PL/SQL Developer Certified Associate
June 1, 2014 - July 13, 2026
Teradata Basics
April 1, 2014 - July 13, 2026
Databricks Data Engineer Professional
July 1, 2023 - July 13, 2026
Databricks Accredited Lakehouse Fundamentals
August 1, 2022 - July 13, 2026
Apache Spark 3 Databricks Certified Associate Developer
August 1, 2022 - July 13, 2026
Mastering Amazon Redshift 2021 Development and Administration
December 1, 2021 - July 13, 2026
Oracle Database 11g Administrator Certified Professional
January 1, 2015 - July 13, 2026
Oracle Database 11g Administrator Certified Associate
October 1, 2014 - July 13, 2026
Oracle Database 11g Performance Tuning Certified Expert
February 1, 2014 - July 13, 2026
Oracle Database SQL Certified Expert
November 1, 2013 - July 13, 2026
Oracle PL/SQL Developer Certified Associate
June 1, 2014 - July 13, 2026
Teradata Basics
April 1, 2014 - July 13, 2026
Databricks Data Engineer Professional
July 1, 2023 - August 20, 2026
Databricks Accredited Lakehouse Fundamentals
August 1, 2022 - August 20, 2026
Apache Spark 3 Databricks Certified Associate Developer
August 1, 2022 - August 20, 2026
Mastering Amazon Redshift 2021 Development and Administration
December 1, 2021 - August 20, 2026
Oracle Database 11g Administrator Certified Professional
January 1, 2015 - August 20, 2026
Oracle Database 11g Administrator Certified Associate
October 1, 2014 - August 20, 2026
Oracle Database 11g Performance Tuning Certified Expert
February 1, 2014 - August 20, 2026
Oracle Database SQL Certified Expert
November 1, 2013 - August 20, 2026
Oracle PL/SQL Developer Certified Associate
June 1, 2014 - August 20, 2026
Teradata Basics
April 1, 2014 - August 20, 2026

Industry Experience

Financial Services, Government, Healthcare, Retail, Professional Services, Other