Senior Cloud Data Engineer and Technical Data Lead with 14+ years of experience designing and optimizing enterprise cloud data platforms and data warehouses across banking, insurance, healthcare, government, and retail. Deep hands-on expertise in ETL/ELT architectures, large-scale migrations (Teradata/Oracle/SQL Server to Snowflake/Redshift/Synapse/Databricks), and performance tuning. Proven leader delivering end-to-end data pipeline architectures using Airflow and ADF with Databricks transforms and dbt modeling, including CDC, Delta Lake, CI/CD, and secure multi-cloud production deployments. Leverages AI models (GPT, Gemini, Claude) to accelerate development, testing, and deployment while mentoring teams and driving measurable improvements in reliability, speed, and data quality.

PRABAKAR MUNUSAMY

Senior Cloud Data Engineer and Technical Data Lead with 14+ years of experience designing and optimizing enterprise cloud data platforms and data warehouses across banking, insurance, healthcare, government, and retail. Deep hands-on expertise in ETL/ELT architectures, large-scale migrations (Teradata/Oracle/SQL Server to Snowflake/Redshift/Synapse/Databricks), and performance tuning. Proven leader delivering end-to-end data pipeline architectures using Airflow and ADF with Databricks transforms and dbt modeling, including CDC, Delta Lake, CI/CD, and secure multi-cloud production deployments. Leverages AI models (GPT, Gemini, Claude) to accelerate development, testing, and deployment while mentoring teams and driving measurable improvements in reliability, speed, and data quality.

Available to hire

Senior Cloud Data Engineer and Technical Data Lead with 14+ years of experience designing and optimizing enterprise cloud data platforms and data warehouses across banking, insurance, healthcare, government, and retail. Deep hands-on expertise in ETL/ELT architectures, large-scale migrations (Teradata/Oracle/SQL Server to Snowflake/Redshift/Synapse/Databricks), and performance tuning.

Proven leader delivering end-to-end data pipeline architectures using Airflow and ADF with Databricks transforms and dbt modeling, including CDC, Delta Lake, CI/CD, and secure multi-cloud production deployments. Leverages AI models (GPT, Gemini, Claude) to accelerate development, testing, and deployment while mentoring teams and driving measurable improvements in reliability, speed, and data quality.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more

Language

Work Experience

Senior Cloud Data Engineer & Technical Data Lead at CAPGEMINI AUSTRALIA
June 1, 2023 - Present
Architected and delivered modern cloud data platform pipelines for Department of Justice and Community Safety using orchestration and transformation chain patterns: Airflow ingestion → Databricks transforms → dbt models targeting Snowflake/Redshift, and ADF ingestion → Databricks transforms → dbt models targeting Azure Synapse. Built production-grade Airflow DAGs and ADF pipelines in Python for automated ingestion, scheduling, retries, dependency management, SLA monitoring, and robust error handling. Developed Databricks PySpark/SQL transformations for cleansing, standardization, SCD handling, and business rule implementation. Implemented dbt staging/intermediate/mart layers with automated tests and lineage. Designed Delta Lake incremental processing for ACID-compliant synchronization and built performance-optimized SQL (cluster/distribution/sort strategies), improving query performance. Established CI/CD using Jenkins across SIT/PREPROD/PROD for deploying DDL/DML, Airflow DAG
Senior Cloud Data Engineer & Technical Data Lead at COGNIZANT AUSTRALIA
May 1, 2022 - June 1, 2023
Delivered a modern cloud data platform for National Australia Bank (Customer Master) using Airflow → Databricks → dbt → Snowflake/Redshift and ADF → Databricks → dbt → Azure Synapse patterns to enable 360-degree customer views. Architected Airflow DAGs and ADF pipelines for ingestion from on-prem/cloud sources into Databricks with scheduling, dependencies, retries, and SLA monitoring. Developed Databricks transformations (PySpark/SQL) for cleansing, standardization, and conformance before dbt modeling. Built dbt models (staging/intermediate/mart) with tests, documentation, and data lineage across multiple warehouses. Designed dimensional schemas (Star/Snowflake) and implemented data profiling and data quality processes. Established historical migration strategy for 5+ years of customer data to Snowflake and Redshift with zero data loss. Optimized clustering keys, distribution/sort keys, and indexing to improve query performance. Implemented Git/Jenkins CI/CD for deploying
Senior Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
July 1, 2020 - May 1, 2022
Led Oracle and Netezza to Snowflake and AWS Redshift migration for Suncorp Claim Centre Upgrade using Airflow → Python → dbt → Snowflake/Redshift architecture. Migrated 500+ database objects with zero data loss, converting Oracle/Netezza SQL to Snowflake/Redshift-compatible SQL with comprehensive testing and data validation. Used bulk loading and AWS DMS to move data efficiently with minimal downtime. Designed ETL/ELT strategies for insurance claims and policy/customer data from multiple source systems. Built Airflow DAGs for full/incremental loading from sources and S3 into Snowflake/Redshift with SLA monitoring and robust dependency/retry/error handling. Implemented Python transformation steps in Airflow, then produced dbt staging/intermediate/mart models with automated tests and lineage. Developed EMR scripts for processing CSV/JSON/Parquet from S3 with validation and exception handling. Set up Git repositories and Jenkins CI/CD to reduce manual deployments by 70%.
Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
December 1, 2019 - July 1, 2020
Executed Teradata to Snowflake cloud migration for Discover Financial Services investment banking data warehouse using Teradata → Airflow ingestion → Python → Snowflake pattern. Migrated 300+ tables and 1000+ stored procedures with SQL conversion, performance optimization, and end-to-end functional validation. Developed Airflow DAGs for full/incremental loads from Teradata and S3 into Snowflake with workflow dependencies, retries, SLA monitoring, and rollback procedures via CI/CD. Implemented CDC within Airflow-based Python pipelines to enable incremental synchronization. Designed Snowflake fact/dimension models with optimized clustering keys, micro-partitions, and compression. Performed data profiling/cleansing/mining in transformation layers and set up Git/Jenkins CI/CD for SIT/PROD deployment of Airflow DAGs, Python scripts, and Snowflake DDLs. Received Customer Focus Award for migration excellence.
Wellmark Blue Cross Blue Shield Projects (Cloud Data Engineer) at COGNIZANT TECHNOLOGY SOLUTIONS
October 1, 2017 - December 1, 2019
Led phased migration from Teradata to AWS Redshift for healthcare insurance, dental, and pharmaceutical data warehouse using Teradata → Airflow ingestion → Python → Redshift. Replaced legacy Teradata ETL and Informatica PowerCenter with cloud-native Python pipelines orchestrated by Airflow. Built Airflow DAGs for full and incremental extractions including dependencies, retries, and error handling. Migrated Informatica PowerCenter ETL workflows (sessions/mappings/parameter files) into Python-based Airflow pipelines preserving integration logic. Implemented CDC processes for incremental ETL and data freshness. Converted Teradata bulk loading utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT) into Python and Redshift COPY-based loading for performance. Tuned Redshift SQL using distribution keys, sort keys, and vacuum/analyze strategies, with production job monitoring and alerting.
Cloud Data Engineer at COGNIZANT TECHNOLOGY SOLUTIONS
October 1, 2017 - December 1, 2019
Led phased migration from Teradata to AWS Redshift for Wellmark Blue Cross Blue Shield healthcare insurance and related data using Teradata → Airflow ingestion → Python → Redshift pattern. Replaced legacy Teradata ETL and Informatica PowerCenter with cloud-native Airflow-orchestrated Python pipelines. Built Airflow DAGs for full and incremental extractions into Redshift staging and dimensional layers with scheduling, dependencies, retries, and error handling. Migrated Informatica PowerCenter ETL workflows (sessions/mappings/parameter files) into Python-based Airflow pipelines, preserving integration logic. Implemented CDC in Airflow Python pipelines for incremental ETL and data freshness. Converted Teradata loading utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT) into Python and Redshift COPY-based bulk loading. Developed complex Redshift SQL and stored procedures using distribution/sort keys and vacuum/analyze tuning. Implemented automated monitoring, error handling, a
Senior Data Engineer at INFOSYS LIMITED
March 1, 2015 - October 1, 2017
Built and supported Informatica PowerCenter → SQL → Teradata architecture for RBS FATCA regulatory reporting (IRS/HMRC). Designed and implemented Informatica PowerCenter workflows, sessions, mappings, and transformations to populate on-prem Teradata with complex business rules. Developed complex Teradata SQL queries, stored procedures, macros, and views as the SQL transformation layer. Loaded data into Teradata staging and tactical tables using Teradata utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT). Implemented CDC processes in Informatica for incremental ETL and audit trail consistency. Created Unix/Windows scripts for automated file transfers, batch scheduling, and FTP/SFTP operations. Provided production support with automated error handling and notifications, maintaining 99.9% availability.
Data Engineer at TATA CONSULTANCY SERVICES LIMITED
December 1, 2011 - March 1, 2015
Developed end-to-end ETL strategies using Informatica PowerCenter → SQL → Teradata for a Teradata warehouse supporting investment deposit, funding accounts, and corporate services. Designed Informatica PowerCenter workflows, sessions, mappings, transformations, and parameter files across Teradata and Oracle source systems. Implemented Teradata SQL (queries, stored procedures, macros, views) for business logic and performance optimization. Loaded data into Teradata staging/tactical tables using high-volume utilities (BTEQ/TPUMP/MULTILOAD/FASTLOAD/FASTEXPORT/TPT). Implemented CDC in Informatica for incremental ETL and consistency across Teradata systems. Automated file transfers, batch scheduling, and FTP/SFTP operations using Unix/Windows shell scripts. Provided production support with automated error handling and notification systems, reducing production incidents by 45%.

Education

Post Graduate Diploma in Data Mining at Annamalai University
January 1, 2019 - January 1, 2019
Bachelor of Engineering in Electronics and Communication Engineering at Alagappa Chettiar College of Engineering and Technology (Affiliated to Anna University, Tamil Nadu)
January 1, 2011 - January 1, 2011
Post Graduate Diploma in Data Mining at Annamalai University
January 1, 2019 - January 1, 2019
Bachelor of Engineering in Electronics and Communication Engineering at Alagappa Chettiar College of Engineering and Technology (Affiliated to Anna University, Tamil Nadu, India)
January 1, 2011 - January 1, 2011

Qualifications

Databricks Data Engineer Professional
July 1, 2023 - July 13, 2026
Databricks Accredited Lakehouse Fundamentals
August 1, 2022 - July 13, 2026
Apache Spark 3 Databricks Certified Associate Developer
August 1, 2022 - July 13, 2026
Mastering Amazon Redshift 2021 Development and Administration
December 1, 2021 - July 13, 2026
Oracle Database 11g Administrator Certified Professional
January 1, 2015 - July 13, 2026
Oracle Database 11g Administrator Certified Associate
October 1, 2014 - July 13, 2026
Oracle Database 11g Performance Tuning Certified Expert
February 1, 2014 - July 13, 2026
Oracle Database SQL Certified Expert
November 1, 2013 - July 13, 2026
Oracle PL/SQL Developer Certified Associate
June 1, 2014 - July 13, 2026
Teradata Basics
April 1, 2014 - July 13, 2026
Databricks Data Engineer Professional
July 1, 2023 - July 13, 2026
Databricks Accredited Lakehouse Fundamentals
August 1, 2022 - July 13, 2026
Apache Spark 3 Databricks Certified Associate Developer
August 1, 2022 - July 13, 2026
Mastering Amazon Redshift 2021 Development and Administration
December 1, 2021 - July 13, 2026
Oracle Database 11g Administrator Certified Professional
January 1, 2015 - July 13, 2026
Oracle Database 11g Administrator Certified Associate
October 1, 2014 - July 13, 2026
Oracle Database 11g Performance Tuning Certified Expert
February 1, 2014 - July 13, 2026
Oracle Database SQL Certified Expert
November 1, 2013 - July 13, 2026
Oracle PL/SQL Developer Certified Associate
June 1, 2014 - July 13, 2026
Teradata Basics
April 1, 2014 - July 13, 2026

Industry Experience

Financial Services, Government, Healthcare, Retail, Professional Services