I’m a senior Data Engineer with 9+ years of experience building enterprise-scale data platforms, scalable ETL pipelines, and cloud-native lakehouse architectures on AWS. I design end-to-end pipelines that consolidate structured, semi-structured, and unstructured data into secure, governed data lakes and analytics warehouses—using tools like AWS Glue, EMR Serverless, Lambda, Redshift Serverless, and orchestration with Step Functions. I’m especially hands-on with PySpark and Spark SQL for high-volume distributed transformations, plus modern lakehouse formats such as Apache Iceberg (and Delta Lake), Parquet, and dbt for reliable, testable analytics layers. I also bring strong governance and compliance experience through AWS Lake Formation and data quality/lineage practices, and I’ve built both batch and real-time streaming pipelines using MSK/Kafka and Kinesis for low-latency processing in regulated domains like banking and healthcare.

Meghana Gude

I’m a senior Data Engineer with 9+ years of experience building enterprise-scale data platforms, scalable ETL pipelines, and cloud-native lakehouse architectures on AWS. I design end-to-end pipelines that consolidate structured, semi-structured, and unstructured data into secure, governed data lakes and analytics warehouses—using tools like AWS Glue, EMR Serverless, Lambda, Redshift Serverless, and orchestration with Step Functions. I’m especially hands-on with PySpark and Spark SQL for high-volume distributed transformations, plus modern lakehouse formats such as Apache Iceberg (and Delta Lake), Parquet, and dbt for reliable, testable analytics layers. I also bring strong governance and compliance experience through AWS Lake Formation and data quality/lineage practices, and I’ve built both batch and real-time streaming pipelines using MSK/Kafka and Kinesis for low-latency processing in regulated domains like banking and healthcare.

Available to hire

I’m a senior Data Engineer with 9+ years of experience building enterprise-scale data platforms, scalable ETL pipelines, and cloud-native lakehouse architectures on AWS. I design end-to-end pipelines that consolidate structured, semi-structured, and unstructured data into secure, governed data lakes and analytics warehouses—using tools like AWS Glue, EMR Serverless, Lambda, Redshift Serverless, and orchestration with Step Functions.

I’m especially hands-on with PySpark and Spark SQL for high-volume distributed transformations, plus modern lakehouse formats such as Apache Iceberg (and Delta Lake), Parquet, and dbt for reliable, testable analytics layers. I also bring strong governance and compliance experience through AWS Lake Formation and data quality/lineage practices, and I’ve built both batch and real-time streaming pipelines using MSK/Kafka and Kinesis for low-latency processing in regulated domains like banking and healthcare.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Beginner
See more

Language

English
Advanced

Work Experience

Senior Data Engineer at Truist Financial
February 1, 2025 - Present
Designed and deployed serverless ETL pipelines on AWS (Glue 4.0, Lambda, Step Functions) to consolidate banking data into Redshift Serverless. Built a scalable S3 data lake using Apache Iceberg with partitioned Parquet, enabling ACID transactions, schema evolution, and time-travel for regulatory reporting. Implemented streaming ingestion with Kafka (AWS MSK) and Spark Structured Streaming, persisting results as Parquet/Iceberg for downstream analytics. Applied enterprise governance with AWS Lake Formation for column- and row-level fine-grained access control and cataloging. Automated infrastructure provisioning with Terraform/CloudFormation and implemented CI/CD with GitHub Actions for reproducible deployments. Standardized transformations using dbt Core within Redshift, and added operational reliability via CloudWatch monitoring/alarms and audit logging.
Senior Data Engineer at GE HealthCare
June 1, 2023 - January 31, 2025
Built HIPAA-compliant healthcare ETL and data lake pipelines integrating patient, clinical, imaging, and operational data into AWS S3. Developed AWS Glue crawlers and ETL jobs for schema discovery and metadata management, and used PySpark/Spark to transform and validate large datasets into optimized Parquet/JSON/CSV formats. Implemented near real-time ingestion and processing using Kafka for patient monitoring and telemetry streams, with event-driven workflows via Lambda and alerting through CloudWatch/SNS. Loaded curated datasets into Redshift for reporting and dashboards. Implemented secure infrastructure with VPC and enforced governance using AWS Lake Formation with fine-grained permissions. Automated delivery with Terraform, Jenkins CI/CD, and version-controlled repositories; used DynamoDB to track job execution state and checkpoints.
Senior Data Engineer at Target Corporation
August 1, 2021 - May 31, 2023
Supported Azure-based production data platforms for retail analytics across 1,900+ store and digital environments. Led migration from on-prem SQL Server to Azure Synapse Analytics, redesigning schemas for performance and cost improvements. Implemented IaC using ARM templates and Terraform to provision Azure Data Lake Gen2, Databricks, Synapse, and Data Factory components. Built streaming ingestion with Event Hubs/Kafka (HDInsight) to land POS data into the lake for near real-time visibility. Optimized PySpark processing on Databricks (partitioning, broadcast joins) and designed Synapse data models using distribution/partition strategies. Established CI/CD with Azure DevOps for automated deployment of Data Factory pipelines, Databricks notebooks, and Functions. Worked with Azure ML for model integration and used Delta Lake features (ACID, versioning, time-travel) to power reliable curated datasets.
Senior Data Engineer at RBC (Royal Bank of Canada)
December 1, 2019 - July 31, 2021
Developed event-driven Azure pipelines that ingest customer/account/transaction data from Synapse, transform it with Databricks (PySpark), and write processed outputs into Neo4j for customer relationship mapping and fraud network analysis. Orchestrated workflows using Azure Data Factory and Logic Apps, with audit logging via Azure SQL Database and Cosmos DB and job triggering through Azure Functions. Built Databricks notebooks for querying and transformation, using Apache Iceberg on Azure Data Lake Gen2 and creating optimized tables for regulatory/risk reporting in Synapse. Provisioned Azure infrastructure (Data Lake Gen2, Databricks, IAM, Service Bus, Cosmos DB, Functions) via Terraform and automated CI/CD with Azure DevOps. Created reusable Python modules packaged as wheels and deployed them to Databricks clusters to standardize transformation logic.
Data Analyst at Zensar Technologies
September 1, 2016 - October 31, 2019
Migrated on-prem SSIS ETL workflows to GCP using Cloud Dataflow (Apache Beam) for multi-million record processing across heterogeneous sources (SQL Server, Oracle, flat files). Built streaming ingestion pipelines with Pub/Sub and Dataflow Streaming to improve scalability and reduce operational overhead. Designed a 3-zone data lake on GCS (raw/curated/consumption) integrated with BigQuery, and developed PySpark jobs on Dataproc for distributed file processing. Implemented BigQuery security controls (row/column-level) and improved analytics performance using partitioning, clustering, materialized views, and query rewrites. Created reporting dashboards in Power BI and Tableau connected to BigQuery and modeled star/snowflake schemas for analytical workloads; performed ETL validation, performance tuning, and data modeling across systems.

Education

Bachelor of Technology – Computer Science at Stanley College of Engineering and Technology for Women
January 1, 2016 - August 28, 2026
Bachelor of Technology – Computer Science at Stanley College of Engineering and Technology for Women
January 1, 2012 - January 1, 2016

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Retail, Professional Services, Other