Senior Data Engineer with 10+ years of experience designing and maintaining large-scale ETL/ELT pipelines, streaming architectures, and modern data platforms (data lakes/warehouses). Proficient in SQL, Python, Apache Spark, Databricks, Snowflake, and cloud ecosystems across AWS/Azure/GCP. Experienced in building real-time ingestion and processing with Kafka/Kinesis/Event Hubs/Pub/Sub, orchestrating workflows with Airflow/Prefect/Control-M, and implementing data governance, quality frameworks, monitoring, CI/CD, and secure, compliant data access for enterprise analytics and reporting.

Neelesh Kasukurthi

Senior Data Engineer with 10+ years of experience designing and maintaining large-scale ETL/ELT pipelines, streaming architectures, and modern data platforms (data lakes/warehouses). Proficient in SQL, Python, Apache Spark, Databricks, Snowflake, and cloud ecosystems across AWS/Azure/GCP. Experienced in building real-time ingestion and processing with Kafka/Kinesis/Event Hubs/Pub/Sub, orchestrating workflows with Airflow/Prefect/Control-M, and implementing data governance, quality frameworks, monitoring, CI/CD, and secure, compliant data access for enterprise analytics and reporting.

Available to hire

Senior Data Engineer with 10+ years of experience designing and maintaining large-scale ETL/ELT pipelines, streaming architectures, and modern data platforms (data lakes/warehouses). Proficient in SQL, Python, Apache Spark, Databricks, Snowflake, and cloud ecosystems across AWS/Azure/GCP.

Experienced in building real-time ingestion and processing with Kafka/Kinesis/Event Hubs/Pub/Sub, orchestrating workflows with Airflow/Prefect/Control-M, and implementing data governance, quality frameworks, monitoring, CI/CD, and secure, compliant data access for enterprise analytics and reporting.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Beginner
See more

Work Experience

Senior Big Data Engineer at Cigna Health
August 1, 2024 - Present
Engineered HIPAA-compliant healthcare ETL/ELT pipelines using Azure Data Factory and Azure Synapse for large-scale claims data processing. Built real-time streaming transformations on Databricks with Spark Structured Streaming and event ingestion via Azure Event Hubs and Kafka/Kafka Streams into Azure Data Lake, supporting sub-minute latency and millions of events daily. Orchestrated 200+ pipelines using Prefect and Control-M with dependency management, retries, and SLA monitoring. Implemented Apache Hudi for incremental processing with ACID transactions and record-level updates/deletes, supporting audit history for regulatory compliance. Containerized services with Docker/Kubernetes, provisioned infrastructure using Terraform, and established CI/CD with GitHub Actions. Migrated and modernized legacy Hadoop/MapReduce and Talend workflows to cloud-native Spark/Databricks and ADF/Azure Synapse, improving processing time and reducing costs. Delivered governance and quality using Azure Pur
Senior Data Engineer at TransUnion
October 1, 2022 - July 1, 2024
Designed streaming and batch data platforms for banking/transaction analytics using Kafka/Kafka Connect and real-time processing with Spark Structured Streaming and Apache Flink. Implemented exactly-once ingestion semantics, schema evolution, and fraud detection pattern/anomaly detection with stateful stream processing at sub-second latency. Built Delta Lake medallion architectures (bronze/silver/gold) on Databricks supporting ACID, time travel, and regulatory compliance. Developed ETL transformations in Spark/Scala for nested JSON and optimized partitioning. Implemented data quality frameworks with Great Expectations and Deequ to prevent downstream corruption. Orchestrated workflows with Airflow across 200+ pipelines including sensors, retries, and SLA monitoring. Implemented CDC pipelines with AWS Kinesis and Kafka Streams and optimized Spark SQL performance via AQE, partition pruning, and broadcast joins. Established IaC with Terraform, deployed containerized workloads on Kubernetes
Senior Data Engineer at Costco
December 1, 2019 - September 1, 2022
Built real-time retail data ingestion pipelines using Kafka and GCP Pub/Sub/Kafka Streams and used Apache Beam for enrichment and aggregation. Developed Spark Structured Streaming jobs with PySpark and persisted processed outputs into BigQuery for immediate analytics. Implemented scalable ETL using Apache NiFi, orchestrated workflows with Dagster, and scheduled batch jobs via Autosys while monitoring with GCP Stackdriver. Designed Google Cloud Storage data lake architecture using Apache Iceberg for ACID transactions and schema evolution. Integrated datasets with Snowflake and built Looker/BI dashboards for executive reporting. Implemented IaC with Terraform for GCP Dataflow pipelines, Kubernetes clusters, and BigQuery datasets. Secured data using GCP IAM and KMS with Apache Ranger for fine-grained authorization. Created CI/CD for PySpark and streaming jobs with GitLab CI/CD and automated data quality validations with custom Python rules and Beam-based checks. Set up Splunk and Apache A
Data Engineer at Delta Airlines
February 1, 2018 - November 1, 2019
Implemented data pipelines with Azure Data Factory to extract from SQL Server and load into Azure SQL Database using incremental loads and robust error handling. Developed PySpark transformations on Azure Databricks with Delta Lake for versioning and performance-efficient staging in Azure Blob Storage. Built real-time telemetry ingestion using Azure Stream Analytics with windowing/aggregations, persisting results to Azure Synapse and alerting via Azure Monitor. Designed dimensional models in Azure Synapse Analytics with fact/dimension tables, partitioning strategies, and columnstore indexes for analytical performance. Created executive dashboards using Power BI with DAX and row-level security. Automated provisioning and CI/CD using Terraform/ARM and Jenkins/Git. Secured secrets and access with Azure Key Vault and Azure Active Directory/managed identities. Established hybrid connectivity using Azure ExpressRoute and private endpoints; implemented logging/monitoring with Log Analytics an
ETL Developer at Ramco Systems
June 1, 2016 - November 1, 2017
Developed ETL workflows in Informatica PowerCenter to extract from MySQL and flat files, perform complex transformations, and load into Teradata with SCD Type 2 logic. Built ETL mappings to migrate enterprise data from Teradata to AWS Redshift using parallel processing and partitioning, optimizing sessions for large volumes and enforcing data quality/integrity. Modeled dimensional schemas in Redshift (star schema) and wrote SQL for validation, using distribution keys to improve performance. Integrated REST APIs to extract nested JSON into staging on AWS S3, supporting incremental loads and error handling for failed transactions. Automated ETL triggering via AWS Lambda when files arrived in S3; implemented logging/monitoring with CloudWatch and failure notifications via SNS. Created Tableau dashboards connected to Redshift/MySQL with scheduled refresh and access controls. Established CI/CD with Jenkins/Git for Informatica deployments across environments; secured AWS access using IAM rol

Education

Bachelors of engineering in Computer Engineering at Padmabhushan Vasantdada Patil Pratishthan's College of Engineering
August 1, 2012 - May 1, 2016

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Financial Services, Retail, Other