Data Engineer with 5+ years of experience designing and building scalable cloud-native data platforms, Lakehouse architectures, and enterprise data warehouses on Azure, GCP, and Databricks. Experienced in developing batch and real-time ELT pipelines using Python, PySpark, Spark SQL, Azure Data Factory, Databricks, BigQuery, and Delta Lake. Strong expertise in data modeling, streaming, performance optimization, data governance, CI/CD automation, and cloud-native orchestration. Proven ability to collaborate with cross-functional teams, translate business requirements into scalable data solutions, and deliver secure, high-performance, analytics and AI-ready data platforms.

Durga Linga

Data Engineer with 5+ years of experience designing and building scalable cloud-native data platforms, Lakehouse architectures, and enterprise data warehouses on Azure, GCP, and Databricks. Experienced in developing batch and real-time ELT pipelines using Python, PySpark, Spark SQL, Azure Data Factory, Databricks, BigQuery, and Delta Lake. Strong expertise in data modeling, streaming, performance optimization, data governance, CI/CD automation, and cloud-native orchestration. Proven ability to collaborate with cross-functional teams, translate business requirements into scalable data solutions, and deliver secure, high-performance, analytics and AI-ready data platforms.

Available to hire

Data Engineer with 5+ years of experience designing and building scalable cloud-native data platforms, Lakehouse architectures, and enterprise data warehouses on Azure, GCP, and Databricks. Experienced in developing batch and real-time ELT pipelines using Python, PySpark, Spark SQL, Azure Data Factory, Databricks, BigQuery, and Delta Lake.

Strong expertise in data modeling, streaming, performance optimization, data governance, CI/CD automation, and cloud-native orchestration. Proven ability to collaborate with cross-functional teams, translate business requirements into scalable data solutions, and deliver secure, high-performance, analytics and AI-ready data platforms.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more

Work Experience

Data Engineer at ConEdison, USA
January 1, 2026 - Present
Contributed to the design and implementation of an enterprise Azure Lakehouse architecture using Azure Databricks, ADLS Gen2, Delta Lake, and Unity Catalog for reporting and AI workloads. Built scalable batch and incremental ELT pipelines using Azure Data Factory and Azure Databricks to ingest structured, semi-structured, and unstructured data into a Medallion (Bronze/Silver/Gold) architecture. Developed PySpark applications processing 20+ TB daily across work order, asset, procurement, invoice, and operational datasets. Implemented Auto Loader and CDC pipelines for efficient incremental ingestion into Delta Lake. Developed Delta Live Tables (DLT) pipelines with expectations and automated data quality checks. Built streaming pipelines using Azure Event Hubs and Apache Kafka for near real-time operational events. Optimized Spark jobs (partitioning, caching, AQE, Z-Ordering, OPTIMIZE/VACUUM) improving performance by ~25%. Orchestrated end-to-end workflows using Azure Data Factory and Dat
Data Engineer at GSK, San Francisco, CA
January 1, 2024 - December 1, 2025
Designed and implemented enterprise-scale healthcare data lakehouse architecture on Google Cloud supporting 500+ million patient and claims records. Built ETL/ELT pipelines using PySpark, Apache Beam, Dataflow, and Dataproc to ingest structured, semi-structured, and unstructured datasets. Implemented real-time streaming pipelines using Google Pub/Sub and Apache Kafka for patient encounters, lab reports, pharmacy transactions, and device telemetry. Integrated data sources including Epic, Cerner, HL7 messages, FHIR APIs, EMR/EHR systems, and insurance claims; implemented CDC pipelines to synchronize clinical databases into BigQuery with near real-time processing. Optimized BigQuery warehouse using SQL tuning, partitioning/clustering, materialized views, and BI Engine acceleration. Developed reusable dbt models for incremental transformations and standardized business logic; built dimensional models (star and snowflake) for clinical, financial, claims, provider, and patient analytics. Red
ETL Developer at Zentron Labs Pvt Ltd, India
July 1, 2021 - July 1, 2023
Developed and maintained scalable ETL workflows using Informatica PowerCenter to load healthcare claims, provider, and patient data into an enterprise warehouse. Built mappings, mapplets, reusable transformations, workflows, worklets, and sessions following ETL best practices. Extracted data from Oracle, SQL Server, flat files, XML, and REST APIs. Implemented incremental loading, CDC, and SCD Type 1/2 to support accurate historical and current views. Processed over 40 million records daily with high data accuracy. Developed complex SQL (joins, window functions, stored procedures) and improved performance through partitioning/pushdown optimization, bulk loading, and indexing (reducing execution time by 35%). Migrated historical datasets to Snowflake via AWS S3 staging and AWS Glue. Created Python utilities for validation and automation; scheduled jobs with Control-M and monitored using Informatica Workflow Monitor. Added robust exception handling and restart mechanisms; supported produc

Education

Master of Science in Applied Statistics and Data Science at University of Texas at Arlington, USA
August 1, 2023 - December 1, 2024
Bachelor's in Electrical Engineering at Indian Institute of Technology Tirupati, India
January 1, 2017 - January 1, 2021

Qualifications

Google Cloud Certified – Professional Data Engineer
January 11, 2030 - August 12, 2026

Industry Experience

Healthcare