Data Engineer with 5+ years of experience designing and deploying production-grade cloud data platforms across AWS, Azure, Databricks, and Snowflake for healthcare and enterprise clients. Skilled in building scalable ETL/ELT pipelines, Spark and PySpark-based Big Data processing, and modern Lakehouse architectures (Delta Lake) that support AI-ready analytics. Proven track record of reducing infrastructure costs by up to 30%, improving pipeline reliability, and enforcing HIPAA-aligned data governance across sensitive claims and clinical datasets. Adept at workflow orchestration (Airflow), CI/CD automation, and cross-functional delivery in fast-paced, regulated environments. Currently driving cloud migration and cost-optimization initiatives at CVS Health, translating complex data engineering work into measurable business impact.

PHANI VARSHITH KODURI

Data Engineer with 5+ years of experience designing and deploying production-grade cloud data platforms across AWS, Azure, Databricks, and Snowflake for healthcare and enterprise clients. Skilled in building scalable ETL/ELT pipelines, Spark and PySpark-based Big Data processing, and modern Lakehouse architectures (Delta Lake) that support AI-ready analytics. Proven track record of reducing infrastructure costs by up to 30%, improving pipeline reliability, and enforcing HIPAA-aligned data governance across sensitive claims and clinical datasets. Adept at workflow orchestration (Airflow), CI/CD automation, and cross-functional delivery in fast-paced, regulated environments. Currently driving cloud migration and cost-optimization initiatives at CVS Health, translating complex data engineering work into measurable business impact.

Available to hire

Data Engineer with 5+ years of experience designing and deploying production-grade cloud data platforms across AWS, Azure, Databricks, and Snowflake for healthcare and enterprise clients. Skilled in building scalable ETL/ELT pipelines, Spark and PySpark-based Big Data processing, and modern Lakehouse architectures (Delta Lake) that support AI-ready analytics.

Proven track record of reducing infrastructure costs by up to 30%, improving pipeline reliability, and enforcing HIPAA-aligned data governance across sensitive claims and clinical datasets. Adept at workflow orchestration (Airflow), CI/CD automation, and cross-functional delivery in fast-paced, regulated environments. Currently driving cloud migration and cost-optimization initiatives at CVS Health, translating complex data engineering work into measurable business impact.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate

Work Experience

Data Engineer at CVS Health
January 1, 2025 - Present
Architected scalable ETL/ELT pipelines on AWS (S3, Glue, EMR) and Databricks to process 2+ TB of daily pharmacy claims, eligibility, and EHR data, reducing end-to-end pipeline latency by 35% through distributed PySpark workflows. Built a Delta Lake-based data lakehouse integrating Snowflake and AWS Glue, reducing annual data warehousing costs by 22% by consolidating siloed claims and pharmacy datasets. Designed and orchestrated 50+ interdependent Airflow DAGs, reducing manual pipeline interventions by 40% via dependency-aware scheduling, retries, and SLA-based alerting. Implemented dbt and Great Expectations transformation and testing layer, increasing automated data quality test coverage by 60%. Enforced HIPAA-aligned data governance with PII/PHI masking, RBAC, and automated data cataloging, achieving zero compliance incidents over 18 months. Led migration from legacy on-prem ETL to cloud-native AWS architecture, cutting infrastructure costs by 30% and improving processing speed by 45
Data Engineer at Infosys
September 1, 2021 - June 1, 2023
Developed ETL pipelines using Azure Data Factory and PySpark to process healthcare claims and provider data, reducing nightly batch processing time by 30% via parallelized ingestion and transformations. Migrated on-prem SQL Server data warehouses to Azure Synapse Analytics, reducing query response times by 20% by redesigning schemas and implementing partitioning and indexing best practices. Built ingestion pipelines to extract and parse semi-structured JSON and Parquet payloads via REST APIs into SQL Server and cloud data lakes, reducing manual data pull efforts by 20%. Implemented automated data quality checks and reconciliation processes across claims datasets, reducing downstream discrepancies by 40%. Collaborated with business analysts and QA teams to translate healthcare data requirements into technical specifications, improving sprint delivery predictability across six release cycles through documented mapping logic and data lineage. Optimized ADF pipeline costs, reducing monthly
Junior Data Engineer at Insight Global
August 1, 2020 - August 1, 2021
Developed SQL-based ETL scripts to extract, transform, and load data into reporting tables, reducing manual reporting effort by 20% through automation of recurring data pull and refresh processes. Assisted in building and maintaining Python data processing scripts for internal analytics, speeding turnaround on ad hoc data requests by creating reusable functions for common data cleaning tasks. Supported migration of spreadsheets and flat-file sources into a centralized SQL Server database, improving data accuracy and establishing a single source of truth via normalized schema design and load scripts. Wrote and optimized SQL queries and stored procedures for validation and reporting, improving query performance by 15% via indexing and query tuning. Contributed to pipeline modernization initiatives by shadowing senior engineers and completing internal training on cloud and big data concepts.

Education

Master of Science in Computer Science at Oklahoma City University
January 11, 2030 - May 1, 2025
Bachelor of Technology in Computer Science at Mallareddy Institute of Technology
January 11, 2030 - May 1, 2021

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate

Hire a Data Engineer

We have the best data engineer experts on Twine. Hire a data engineer today.