Senior Data Engineer with 5+ years building production data platforms across healthcare and financial services, owning 25 to 35 end-to-end pipelines that process millions of records daily. I focus on performance, reliability, governance, and cost optimization—cutting runtimes and compute costs using Spark, incremental ingestion, and Delta Lake patterns. I’ve engineered robust ETL/ELT workflows with Airflow, dbt, and data quality automation to drive strong SLA performance, and I design governed dimensional/Medallion models for trusted BI and self-service analytics (including Databricks Genie). I’m experienced operating in regulated environments and translating stakeholder requirements into scalable, auditable data architectures on AWS and Azure.

Vamsi Krishna Mamillapalli

Senior Data Engineer with 5+ years building production data platforms across healthcare and financial services, owning 25 to 35 end-to-end pipelines that process millions of records daily. I focus on performance, reliability, governance, and cost optimization—cutting runtimes and compute costs using Spark, incremental ingestion, and Delta Lake patterns. I’ve engineered robust ETL/ELT workflows with Airflow, dbt, and data quality automation to drive strong SLA performance, and I design governed dimensional/Medallion models for trusted BI and self-service analytics (including Databricks Genie). I’m experienced operating in regulated environments and translating stakeholder requirements into scalable, auditable data architectures on AWS and Azure.

Available to hire

Senior Data Engineer with 5+ years building production data platforms across healthcare and financial services, owning 25 to 35 end-to-end pipelines that process millions of records daily. I focus on performance, reliability, governance, and cost optimization—cutting runtimes and compute costs using Spark, incremental ingestion, and Delta Lake patterns.

I’ve engineered robust ETL/ELT workflows with Airflow, dbt, and data quality automation to drive strong SLA performance, and I design governed dimensional/Medallion models for trusted BI and self-service analytics (including Databricks Genie). I’m experienced operating in regulated environments and translating stakeholder requirements into scalable, auditable data architectures on AWS and Azure.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

Work Experience

Senior Data Engineer | Analytics Engineer | BI and Data Warehouse Engineer at Helix Biomedix
August 1, 2024 - Present
Reduced pipeline compute costs by ~25% via Python-based change data capture to replace full table reloads with incremental, row-level ingestion. Improved SLA compliance to ~95% by orchestrating Airflow (Astronomer) with dependency management, retries, backfills, and proactive alerting. Accelerated heavy Spark batch jobs by ~30% through skew handling (salting, broadcast joins, partition tuning). Enforced data contracts between Medallion layers to prevent upstream schema changes from breaking downstream consumers, reducing incidents by ~50%. Built governed Bronze/Silver/Gold (Medallion) Delta Lake datasets and owned 25–35 end-to-end production pipelines. Developed Spark Structured Streaming pipelines over Kafka with checkpointing and idempotent Delta writes for exactly-once processing. Improved Databricks Genie accuracy to ~85–90% by curating metadata and grounding instructions for self-service analytics. Strengthened governance and data quality using Unity Catalog (RBAC, lineage, co
Data Engineer at Sundaram Finance
February 1, 2019 - December 1, 2022
Shortened batch processing time by ~35% by migrating legacy SQL stored procedures to distributed PySpark pipelines on EMR for same-day regulatory reporting. Built and maintained ETL pipelines ingesting millions of loan/transaction/customer records into a Redshift warehouse used by risk, collections, and finance teams. Architected an AWS data warehouse on S3, Glue, and Redshift, consolidating fragmented sources into a centralized analytics platform. Sped up month-end close by modeling star-schema dimensional warehouses with slowly changing dimensions and replacing ad hoc extracts with governed data marts. Ensured regulatory submission accuracy with reconciliation and validation checks between source systems and the warehouse, supporting SOX controls, PCI DSS handling for payment data, and data privacy requirements. Orchestrated pipelines in Apache Airflow with SLA alerting, dependency management, and automated retries. Tuned dashboard performance via Redshift distribution/sort keys and

Education

Master of Science, Business Analytics at University of Scranton, PA
January 1, 2023 - May 1, 2024

Qualifications

Databricks Certified Data Engineer Professional
January 11, 2030 - August 7, 2026
Microsoft Certified: Azure Data Engineer Associate
January 11, 2030 - August 7, 2026
Microsoft Certified: Azure Fundamentals (AZ-900)
January 11, 2030 - August 7, 2026
DataCamp Certified SQL Associate
January 11, 2030 - August 7, 2026

Industry Experience

Healthcare, Financial Services, Professional Services