Senior Data Engineer with 8+ years of experience designing and optimizing enterprise data platforms across Workforce Analytics, Healthcare, and Telecommunications. Expertise in cloud-native Data Lake/Lakehouse architectures using Azure Databricks, AWS, Snowflake, Delta Lake, and Apache Spark to process large-scale structured, semi-structured, and streaming datasets. Strong background building scalable ETL/ELT pipelines with PySpark, Spark SQL, Azure Data Factory, AWS Glue, Apache Airflow, Kafka, and Delta Live Tables, including Medallion Architecture, CDC, and SCD Type 2. Hands-on in data governance, quality (Great Expectations, Unity Catalog), performance tuning, and infrastructure automation with Terraform, Docker, and CI/CD tools, collaborating with cross-functional teams to deliver secure, high-performance, and cost-optimized data solutions.

Surakshya Aryal

Senior Data Engineer with 8+ years of experience designing and optimizing enterprise data platforms across Workforce Analytics, Healthcare, and Telecommunications. Expertise in cloud-native Data Lake/Lakehouse architectures using Azure Databricks, AWS, Snowflake, Delta Lake, and Apache Spark to process large-scale structured, semi-structured, and streaming datasets. Strong background building scalable ETL/ELT pipelines with PySpark, Spark SQL, Azure Data Factory, AWS Glue, Apache Airflow, Kafka, and Delta Live Tables, including Medallion Architecture, CDC, and SCD Type 2. Hands-on in data governance, quality (Great Expectations, Unity Catalog), performance tuning, and infrastructure automation with Terraform, Docker, and CI/CD tools, collaborating with cross-functional teams to deliver secure, high-performance, and cost-optimized data solutions.

Available to hire

Senior Data Engineer with 8+ years of experience designing and optimizing enterprise data platforms across Workforce Analytics, Healthcare, and Telecommunications. Expertise in cloud-native Data Lake/Lakehouse architectures using Azure Databricks, AWS, Snowflake, Delta Lake, and Apache Spark to process large-scale structured, semi-structured, and streaming datasets.

Strong background building scalable ETL/ELT pipelines with PySpark, Spark SQL, Azure Data Factory, AWS Glue, Apache Airflow, Kafka, and Delta Live Tables, including Medallion Architecture, CDC, and SCD Type 2. Hands-on in data governance, quality (Great Expectations, Unity Catalog), performance tuning, and infrastructure automation with Terraform, Docker, and CI/CD tools, collaborating with cross-functional teams to deliver secure, high-performance, and cost-optimized data solutions.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Work Experience

Senior Data Engineer at Skill Squirrel
November 1, 2024 - Present
Designed and developed scalable ETL/ELT pipelines using Azure Databricks, PySpark, Azure Data Factory, and AWS Glue, processing over 5 TB of workforce analytics data daily. Built enterprise Lakehouse architecture using Azure Data Lake Storage Gen2, Amazon S3, Delta Lake, and Snowflake following Medallion Architecture. Developed Kafka and Spark Structured Streaming pipelines processing more than 1 million candidate activity events daily. Implemented Apache Airflow DAGs and Delta Live Tables (DLT), improving pipeline reliability and reducing manual intervention by 40%. Optimized Spark workloads using partitioning, caching, AQE, and broadcast joins, reducing processing time by 45%. Integrated Great Expectations data quality framework and enterprise data observability to improve data quality to 99.5%. Built CI/CD pipelines using Terraform, Azure DevOps, Jenkins, and GitHub Actions; collaborated with AI/ML teams and supported MLflow-enabled pipelines.
Data Engineer at UnitedHealth Group (UHG)
February 1, 2022 - August 31, 2024
Designed and developed ETL/ELT pipelines for healthcare claims, pricing, and membership data using Azure Databricks, PySpark, Spark SQL, and Delta Lake. Implemented Medallion Architecture (Bronze/Silver/Gold) for enterprise healthcare analytics, leveraging Delta Lake ACID, schema evolution, CDC, and file optimization. Built SCD Type 2 data models for historical tracking. Implemented Unity Catalog and role-based access controls for HIPAA-compliant governance. Tuned Spark workloads using partitioning, caching, broadcast joins, Z-Ordering, and compaction to improve query performance. Collaborated with data scientists, analysts, and product teams to deliver production-ready datasets and improved cluster utilization and cost efficiency.
Data Engineer at T-Mobile
June 1, 2020 - January 31, 2022
Developed scalable PySpark and Spark SQL pipelines processing over 10 TB of telecom data daily for customer, billing, subscriber, and network analytics. Built streaming ingestion frameworks using AWS Glue, Kafka, and Spark Structured Streaming for near real-time analytics. Designed enterprise Lakehouse architecture using Amazon S3, Delta Lake, Snowflake, and Azure Databricks for Customer 360 analytics (50M+ subscribers). Developed Kafka-based streaming applications reducing network alert processing latency from two hours to under fifteen minutes. Orchestrated 500+ ETL jobs with Apache Airflow and Azure Data Factory, maintaining 99.9% SLA. Optimized analytics performance by 60% through Snowflake/Spark tuning and metadata-driven CDC incremental processing. Automated CI/CD and infrastructure with Terraform, Jenkins, Docker, and GitHub Actions.
Data Engineer at Cedar Gate Technologies
October 1, 2017 - May 31, 2020
Built enterprise ETL pipelines using AWS Glue, Amazon EMR, Spark, and PySpark processing 100M+ healthcare records. Designed cloud data lake architecture using Amazon S3 and Snowflake for healthcare analytics. Processed multi-terabyte claims/provider/member datasets using distributed Spark. Automated ETL workflows with Apache Airflow, achieving 99.9% SLA compliance. Developed dimensional models and CDC pipelines for reporting and analytics. Delivered reporting solutions with Amazon Redshift and Snowflake; built Kafka streaming pipelines for near real-time healthcare events. Implemented IAM security policies and HIPAA-compliant governance controls. Optimized Spark and Snowflake workloads to reduce ETL failures by 35%, with automated deployment/monitoring via Terraform, Jenkins, and CloudWatch.

Education

Add your educational history here.

Qualifications

Masters of Data Analytics
January 11, 2030 - August 13, 2026
Bachelor of Engineering in Electronics and Communications
January 11, 2030 - August 13, 2026
AWS Cloud Practitioner
January 11, 2030 - August 13, 2026
Microsoft Azure fundamental
January 11, 2030 - August 13, 2026
Masters of Data Analytics
January 11, 2030 - August 13, 2026
Bachelor of Engineering in Electronics and Communications
January 11, 2030 - August 13, 2026
AWS Cloud Practitioner
January 11, 2030 - August 13, 2026
Microsoft Azure fundamental
January 11, 2030 - August 13, 2026

Industry Experience

Healthcare, Telecommunications, Other