Data Engineer with 10+ years of experience building and optimizing scalable ETL/ELT pipelines, data warehouses, and data lakes for enterprise analytics across AWS, Azure, and GCP. Strong hands-on expertise in SQL, Python, Apache Spark/PySpark, and modern big data tools to deliver reliable, high-performance data products.

DHANYA M

Data Engineer with 10+ years of experience building and optimizing scalable ETL/ELT pipelines, data warehouses, and data lakes for enterprise analytics across AWS, Azure, and GCP. Strong hands-on expertise in SQL, Python, Apache Spark/PySpark, and modern big data tools to deliver reliable, high-performance data products.

Available to hire

Data Engineer with 10+ years of experience building and optimizing scalable ETL/ELT pipelines, data warehouses, and data lakes for enterprise analytics across AWS, Azure, and GCP. Strong hands-on expertise in SQL, Python, Apache Spark/PySpark, and modern big data tools to deliver reliable, high-performance data products.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

English
Intermediate

Work Experience

Sr. Big Data Engineer at Visa
October 1, 2023 - Present
Built large-scale ingestion and transformation pipelines using Python and PySpark on Amazon EMR, applying functional patterns for reusable and maintainable code under strict data quality requirements. Designed and optimized SQL queries using Presto over petabyte-scale datasets in S3 for faster reconciliation and reporting. Configured and migrated EMR workloads from EMR on EC2 to EMR on EKS via Docker containerization for scalable processing and reduced costs. Deployed and managed EKS for containerized PySpark/Python services with autoscaling and secure IAM controls. Orchestrated event-driven workflows using AWS Lambda and Step Functions for fraud detection ingestion, including real-time S3 triggers and enriched Presto-queryable tables. Developed generative AI and RAG solutions using Amazon Bedrock/Agents with Knowledge Bases integrated with enterprise data. Implemented CI/CD with Jenkins and Argo CD, Terraform/CloudFormation provisioning, and observability with CloudWatch/Grafana/Prome
Sr Data Engineer at Macy's
December 1, 2020 - September 30, 2023
Delivered end-to-end retail ingestion pipelines using PySpark on Amazon EMR on EC2 and containerized PySpark transformations for EMR on EKS. Stored processed datasets to partitioned S3 paths for downstream analytics. Built and optimized Trino SQL for multi-million row retail datasets to support merchandising and supply chain reporting with reduced query time. Created event-driven pipelines using AWS Lambda and Step Functions to trigger EMR jobs when new S3 files landed, enabling near real-time availability. Migrated on-prem ETL to EMR on EKS by repackaging legacy PySpark jobs into Docker orchestrated on Kubernetes. Established CI/CD using GitHub Actions and Argo CD for build/test/deploy workflows. Enforced least-privilege access using AWS IAM and Guardrails; monitored clusters and workflows via CloudWatch; audited via CloudTrail and centralized logs in ELK/Splunk-style dashboards. Built governance using Data Catalog with schema/versioning and classification tags, and supported ML featu
Data Engineer at United Health Group
June 1, 2018 - November 30, 2020
Built healthcare claims ingestion pipelines using PySpark on EMR on EC2, loading curated outputs into Apache Hive for analytics. Implemented distributed Spark batch processing for healthcare encounter/billing records with executor tuning and dynamic partitioning managed through version-controlled bootstrap scripts. Migrated Hive ETL workloads to EMR on EKS by containerizing PySpark using Docker and orchestrating on Kubernetes. Automated event-driven execution using AWS Lambda and Step Functions with retry/failure handling and lineage audit trails in S3. Provisioned EMR, EKS node groups, and Lambda via Terraform; deployed via GitLab CI and Argo CD with build/test automation. Enforced PHI security and HIPAA-compliant access using AWS Guardrails and IAM roles, and monitored production with CloudWatch and Splunk. Managed governance via Data Catalog schema/classification/lineage metadata and validated transformations through test automation on synthetic datasets.
Data Engineer at UPS
February 1, 2017 - May 31, 2018
Architected a transportation analytics platform using AWS S3 for storage and EC2/EMR for distributed processing of high-volume vehicle telemetry. Developed real-time shipment tracking using Apache Flume, processing with MapReduce and aggregating in Redshift for operational dashboards (Tableau). Automated freight management workflows by extracting data from Oracle, transforming in Informatica, and loading into Aurora. Implemented predictive maintenance with EMR + Hive queries on sensor data and scheduled preventive actions via cron. Built streaming route optimization with Apache Storm topology, correlating live traffic and weather with recommendations delivered through messaging queues (ActiveMQ). Set up CI/CD with Git and Docker, deployed microservices, and monitored health with CloudWatch alerts. Migrated legacy SQL Server warehousing to RDS with dimensional modeling; performed batch reconciliation and loaded to Teradata. Implemented streaming logistics event processing with Flume/Had
ETL Developer at Ramco Systems
November 1, 2015 - December 31, 2016
Built an analytics platform on GCP using Compute Engine and BigQuery to process daily transaction records and serve insights through Tableau dashboards. Developed ETL workflows with Informatica PowerCenter to extract via FTP, apply business rules/models, and load data into Cloud Storage. Designed a star-schema dimensional warehouse in ER/Studio and implemented fact/dimension tables in BigQuery with complex SQL aggregations. Implemented real-time inventory synchronization using Kafka ingestion and Hadoop/Hive for reconciled stock availability. Orchestrated batch processing using MapReduce and scheduled scripts. Collaborated on migrating financial reporting infrastructure to GCP, transforming Excel into BigQuery datasets and validating with Tableau to achieve ~99.8% accuracy for stakeholders. Used SVN for version control and followed SDLC/Scrum deliverables with testing and QA.

Education

BTech in Computer Science at ANDHRA UNIVERSITY COLLEGE OF ENGINEERING
June 1, 2011 - April 30, 2015

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Retail, Transportation & Logistics, Software & Internet, Professional Services