Senior Data Engineer with around 7 years of experience designing, developing, and implementing enterprise-scale data engineering, analytics, and cloud modernization solutions across AWS and Microsoft Azure. Expertise in building scalable ETL/ELT pipelines, cloud-native data platforms, data lakes/warehouses, and real-time streaming architectures for BI, advanced analytics, and ML enablement. Hands-on with Python, PySpark/Spark SQL, Kafka, Airflow, Databricks, Snowflake, and major cloud services across AWS/Azure. Experienced in data modeling (Star/Snowflake schemas, SCD Type 1 & 2), Delta Lake, governance, lineage, and data quality frameworks, plus building secure APIs/microservices and production-grade CI/CD, monitoring, and orchestration.

Neha Reddy

Senior Data Engineer with around 7 years of experience designing, developing, and implementing enterprise-scale data engineering, analytics, and cloud modernization solutions across AWS and Microsoft Azure. Expertise in building scalable ETL/ELT pipelines, cloud-native data platforms, data lakes/warehouses, and real-time streaming architectures for BI, advanced analytics, and ML enablement. Hands-on with Python, PySpark/Spark SQL, Kafka, Airflow, Databricks, Snowflake, and major cloud services across AWS/Azure. Experienced in data modeling (Star/Snowflake schemas, SCD Type 1 & 2), Delta Lake, governance, lineage, and data quality frameworks, plus building secure APIs/microservices and production-grade CI/CD, monitoring, and orchestration.

Available to hire

Senior Data Engineer with around 7 years of experience designing, developing, and implementing enterprise-scale data engineering, analytics, and cloud modernization solutions across AWS and Microsoft Azure. Expertise in building scalable ETL/ELT pipelines, cloud-native data platforms, data lakes/warehouses, and real-time streaming architectures for BI, advanced analytics, and ML enablement.

Hands-on with Python, PySpark/Spark SQL, Kafka, Airflow, Databricks, Snowflake, and major cloud services across AWS/Azure. Experienced in data modeling (Star/Snowflake schemas, SCD Type 1 & 2), Delta Lake, governance, lineage, and data quality frameworks, plus building secure APIs/microservices and production-grade CI/CD, monitoring, and orchestration.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate

Work Experience

Sr. Data Engineer at RBC, Canada
August 1, 2023 - Present
Designed and developed end-to-end data engineering solutions across enterprise cloud platforms, including data sourcing, ingestion, transformation, lineage, quality validation, storage, and consumption. Built scalable batch and real-time pipelines using Python, PySpark, Apache Spark, Kafka, AWS S3, AWS Glue, Lambda, and Redshift. Developed reusable Python frameworks for ingestion, transformation, validation, and enrichment across heterogeneous data sources. Implemented REST APIs and microservices using Flask for secure data consumption and analytics applications. Containerized services with Docker and Kubernetes/OpenShift; orchestrated pipelines with Apache Airflow for scheduling, dependencies, monitoring, and recovery. Partnered with Data Scientists and business stakeholders to deliver analytics-ready datasets for ML/predictive analytics/operational intelligence, including integration with AWS SageMaker. Implemented data quality, lineage, metadata management, and governance frameworks
Sr. Data Engineer at PwC Canada
June 1, 2021 - July 31, 2023
Designed and developed scalable ETL/ELT pipelines using Azure Databricks, Azure Data Factory, PySpark, Spark SQL, and Azure Synapse Analytics for enterprise reporting and analytics. Built cloud-native ingestion frameworks integrating SQL Server, Teradata, MongoDB, Azure Blob Storage/ADLS Gen2, Cosmos DB, REST APIs, and Kafka. Developed and optimized PySpark applications for transformations and loading into Synapse and Lakehouse environments. Implemented dimensional models (Star/Snowflake schemas) and SCD Type 1 & 2; built Delta Lake capabilities such as ACID transactions, schema evolution, and time travel. Optimized Spark workloads with partitioning, caching, broadcast joins, and execution-plan tuning. Orchestrated workflows using ADF pipelines, Databricks Workflows, and Synapse Pipelines. Implemented data quality/validation/reconciliation frameworks and operational monitoring for SLA compliance. Developed solutions using Synapse dedicated/serverless SQL pools, integrated AAD and RBAC
Data Engineer at Sanofi, Canada
October 1, 2019 - May 31, 2021
Designed and developed cloud-native data pipelines using AWS S3, AWS Glue, AWS Lambda, Redshift, and Aurora for enterprise data integration, reporting, and analytics. Built ETL/ELT frameworks using Glue and PySpark/Python/SQL, including streaming data processing. Implemented serverless, event-driven architectures using Lambda and API Gateway for automated ingestion, transformation, and orchestration. Created near real-time streaming pipelines with Kinesis Data Streams/Firehose/Data Analytics. Designed and tuned Redshift warehouse solutions including dimensional modeling and star schemas. Developed complex SQL and stored procedures/views for reporting and BI. Integrated Glue with Step Functions for scheduling, dependency management, monitoring, and recovery. Built ingestion frameworks for APIs, relational databases, files, JSON, and cloud apps into AWS platforms. Utilized AWS CDK and CloudFormation for Infrastructure as Code. Implemented CI/CD with Git, CodeCommit, CodePipeline, and Jen

Education

Computer Science certificate at UCLID IT India Pvt. Ltd
January 1, 2019 - January 1, 2019

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Professional Services, Healthcare, Software & Internet