I’m an Azure- and AWS-focused data engineer with 6.7+ years of experience designing and optimizing secure, scalable data platforms and production-grade ETL/ELT pipelines. I build trusted data assets using services like Azure Data Factory, Data Bricks, Synapse Analytics, Snowflake, and PySpark/SQL—supporting analytics, forecasting, and revenue optimization use cases. I also partner closely with data scientists and business stakeholders to enable retail pricing and promotion optimization (including LGM-based simulations) and to deliver reliable machine learning data pipelines. In addition to core engineering, I’ve implemented strong CI/CD automation using Azure DevOps, YML, and governance controls such as RBAC, masking, and encryption—improving pipeline reliability, reducing cloud compute costs, and accelerating downstream analytics.

Dhiraj Tiwari

I’m an Azure- and AWS-focused data engineer with 6.7+ years of experience designing and optimizing secure, scalable data platforms and production-grade ETL/ELT pipelines. I build trusted data assets using services like Azure Data Factory, Data Bricks, Synapse Analytics, Snowflake, and PySpark/SQL—supporting analytics, forecasting, and revenue optimization use cases. I also partner closely with data scientists and business stakeholders to enable retail pricing and promotion optimization (including LGM-based simulations) and to deliver reliable machine learning data pipelines. In addition to core engineering, I’ve implemented strong CI/CD automation using Azure DevOps, YML, and governance controls such as RBAC, masking, and encryption—improving pipeline reliability, reducing cloud compute costs, and accelerating downstream analytics.

Available to hire

I’m an Azure- and AWS-focused data engineer with 6.7+ years of experience designing and optimizing secure, scalable data platforms and production-grade ETL/ELT pipelines. I build trusted data assets using services like Azure Data Factory, Data Bricks, Synapse Analytics, Snowflake, and PySpark/SQL—supporting analytics, forecasting, and revenue optimization use cases.

I also partner closely with data scientists and business stakeholders to enable retail pricing and promotion optimization (including LGM-based simulations) and to deliver reliable machine learning data pipelines. In addition to core engineering, I’ve implemented strong CI/CD automation using Azure DevOps, YML, and governance controls such as RBAC, masking, and encryption—improving pipeline reliability, reducing cloud compute costs, and accelerating downstream analytics.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner

Work Experience

Senior Data Engineer at Mars (Contract)
December 1, 2023 - Present
Architected and optimized scalable Azure-based data pipelines for a Revenue Growth Management (RGM) platform, processing large-scale sales, pricing, promotional, and financial datasets to support retail pricing strategies while reducing cloud compute costs by ~30% through workload optimization and efficient resource utilization. Designed enterprise-grade ETL/ELT pipelines using Azure Data Factory, Azure Databricks, PySpark, SQL, Delta Lake, Synapse Analytics, and Snowflake to deliver trusted datasets for revenue forecasting, profitability analysis, and business reporting. Collaborated with data scientists and business stakeholders to enable retail pricing and promotional optimization using the LGM model and forecasting models built with Python, Prophet, and Azure Machine Learning, ensuring production-ready pipelines for pricing simulations and revenue optimization. Implemented robust CI/CD pipelines with Azure DevOps and YAML, leveraging Data Bricks Python wheel deployments for automat
Data Engineer at QBE (Contract) / GSK Financial Data Modernization (Contract)
February 1, 2022 - December 1, 2023
Developed and maintained scalable ETL/ELT pipelines using Azure Data Factory, Azure Databricks, PySpark, and ADLS Gen2 to ingest and process banking, payment, accounts receivable (AR), accounts payable (AP), general ledger (GL), and financial transaction data from multiple enterprise systems. Built data transformation workflows using Python, PySpark, Spark SQL, and Delta Lake to clean, validate, reconcile, and standardize financial data for accurate financial reporting and regulatory-ready analytics. Implemented centralized Python logging, monitoring, and automated alerting across Databricks and Azure Data Factory, improving pipeline reliability and reducing production downtime by ~80%. Optimized Azure SQL database and Spark workloads via query tuning, partitioning, and performance enhancements to reduce query response time by ~50% for finance reporting and analytics teams. Delivered curated financial datasets for cash flow analysis, bank reconciliation, payment tracking, financial rep
Senior Data Engineer at DHIRAJ TIWARI (Senior Data Engineer) / GSK (Contract)
January 1, 2020 - February 1, 2022
Designed scalable, secure, cloud-native data platform capabilities across Azure and AWS with expertise in Azure Data Factory, Databricks, Synapse Analytics, Snowflake, Azure Machine Learning, and AWS services, implementing hands-on feature engineering, ML data pipelines, MLOps integration, data governance, CI/CD automation, and cost optimization. Built high-performance ETL/ELT pipelines to ingest, transform, and process high-volume insurance and financial data using Azure Data Factory, Azure Databricks, ADLS Gen2, Snowflake, PySpark, and SQL. Implemented a Data Vault 2.0 model (hub/link/satellite) in Snowflake to enable scalable data integration, historical tracking, and trusted enterprise data assets for reporting and analytics. Implemented enterprise-grade security and governance using Snowflake RBAC, Azure RBAC, dynamic data masking, AES encryption, and fine-grained access controls for sensitive PII and insurance data. Automated code deployment and release management using Git, Azur

Education

Bachelor of Technology in Information Technology at Indian Institute of Technology
August 1, 2015 - June 1, 2019

Qualifications

DataBricks Certified Data Engineer Associate
January 11, 2030 - August 31, 2026
Building AI Agents with LangChain (Analytics Vidhya)
January 11, 2030 - August 31, 2026
Building AI Agents from Scratch (Analytics Vidhya)
January 11, 2030 - August 31, 2026
Star Award (Accenture)
January 11, 2030 - August 31, 2026

Industry Experience

Retail, Financial Services, Life Sciences, Software & Internet, Professional Services