Available to hire
Data engineer with close to 10 years of experience building end-to-end pipelines for analytics, regulatory reporting, and machine learning workloads across AWS, Azure, and Snowflake. Experienced in batch and streaming ingestion, Spark/PySpark transformations, and warehouse tuning in regulated domains like banking, insurance, and healthcare.
Strong focus on data accuracy and timing for risk and regulatory use cases, with hands-on work integrating APIs, mainframe extracts, and message queues. Experienced with governance for PII/PHI, SLAs and data quality checks, and CI/CD-driven deployments using Git and Azure DevOps/Airflow.
Work Experience
Senior Data Engineer at BNP Paribas
June 1, 2025 - PresentBuilt real-time trade data pipelines using Kafka and Spark Structured Streaming for equity and fixed-income events to support risk teams within minutes. Standardized Bloomberg/Reuters market data and reference data across asset classes in Snowflake for downstream consumption. Developed PySpark jobs to compute intraday position exposures for market and counterparty risk. Contributed regulatory reporting data feeding Dodd-Frank and MiFID II transaction reports. Implemented a Python reconciliation framework comparing trade capture against front-/back-office systems to detect mismatches before regulatory release. Built AWS Glue and Lambda-based workflows for end-of-day settlement files. Implemented row-level security and masking in Snowflake as part of internal data governance. Tuned Snowflake warehouse sizing and clustering keys for high-cost risk dashboard tables. Improved intraday SLAs by addressing skewed joins and shuffle-heavy transformation stages. Added data quality checks (row cou
Senior Data Engineer at MetLife
October 1, 2023 - May 31, 2025Led Azure migration-focused work moving from AWS-centric pipelines to an Azure stack using ADF and Databricks. Built standardized PySpark pipelines in Databricks to unify claims data across life, disability, and group benefits lines. Worked with fraud teams to aggregate claims history and payment data into modeling-ready features. Used Delta Lake incremental loads to keep claims marts closer to real time. Collaborated with actuaries to align pipeline structures with reserve calculation consumption patterns. Consolidated claims, billing, and policy data into a unified Synapse model to replace multiple conflicting reporting sources. Implemented Python validations to detect schema drift in vendor feeds. Delivered Power BI dashboards for claims operations KPIs. Set up CI/CD using Azure DevOps for ADF pipelines and Databricks notebooks. Added governance controls for column classification and PII/PHI masking in claims data ahead of audit. Migrated legacy SSIS pipeline to ADF and documented d
Data Engineer at Advocate Health
June 1, 2021 - September 30, 2023Joined the data team during an Epic EHR rollout and built AWS-based pipelines ingesting HL7 and FHIR clinical data into a centralized warehouse. Used EMR and PySpark to process patient encounter, lab, and medication records at large scale, handling edge cases from legacy systems. Partnered with compliance to de-identify datasets for research partners in line with HIPAA Safe Harbor requirements. Consolidated data from multiple hospital/clinic systems after a merger and reconciled patient identifiers across sources. Built Redshift data marts supporting HEDIS and CMS Star Ratings. Automated payer claims and eligibility file ingestion using S3 and Lambda, replacing manual processes. Mapped EHR fields to ICD-10, CPT, and LOINC codes with clinical informatics input. Orchestrated nightly batch loads via Airflow into the warehouse. Built Tableau dashboards for readmission and length-of-stay metrics. Retested pipelines through Epic interface changes and supported audit trails/access controls fo
Data Engineer at Wayfair
January 1, 2019 - May 31, 2021Transitioned from on-prem ETL to a cloud-based AWS data stack with Glue, S3, and Redshift. Built ingestion pipelines for product catalog, pricing, and inventory feeds from suppliers into S3 as CSV/JSON and cleaned/structured them using Glue. Streamed clickstream events with Kafka into a raw layer for downstream data science. Developed PySpark jobs to aggregate daily sales and returns at scale with performance tuning. Orchestrated nightly batch pipelines using Airflow with proactive alerting. Cleaned third-party product feed inconsistencies (format mismatches and missing fields). Tuned Redshift tables using improved distribution and sort keys to reduce merchandising dashboard query times. Cataloged tables in AWS Glue Data Catalog and introduced dbt for transformation models to add testing/version control. Created Power BI reports for marketing using Google Analytics exports. Supported data science feature pipelines for churn and LTV. Participated in on-call rotation handling pipeline fa
Data Engineer / ETL Developer at ICICI Bank
August 1, 2016 - December 31, 2018Built ETL workflows in Informatica PowerCenter and SSIS to move core banking data (account balances, transactions, and loan details) from mainframe and Oracle sources into an enterprise warehouse. Wrote T-SQL and PL/SQL stored procedures for daily reconciliation between source systems and the warehouse, detecting mismatches before downstream reporting. Worked on RBI-mandated regulatory reporting extracts ensuring field-level definitions matched compliance requirements. Converted legacy flat-file feeds into structured SSIS packages to improve nightly load reliability and troubleshooting. Built SSRS reports for branch managers covering deposits, disbursements, and account activity. Gathered requirements from business analysts and produced source-to-target mapping documentation. Tested ETL jobs before releases, logged discrepancies, and coordinated with QA to resolve issues ahead of production deployment. Provided after-hours production support during month-end close and maintained data m
Education
Qualifications
Hire a Data Engineer
We have the best data engineer experts on Twine. Hire a data engineer in New York today.