Data Engineer with 5+ years of experience building cloud data platforms and batch/real-time pipelines across finance, healthcare, and enterprise environments. Hands-on with Python, SQL, PySpark, Kafka, Spark, Databricks, Snowflake, AWS, Airflow, and dbt, modernizing legacy ETL and optimizing distributed workloads at scale. Also experienced in integrating RAG and LLM systems with enterprise data for secure AI-powered search, knowledge retrieval, and analytics. Skilled in building analytics-ready data platforms with strong data quality, governance, orchestration, and production reliability.

Priyam Deepak Choksi

Data Engineer with 5+ years of experience building cloud data platforms and batch/real-time pipelines across finance, healthcare, and enterprise environments. Hands-on with Python, SQL, PySpark, Kafka, Spark, Databricks, Snowflake, AWS, Airflow, and dbt, modernizing legacy ETL and optimizing distributed workloads at scale. Also experienced in integrating RAG and LLM systems with enterprise data for secure AI-powered search, knowledge retrieval, and analytics. Skilled in building analytics-ready data platforms with strong data quality, governance, orchestration, and production reliability.

Available to hire

Data Engineer with 5+ years of experience building cloud data platforms and batch/real-time pipelines across finance, healthcare, and enterprise environments. Hands-on with Python, SQL, PySpark, Kafka, Spark, Databricks, Snowflake, AWS, Airflow, and dbt, modernizing legacy ETL and optimizing distributed workloads at scale.

Also experienced in integrating RAG and LLM systems with enterprise data for secure AI-powered search, knowledge retrieval, and analytics. Skilled in building analytics-ready data platforms with strong data quality, governance, orchestration, and production reliability.

See more

Experience Level

Language

Work Experience

Data Engineer at Intuit
February 1, 2026 - Present
Architected a cloud-native Lakehouse and Snowflake analytics platform using Databricks, PySpark, Delta Lake, AWS S3, AWS Glue, and Snowflake, consolidating finance, payments, customer, accounting, and product usage data from 25+ enterprise systems for analytics and executive reporting across 200+ stakeholders. Engineered real-time ingestion and transformation pipelines using Kafka, Spark Structured Streaming, AWS Lambda, and Airflow to process customer transactions, payment events, tax filings, and product activity, reducing end-to-end latency from 95 minutes to under 18 minutes while handling 15M+ events daily. Implemented vector-based knowledge retrieval with embeddings, metadata filtering, and RAG for secure enterprise search across structured and unstructured datasets. Partnered with AI Platform, Data Science, and Product Engineering to build an AI-powered financial knowledge assistant (OpenAI, LangChain, RAG, AWS Bedrock), reducing incident resolution time from 3 hours to under 70
Graduate Teaching Assistant - Business Process Engineering at Northeastern University
April 1, 2025 - December 31, 2025
Supported graduate-level Business Process Engineering courses (INFO 7374 and INFO 7260), coordinating assignment design, grading, and student feedback for 150+ students. Facilitated process-mapping and workflow-analysis activities to help students identify operational bottlenecks and translate observations into structured improvement recommendations. Introduced AI-assisted process analysis concepts (prompt engineering and LLM workflows) to demonstrate how generative AI can summarize process documentation and support early-stage workflow redesign with human validation. Coordinated industry guest-speaker activities and provided instructional support across the academic term.
Research Data Analyst Co-op at Harvard Medical School, Brigham and Women's Hospital
September 1, 2024 - December 31, 2024
Rebuilt a genomic pipeline from 48-hour batch cycles to 15-minute distributed runs, replacing a sequential CSV-and-notebook workflow with PySpark on a Slurm HPC cluster across 5.8TB of UK Biobank, MGB EHR, and metabolomic data; used Python schema validation to catch integrity failures across 133M rows before publication. Consolidated 220+ clinical features from 6 heterogeneous sources into a curated analytics-ready egress layer, cutting weekly research preparation by 85% and migrating 2TB+ of CSV/Excel data to Parquet to reduce storage/query overhead. Delivered ensemble models (XGBoost, Random Forest) over a 48,628-participant cohort, achieving 96% AUC and surfacing 6 novel biomarkers supporting peer-reviewed publications and an ESC Young Investigator Award. Built HIPAA-compliant clinical data pipelines in SQL, Python, Pandas, and REDCap for 500,000+ patient observations. Designed a RAG assistant over the research data workflow, cutting manual data-curation effort by 80% and enabling c
Senior Data Engineer at Heeva Infra
May 1, 2021 - July 31, 2023
Modernized a legacy PostgreSQL and Excel reporting stack into a version-controlled dbt and Airflow platform, replacing 217 ad hoc stored procedures with staged/testing models and an ETL pipeline from SAP into Snowflake, reducing deployment cycles from 3 days to 4 hours. Engineered real-time Kafka and AWS Lambda pipelines to process 17.5M monthly transactions into Amazon Redshift with sub-5-minute freshness by tuning sort keys and distribution styles across 8 high-volume tables to reduce query execution time and compute costs. Modeled a Kimball star-schema warehouse across 18 PostgreSQL/MySQL and REST/gRPC sources, automating 50+ dbt models through Airflow and Great Expectations checks while reducing pipeline failures from 12 to 2 and retiring Excel reporting for 3 departments. Recovered $894K across a 12-client portfolio using Python ML cost-forecasting models to identify 7 engagements 3 weeks before budget overruns. Collaborated in Agile/Scrum with software engineers, QA analysts, and
Data Engineer at CodeNest Solution
August 1, 2019 - April 30, 2021
Built scalable ETL pipelines using Python, SQL, Talend, Apache Spark, and MySQL to ingest and transform data from ERP systems, CRM platforms, and REST APIs; processed 750+ GB of business data each week and supported centralized reporting across multiple client engagements. Developed dimensional data models and optimized SQL transformation workflows for sales, customer, and finance datasets, reducing dashboard refresh time from 55 minutes to under 20 minutes for 80+ operational/executive reports. Integrated structured and semi-structured data from relational databases, flat files, and third-party APIs using Python, SQL, Apache NiFi, and batch frameworks across 15+ production data sources to improve reliability. Implemented automated data validation and reconciliation to validate 1.2M+ records per production cycle, improving data accuracy and reducing downstream analytics issues. Enhanced pipeline performance with SQL optimization, indexing, partitioning, and Spark tuning, reducing ETL r

Education

Master of Science in Information Systems at Northeastern University
January 1, 2025 - December 31, 2025
Bachelor of Science in Information Technology at University of Mumbai
January 1, 2017 - May 31, 2021
Master of Science in Information Systems at Northeastern University
December 1, 2025 - December 1, 2025
Bachelor of Science in Information Technology at University of Mumbai
January 1, 2021 - May 1, 2021

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Professional Services, Software & Internet, Education