Available to hire
Data Engineer with 3+ years of experience designing and supporting scalable ETL/ELT pipelines, cloud data platforms, and enterprise data solutions. Skilled in Python, SQL, PySpark, Databricks, Snowflake, and Apache Spark, along with dimensional modeling (Star Schema, Snowflake Schema, SCD Type 1/2) to deliver consistent analytics.
Experienced with AWS (S3, Glue, Lambda, Redshift, CloudWatch, IAM), Apache Airflow, and dbt for integration, orchestration, automation, and data quality. Hands-on with REST APIs, pipeline monitoring, CI/CD, and Agile delivery to improve reliability, performance, and near real-time data availability for business stakeholders.
Skills
See more
Language
Work Experience
Data Engineer at Zebra Technologies
January 1, 2026 - PresentDesigned and maintained PySpark and SQL data pipelines in Databricks, processing over 5M records daily from ERP, IoT, and operational systems for enterprise reporting and analytics. Developed scalable ELT workflows using AWS Glue, S3, and Snowflake, reducing end-to-end processing time by 32% while improving pipeline reliability. Built reusable dbt models and implemented dimensional data models for consistent reporting across 15 business datasets. Automated workflow scheduling and dependency management using Apache Airflow, achieving 98% pipeline success rates with retry policies and proactive monitoring. Integrated REST APIs, JSON, Parquet, and relational databases to enable near real-time data availability with 40% faster ingestion. Implemented data quality validations that reduced production data issues by 30%, and optimized Snowflake performance via partitioning and clustering, reducing query runtime by 28%. Collaborated with stakeholders and delivered production enhancements while
Data Engineer at Zebra Technologies, IL, USA
January 1, 2026 - PresentDesigned and maintained PySpark and SQL data pipelines in Databricks, processing over 5M records daily from ERP, IoT, and operational systems for enterprise reporting and analytics. Developed scalable ELT workflows using AWS Glue, S3, and Snowflake, reducing end-to-end processing time by 32% while improving pipeline reliability. Built reusable dbt models and dimensional data models to support consistent reporting across 15 business datasets and reduce duplicate transformation logic. Automated scheduling and dependency management using Apache Airflow, achieving 98% success rates through retries and proactive monitoring. Integrated REST APIs, JSON, Parquet, and relational databases into centralized data platforms for near real-time availability, with 40% faster ingestion. Implemented data quality validation and reconciliation checks, reducing production data issues by 30%. Optimized Snowflake performance using partitioning and clustering, decreasing analytical query runtime by 28%. Colla
Data Engineer Intern at Labelmaster
January 1, 2025 - May 31, 2025Developed Python and SQL ETL workflows to cleanse, transform, and load logistics data, improving processing efficiency by 25%. Built AWS Glue pipelines integrating data from S3 and relational databases, reducing manual ingestion effort by 30%. Performed data profiling, validation, and reconciliation across operational datasets to improve data quality for downstream analytics. Created Power BI dashboards using curated warehouse datasets to monitor inventory, shipping performance, and operational KPIs. Collaborated with data engineers on automated pipeline deployments using Git, CI/CD practices, REST APIs, and Agile development workflows.
Data Engineer Intern at Labelmaster, IL, USA
January 1, 2025 - May 31, 2025Developed Python and SQL ETL workflows to cleanse, transform, and load logistics data, improving processing efficiency by 25%. Built AWS Glue pipelines integrating data from S3 and relational databases, reducing manual ingestion effort by 30%. Performed data profiling, validation, and reconciliation across operational datasets to improve data quality for downstream analytics. Created Power BI dashboards using curated warehouse datasets to monitor inventory, shipping performance, and operational KPIs. Collaborated with data engineers on automated pipeline deployments using Git, CI/CD practices, REST APIs, and Agile development.
Junior Data Engineer at LTIMindtree
December 1, 2020 - July 31, 2023Enhanced and supported Python and SQL ETL workflows to integrate data from 12 enterprise source systems, improving reporting and analytics data availability. Built transformation logic using PySpark and Apache Spark to process structured and semi-structured data exceeding 8M records per day, reducing batch execution time by 25%. Assisted with ingestion pipelines using AWS S3, Glue, and Redshift, increasing systematic data loads by 35% while minimizing manual intervention. Conducted data validation, profiling, and reconciliation, improving data accuracy by 22% and ensuring alignment with business quality standards. Wrote/optimized SQL queries, stored procedures, and views, reducing report generation time by 30%. Supported production deployments, monitored pipelines using CloudWatch, and resolved ETL failures to maintain 97% on-time completion. Contributed to Git-based version control, code reviews, testing, and release activities across 40 production deployments with minimal post-releas
Junior Data Engineer at LTIMindtree, Hyderabad, India
December 1, 2020 - July 31, 2023Enhanced and supported Python and SQL ETL workflows to integrate data from 12 enterprise source systems, improving availability for reporting and analytics. Built transformation logic using PySpark and Apache Spark to process structured and semi-structured datasets exceeding 8M records per day, reducing batch execution time by 25%. Assisted with ingestion pipelines using AWS S3, Glue, and Redshift, increasing systematic data loads by 35% while minimizing manual intervention. Performed data validation, profiling, and reconciliation, improving data accuracy by 22% and supporting business quality standards. Wrote and revamped SQL queries, stored procedures, and views, reducing report generation time by 30%. Supported production deployments, monitored pipeline execution using CloudWatch, and resolved ETL failures, maintaining 97% on-time completion. Contributed to Git-based version control, code reviews, testing, and releases, enabling 40 production deployments with minimal post-release is
Education
Master of Science in Data Science at Illinois Institute of Technology, Chicago, IL
January 1, 2025 - December 31, 2025Master of Science in Data Science at Illinois Institute of Technology
January 11, 2030 - December 1, 2025Master of Science in Data Science at Illinois Institute of Technology
December 1, 2025 - December 1, 2025Qualifications
Industry Experience
Computers & Electronics, Software & Internet, Transportation & Logistics, Professional Services, Manufacturing
Skills
See more
Hire a Data Engineer
We have the best data engineer experts on Twine. Hire a data engineer today.