Data Engineer specializing in enterprise ETL pipelines, cloud architecture, and data infrastructure optimization. Experienced in designing scalable solutions using Matillion, PySpark, and AWS, and in improving performance and operational efficiency through robust data engineering practices. Collaborates with stakeholders to translate complex business requirements into reliable technical solutions. Builds production-grade data platforms with CI/CD automation, schema evolution, and search/knowledge retrieval capabilities using AWS services and AI tooling.

Pratiksha Vyas

Data Engineer specializing in enterprise ETL pipelines, cloud architecture, and data infrastructure optimization. Experienced in designing scalable solutions using Matillion, PySpark, and AWS, and in improving performance and operational efficiency through robust data engineering practices. Collaborates with stakeholders to translate complex business requirements into reliable technical solutions. Builds production-grade data platforms with CI/CD automation, schema evolution, and search/knowledge retrieval capabilities using AWS services and AI tooling.

Available to hire

Data Engineer specializing in enterprise ETL pipelines, cloud architecture, and data infrastructure optimization. Experienced in designing scalable solutions using Matillion, PySpark, and AWS, and in improving performance and operational efficiency through robust data engineering practices.

Collaborates with stakeholders to translate complex business requirements into reliable technical solutions. Builds production-grade data platforms with CI/CD automation, schema evolution, and search/knowledge retrieval capabilities using AWS services and AI tooling.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Language

Work Experience

Data Engineer at Merkle
November 8, 2022 - Present
Architected and optimized enterprise ETL pipeline handling large amounts of growing data with automated scaling. Implemented serverless data processing workflow (S3 → EventBridge → Step Functions → EMR cluster) achieving 35% performance improvement through intelligent cluster provisioning and PySpark optimization. • Partnered closely with business stakeholders including train operators, investment banks, and financial institutions to understand complex data requirements and deliver tailored solutions that aligned with their strategic objectives. • Engineered end-to-end data pipeline infrastructure using Matillion and AWS services, enabling enterprise clients to process complex data transformations at scale with reduced operational overhead. • Established CI/CD practices using GitHub Actions and Docker for systematic testing and automated deployment of data pipeline changes. • Leveraged AWS Glue for intelligent data cataloging and schema evolution, enabling seamless integration of new data sources across enterprise systems. • Developed real-time search and knowledge retrieval system using AWS OpenSearch and Bedrock, improving data accessibility and analytics capabilities.
Data Engineer at Merkle (Dentsu) UK
November 1, 2022 - Present
Architected and optimized enterprise ETL pipelines handling large volumes of growing data with automated scaling. Implemented a serverless workflow (S3 → EventBridge → Step Functions → EMR) and improved performance by 35% using intelligent cluster provisioning and PySpark optimizations. Built end-to-end pipeline infrastructure using Matillion and AWS services to support complex transformations at scale with reduced operational overhead. Established CI/CD practices using GitHub Actions and Docker for automated testing and deployment of data pipeline changes. Used AWS Glue for data cataloging and schema evolution to integrate new sources across enterprise systems. Developed real-time search and knowledge retrieval using AWS OpenSearch and Amazon Bedrock to improve data accessibility and analytics capabilities.
AI and Data Engineer Accelerator at AiCore
May 1, 2022 - October 1, 2022
Built a data collection and processing pipeline for e-commerce analysis using Python OOP, Selenium web scraping, and Pandas transformations. Implemented an automated testing framework and configured CI/CD using GitHub Actions and Docker to enable containerized deployments.
Research Engineer Intern at InterDigital Europe Ltd
September 1, 2021 - May 1, 2022
Designed and developed RESTful APIs for a service-based platform using Flask. Implemented multiprocessing in Python to improve processing efficiency. Worked in an Agile (Kanban) environment using GitHub for version control and issue tracking. Assisted with deployment of Virtual Network Functions on OpenStack infrastructure, gaining hands-on experience with cloud infrastructure and containerization technologies.
Software Engineer at Advantmed LLP
April 1, 2016 - November 1, 2019
Designed and developed web applications using PHP and ASP.NET with relational database management (MySQL). Crafted robust APIs and web services using ASP.NET to support application integration and complex data processing workflows. Performed debugging and troubleshooting while working collaboratively within a team to deliver production features.

Education

Master of Science – Artificial Intelligence (Industrial Placement) at University of East London
September 1, 2020 - June 1, 2022
Master of Science – Information Technology at Kadi Sarva Vishwavidyalaya
June 1, 2014 - June 1, 2016
Bachelor of Science – Information Technology at Ganpat University
July 1, 2011 - June 1, 2014

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Computers & Electronics, Software & Internet, Transportation & Logistics, Professional Services