Available to hire
I’m Shrey Patel, an AI Data Engineer with around 4 years of experience designing cloud-native, AI-driven data platforms. I specialize in building scalable ETL/ELT pipelines and real-time streaming architectures using Azure Data Factory, Databricks, Event Hubs, AWS Glue, Lambda, and Kinesis to empower data-driven decision making.
I enjoy turning complex data into actionable insights, applying ML and anomaly detection, and ensuring governance and security across AWS and Azure. I thrive on collaboration with cross-functional teams to deliver cost-effective solutions, maintain reliability, and drive continuous performance optimization.
Language
English
Fluent
Work Experience
AI Engineer at Walmart
November 1, 2024 - PresentArchitected and deployed enterprise-scale Azure Lakehouse architecture processing 5+ TB of structured and unstructured data daily across behavioral, financial, and telemetry domains, enabling scalable analytics and AI workloads supporting $500M+ annual transaction ecosystems. Designed and optimized end-to-end ETL/ELT pipelines using Azure Data Factory and Databricks (PySpark) with Delta Lake, reducing data latency by 65% and accelerating executive reporting cycles from 48 hours to under 6 hours. Implemented a multi-agent underwriting workflow in LangGraph, automated ~40% of routine pre-underwriting checks, and stood up an LLM evaluation harness with CI/CD governance. Optimized storage and analytics through partitioning, Z-order indexing, and query tuning, boosting enterprise BI performance by 55% and reducing compute costs. Led Infrastructure-as-Code and DevOps automation with Terraform and GitHub Actions, enforcing RBAC and regulatory compliance across regulated datasets.
Software Data Engineer at IBM
January 1, 2021 - July 1, 2023Architected and optimized AWS Glue + PySpark ETL pipelines processing 100M+ financial transactions annually (multi-terabyte datasets), achieving 99.9% data accuracy, 45% faster processing, and 30% reduction in infrastructure costs through Spark optimization and partition tuning. Designed and implemented a scalable Amazon Redshift data warehouse integrating 15+ cross-functional systems, enabling C-suite analytics and dashboards. Engineered event-driven, serverless ingestion using AWS Lambda, S3, and Step Functions, automating batch and streaming workflows and improving pipeline reliability to 99.95% SLA. Built real-time transaction monitoring with AWS Kinesis + Spark Streaming for near real-time fraud detection; deployed fraud scoring and churn models on SageMaker; developed REST microservices on Kubernetes and created dashboards for risk analysts.
Education
Master of Science in Computer Software Engineering Systems at Northeastern University
January 11, 2030 - June 30, 2026Qualifications
OCI Certified Multicloud Architect Professional
January 11, 2030 - June 30, 2026OCI Certified Data Science Professional
January 11, 2030 - June 30, 2026OCI Certified Generative AI Professional
January 11, 2030 - June 30, 2026NVIDIA Certified Professional: Gen AI LLMs
January 11, 2030 - June 30, 2026Databricks Certified Machine Learning Professional
January 11, 2030 - June 30, 2026Industry Experience
Software & Internet, Retail, Financial Services
Hire a Data Scientist
We have the best data scientist experts on Twine. Hire a data scientist today.