Available to hire
I’m a data engineer specializing in scalable data platforms and ML-enabled data pipelines. I love turning complex data into reliable, high-performance solutions that power AI-driven decision making.
I thrive in collaborative environments across finance, retail, social media, and tech, delivering production-grade data solutions with strong data quality, governance, and measurable business impact.
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Language
English
Fluent
Work Experience
AI & ML Data Engineer at Meta
October 1, 2024 - PresentOwned end-to-end ML data pipelines, improving data freshness and retraining speed by optimizing large-scale pipelines processing ~50 TB of training data daily; built real-time recommendation frameworks with feature stores and low-latency streaming; unified multi-modal metadata across Facebook and Instagram using PySpark and TFT; achieved 99.9% pipeline reliability and reduced operational overhead by 40% via orchestration, dependency management, retries, and monitoring; accelerated ML development with Python-based data processing frameworks for automated feature engineering, validation, and QA; mentored peers and enforced CI/CD best practices for data workflows.
Data Engineer at Goldman Sachs
September 1, 2023 - October 1, 2024Designed and implemented high-performance Snowflake data models with partitioning and clustering; built a modular cloud-native data platform using Delta Lake on S3, Spark on AWS EMR, and Apache Hudi to support batch and streaming workloads; automated CI/CD using Jenkins, Docker, Kubernetes (EKS), and Terraform; developed 30+ ETL/ELT pipelines with AWS Glue for data migration from Oracle and SFTP to S3; optimized Python/SQL pipelines, improving data processing times by ~50%; established data quality frameworks and ML model tracking with MLflow; delivered scalable recommendation prototypes using TensorFlow, Airflow, and PostgreSQL.
Data Engineer at Cognizant
February 1, 2019 - August 1, 2021Designed and maintained 50+ scalable ETL pipelines on Microsoft Azure (Azure Data Factory, Databricks, Synapse) with cost-savings through partitioning and resource scaling; built Delta Lake architectures with PySpark Autoloader and Structured Streaming for near real-time IoT and retail analytics with 99.9% data accuracy; modernized legacy SQL Server environments with partitioning and indexing; delivered large-scale big data solutions using Hadoop ecosystem (HDFS, Hive, HBase); deployed sharded MongoDB clusters to improve read/write performance; created feature extraction workflows supporting ML model training; automated shell scripting to enhance operations.
Education
Master’s in Business Analytics at University of Texas at Dallas
January 11, 2030 - June 29, 2026Master’s in Business Analytics at University of Texas at Dallas, Richardson, TX, USA
January 11, 2030 - June 29, 2026Qualifications
Industry Experience
Financial Services, Software & Internet, Media & Entertainment, Retail, Other, Professional Services
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Hire a Data Scientist
We have the best data scientist experts on Twine. Hire a data scientist in Richardson today.