Available to hire
I am a data engineer who enjoys turning complex data into reliable, ML-ready datasets and scalable pipelines. I thrive in cross-functional teams where collaboration between data scientists and engineers accelerates the journey from experimentation to production, delivering measurable business impact.
I’m passionate about AI and ML capabilities, from feature engineering and model deployment to retrieval-augmented generation and governance. I continuously seek ways to optimize performance, embrace scalable cloud-native solutions, and share reusable frameworks that expedite analytics and ML initiatives across industries.
Skills
Work Experience
Data Engineer at Velocity Global
October 1, 2024 - October 1, 2024Designed ML-ready data pipelines in Azure Databricks with PySpark to cleanse payroll and compliance data from multiple countries, creating high-quality training datasets for predictive workforce and cost models. Engineered feature pipelines combining structured HR/payroll data with semi-structured JSON and unstructured compliance documents to enable richer feature sets and improved model accuracy. Implemented automated validation and transformation pipelines in Delta Lake with schema checks and anomaly detection for consistent training data. Collaborated with data scientists to operationalize ML models in Azure ML, building retraining pipelines with experiment tracking, hyperparameter tuning, and version control. Developed Python-based APIs and microservices to serve ML predictions within HR dashboards and compliance reports. Optimized Spark jobs and SQL transformations to reduce training and inference times, leveraging distributed compute and caching. Built semantic SQL models in Syna
Senior Data Engineer at Premier Inc
October 1, 2024 - November 25, 2025Designed and productionized PySpark-based data pipelines on Azure Databricks to cleanse, normalize, and transform HL7/FHIR healthcare datasets, producing ML-ready data used for predictive models across 100+ hospitals. Built Delta Lake ingestion pipelines on ADLS Gen2 with audit logging, lineage tracking, and validation layers to ensure reproducibility and compliant ML workloads. Collaborated with data scientists to engineer features from structured, semi-structured, and unstructured data (EHR, claims, IoT streams, text) for clinical risk and cost models. Automated model training and retraining pipelines using MLflow and Azure ML, with experiment tracking and version control. Implemented real-time streaming pipelines with Event Hub and Spark Streaming to feed anomaly detection and early warning systems. Deployed APIs and microservices for model serving integrated into clinical dashboards and workflows. Applied governance with Purview and RBAC for auditable, explainable outputs. Optimize
Data Engineer at USAA
April 1, 2022 - April 1, 2022Developed real-time fraud detection models in Azure Synapse and Snowflake, combining SQL aggregations, anomaly detection queries, and PySpark feature engineering to improve fraud alert precision and reduce false positives by 20%. Built Python ETL pipelines in Databricks and Azure Data Factory, ingesting banking transactions from SQL Server, SAS, and IoT streams into ADLS Gen2 and Snowflake to create ML-ready datasets. Collaborated with data scientists to support predictive credit risk models, preparing features in PySpark and integrating outputs into BI and fraud workflows. Designed metadata-driven ELT frameworks with audit tables and reconciliation scripts to ensure regulatory compliance (PCI DSS, HIPAA). Optimized PySpark jobs and SQL transformations in Databricks and Synapse, tuning caching and partitioning for faster data refresh. Operated Python-based anomaly detection workflows with validations and statistical checks to improve early fraud signal detection. Implemented CI/CD in A
Data Analyst at LPL Financial
September 1, 2019 - September 1, 2019Developed and optimized SQL queries for reconciliation and reporting across Oracle, Teradata, and AWS Redshift, improving execution times by 35% and the trustworthiness of financial reports. Automated reporting workflows with Python to transform and validate large transactional datasets, reducing manual reconciliation tasks by 60% and increasing accuracy for monthly compliance reports. Migrated legacy reporting from SQL Server/Teradata into Redshift and Power BI, delivering dynamic dashboards with drill-through analysis. Identified patterns in investment transactions for compliance monitoring and provided mentorship on SQL optimization and Python scripting in Agile teams. Validated third-party datasets and created metadata-driven audit logs for Power BI dataset refreshes to enhance data lineage and auditability.
Education
Bachelor of Science in Computer Science at Texas Woman’s University
January 11, 2030 - November 25, 2025Qualifications
Microsoft Certified: Fabric Data Analyst Associate (DP-700)
August 1, 2025 - August 1, 2026Industry Experience
Financial Services, Healthcare, Software & Internet, Professional Services
Skills
Hire a Data Analyst
We have the best data analyst experts on Twine. Hire a data analyst in Euless today.