Data Scientist with 5+ years of experience engineering scalable machine learning models, clinical NLP pipelines, and production-grade AI solutions across enterprise healthcare and insurance environments. Expertise in transforming EHR and claims data into clinical insights using PySpark, Databricks, Snowflake, XGBoost, and LLM/RAG frameworks. Proven track record of deploying end-to-end predictive models and automating MLOps workflows on AWS/Azure, improving risk stratification, reducing operational costs, and enabling clinical decision support through explainability, drift monitoring, and compliant data handling.

Tushar Ahuja

Data Scientist with 5+ years of experience engineering scalable machine learning models, clinical NLP pipelines, and production-grade AI solutions across enterprise healthcare and insurance environments. Expertise in transforming EHR and claims data into clinical insights using PySpark, Databricks, Snowflake, XGBoost, and LLM/RAG frameworks. Proven track record of deploying end-to-end predictive models and automating MLOps workflows on AWS/Azure, improving risk stratification, reducing operational costs, and enabling clinical decision support through explainability, drift monitoring, and compliant data handling.

Available to hire

Data Scientist with 5+ years of experience engineering scalable machine learning models, clinical NLP pipelines, and production-grade AI solutions across enterprise healthcare and insurance environments. Expertise in transforming EHR and claims data into clinical insights using PySpark, Databricks, Snowflake, XGBoost, and LLM/RAG frameworks.

Proven track record of deploying end-to-end predictive models and automating MLOps workflows on AWS/Azure, improving risk stratification, reducing operational costs, and enabling clinical decision support through explainability, drift monitoring, and compliant data handling.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

English
Advanced

Work Experience

Data Scientist at CVS Health
May 1, 2025 - Present
Architected patient risk stratification models using Python and XGBoost on 90M+ electronic health records and longitudinal claims data, improving high-risk patient identification accuracy by 29% and lowering downstream care management costs. Built clinical NLP extraction pipelines using Transformers and semantic embeddings to process unstructured physician notes, reducing manual chart review time by 58%. Developed an enterprise RAG clinical retrieval platform using LangChain, FAISS, and semantic search across 5M+ policy documents, reducing query latency by 42%. Established production MLOps governance with MLflow, SHAP interpretability, and automated drift monitoring, reducing false-positive prediction errors by 24%. Created distributed feature engineering pipelines in PySpark/Databricks/Snowflake for multi-terabyte datasets, reducing execution times by 35%. Ran A/B experiments and hypothesis testing for preventive care outreach, increasing member engagement by 18%. Partnered with opera
Data Scientist at Mindtree
January 1, 2020 - December 31, 2023
Engineered supervised ML models using Python, SQL, and LightGBM to automate insurance claim classification, improving accuracy by 31% and reducing manual review costs. Built distributed ETL pipelines with Apache Spark and SQL for semi-structured healthcare datasets, cutting data preparation times by 47%. Developed clinical NLP models with NER and text classification to parse unstructured medical records, boosting ingestion throughput by 52%. Built statistical regression frameworks to analyze healthcare utilization trends, reducing resource allocation expenses by 12%. Deployed containerized scoring pipelines via Docker on Azure ML, shortening deployment cycles by 36%. Built dashboards in Tableau/Power BI for 20+ executive stakeholders to monitor KPIs, accelerating decision-making by 30%. Implemented REST APIs for real-time inference over 50k+ daily transactions with response times under 85ms. Optimized database schemas and complex SQL in PostgreSQL warehouses (15k+ records) to reduce da

Education

Master of Science in Data Analytics and Visualization at Yeshiva University
January 1, 2024 - December 31, 2025
Bachelor in Computer Science at Chitkara University
June 1, 2018 - June 1, 2022
Master of Science in Data Analytics and Visualization at Yeshiva University
January 1, 2024 - December 31, 2025
Bachelor in Computer Science at Chitkara University
June 1, 2018 - June 1, 2022

Qualifications

Microsoft Azure Data Scientist Associate
January 11, 2030 - July 24, 2026
AWS Certified Cloud Practitioner
January 11, 2030 - July 24, 2026
Google Data Analytics Certificate
January 11, 2030 - July 24, 2026
Data Visualization with Tableau Specialization
January 11, 2030 - July 24, 2026
Python for Financial Analysis and Algorithmic Trading
January 11, 2030 - July 24, 2026
Microsoft Azure Data Scientist Associate
January 11, 2030 - July 24, 2026
AWS Certified Cloud Practitioner
January 11, 2030 - July 24, 2026
Google Data Analytics Certificate
January 11, 2030 - July 24, 2026
Data Visualization with Tableau Specialization
January 11, 2030 - July 24, 2026
Python for Financial Analysis and Algorithmic Trading
January 11, 2030 - July 24, 2026

Industry Experience

Healthcare, Financial Services, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more