Data Scientist with around 4 years of experience building production machine learning, deep learning, and Generative AI solutions using Python, SQL, Spark, Databricks, Snowflake, TensorFlow, PyTorch, and AWS. Skilled in predictive modeling, feature engineering, MLOps, and ML deployment using MLflow, Kubeflow, Airflow, Docker, Kubernetes, and SageMaker. Experienced in LLM-powered applications using LangChain, LangGraph, OpenAI, Hugging Face Transformers, and Retrieval-Augmented Generation (RAG). Proven track record delivering data-driven solutions including fraud detection and real-time risk scoring, end-to-end MLOps automation for faster deployments, and RAG systems for faster investigation workflows. Strong background across scalable data pipelines, model monitoring and drift detection, NLP embedding generation, and collaboration with cross-functional teams to deploy ML across multiple business workflows.

Mehfooz Alam

Data Scientist with around 4 years of experience building production machine learning, deep learning, and Generative AI solutions using Python, SQL, Spark, Databricks, Snowflake, TensorFlow, PyTorch, and AWS. Skilled in predictive modeling, feature engineering, MLOps, and ML deployment using MLflow, Kubeflow, Airflow, Docker, Kubernetes, and SageMaker. Experienced in LLM-powered applications using LangChain, LangGraph, OpenAI, Hugging Face Transformers, and Retrieval-Augmented Generation (RAG). Proven track record delivering data-driven solutions including fraud detection and real-time risk scoring, end-to-end MLOps automation for faster deployments, and RAG systems for faster investigation workflows. Strong background across scalable data pipelines, model monitoring and drift detection, NLP embedding generation, and collaboration with cross-functional teams to deploy ML across multiple business workflows.

Available to hire

Data Scientist with around 4 years of experience building production machine learning, deep learning, and Generative AI solutions using Python, SQL, Spark, Databricks, Snowflake, TensorFlow, PyTorch, and AWS. Skilled in predictive modeling, feature engineering, MLOps, and ML deployment using MLflow, Kubeflow, Airflow, Docker, Kubernetes, and SageMaker. Experienced in LLM-powered applications using LangChain, LangGraph, OpenAI, Hugging Face Transformers, and Retrieval-Augmented Generation (RAG).

Proven track record delivering data-driven solutions including fraud detection and real-time risk scoring, end-to-end MLOps automation for faster deployments, and RAG systems for faster investigation workflows. Strong background across scalable data pipelines, model monitoring and drift detection, NLP embedding generation, and collaboration with cross-functional teams to deploy ML across multiple business workflows.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Work Experience

Data Scientist at Capital One
April 1, 2025 - Present
Developed real-time credit card fraud detection models using Python, XGBoost, and LightGBM on Databricks and Spark, leveraging behavioral, transactional, merchant, and device features to evaluate 20M+ payment transactions per month. Built scalable feature pipelines with PySpark, SQL, Snowflake, and Databricks to enhance model performance for fraud workflows. Automated end-to-end ML workflows using AWS SageMaker, MLflow, and Airflow, improving model delivery time by 45%. Implemented production RAG using LangChain, LangGraph, OpenAI, and FAISS to retrieve and summarize fraud investigation records and policies, reducing case review time by 34%. Monitored production models with MLflow and SageMaker Model Monitor to track drift and performance stability. Applied transformer-based NLP (Hugging Face, PyTorch) for contextual embeddings from merchant/dispute data, and collaborated across teams for champion-challenger evaluations and deployment across 20+ risk decision workflows.
Data Scientist at Tatvasoft
April 1, 2021 - July 1, 2023
Built predictive patient readmission models using Python, CatBoost, scikit-learn, and SQL from 3M+ electronic health records. Designed healthcare data pipelines with Pandas, NumPy, Spark, and SQL integrating 10+ clinical sources into ML-ready datasets. Explained predictions using SHAP and feature importance, and supported transparency through statistical hypothesis testing. Deployed real-time inference services using FastAPI, AWS Lambda, and API Gateway for patient risk scoring integrated into hospital workflows. Developed deep learning models (TensorFlow/Keras CNNs and LSTMs) for spatial and temporal patterns in EHR data. Validated models using cross-validation and ROC-AUC/Precision-Recall/F1 and calibration across cohorts. Forecasted admission trends using Prophet/ARIMA with Spark and SQL across 36+ months of data. Established production MLOps with Kubeflow, Docker, Kubernetes, CI/CD, and a centralized feature store for scalable deployment.

Education

Master of Science in Data Science at Kent State University
January 11, 2030 - August 20, 2026
Master of Science in Data Science at Kent State University
January 11, 2030 - August 20, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Professional Services