Available to hire
Hi, I’m Saffa Samreen, an AI/ML Engineer with around 4 years of experience building and deploying production-grade machine learning systems across fraud detection, NLP, search ranking, and recommendation systems. I enjoy designing large-scale ML pipelines, feature engineering, and real-time data processing using Python, SQL, Spark, and AWS.
I specialize in transformer-based NLP, graph-based modeling, and LLM-based applications including Retrieval-Augmented Generation (RAG) and agent workflows. I also focus on end-to-end MLOps—CI/CD, model registry, monitoring, and scalable deployment of ML systems in distributed production environments.
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Language
English
Fluent
Work Experience
AI Engineer at Goldman Sachs
August 1, 2025 - PresentArchitected a real-time fraud detection and transaction monitoring system processing 30M+ financial transactions using Python, SQL, and distributed feature engineering pipelines for behavioral anomaly detection in streaming data. Built ensemble ML models (XGBoost, Random Forest, LightGBM) for fraud and insider trading detection, improving fraud detection precision by 18–25% and reducing false positives by 20% across real-time trading systems. Designed graph-based learning models using Graph Neural Networks to identify fraud rings, increasing detection recall of coordinated fraud clusters by 22% across multi-hop transaction networks. Implemented domain-adapted LLM fine-tuning using LoRA-based methods on financial fraud datasets, reducing hallucination rate by 30% and improving classification consistency across high-volume investigation workflows. Developed NLP-based fraud intelligence system using BERT and embeddings to process 2M+ unstructured documents, improving signal extraction e
Data Scientist at Goldman Sachs
August 1, 2025 - PresentLed development of a production-scale Retrieval-Augmented Generation platform integrating LLaMA-3, LangChain, FAISS, and PostgreSQL to process 50M+ research documents for grounded outputs; improved analyst turnaround by 38%. Fine-tuned GPT-4 APIs with PEFT and LoRA to automate memo drafting and compliance workflows, reducing documentation overhead by 41%. Built multi-asset risk and stress-testing models with XGBoost and Monte Carlo simulations to enhance downside risk prediction. Implemented real-time market anomaly detection with PyTorch, Kafka, and Spark Structured Streaming, reducing false positives by 26% and accelerating regulatory alerts by 32%. Created document intelligence pipelines with spaCy and ONNX Runtime for 2M+ filings (90% precision). Established MLOps with MLflow, Docker, Kubernetes, and GitHub Actions for 20+ production models with 99.5% uptime. Delivered Tableau dashboards with Snowflake for risk metrics and AI insights, reducing reporting latency by 27%.
AI/ML Engineer at LTIMindTree
April 1, 2023 - July 1, 2024Engineered NLP chatbot using PyTorch and Hugging Face Transformers (BERT, RoBERTa) for intent classification and entity recognition across 1M+ HR and customer support queries. Built NLP feature pipelines including tokenization, contextual embeddings, and attention-based representations to improve intent classification stability in production workloads. Designed RAG-based retrieval system using FAISS and Pinecone vector databases for semantic search over 400K+ enterprise documents enabling context-aware responses. Developed LLM-based conversational agents using GPT APIs with structured prompting and agent workflows for reasoning and automation. Implemented FastAPI-based model serving layer for NLP models with optimized inference latency and integration into enterprise CRM systems for real-time usage. Built end-to-end MLOps pipelines on AWS with CI/CD, automated training workflows, and MLflow model versioning, reducing deployment cycles by 40% and improving experiment reproducibility acr
Data Scientist at LTIMindTree
April 1, 2023 - July 1, 2024Architected end-to-end retail analytics for a large e-commerce client, delivering churn prediction and demand forecasting with XGBoost, LightGBM, CatBoost, and Scikit-learn; improved customer retention by 22% and forecast accuracy by 18%, contributing to a 12% quarterly revenue uplift. Built NLP pipelines with Transformer models (BERT, GPT-based) analyzing 2M+ reviews and transcripts, achieving 91% F1 in sentiment/intents. Developed Generative AI-powered retail assistant (LangChain + GPT) for contextual product recommendations and automated content creation, boosting engagement by 30% and conversions by 14%. Designed scalable time-series forecasting pipelines (ARIMA, Prophet) with ensemble methods to optimize inventory, reducing stock-outs by 26% and excess inventory costs by 15%. Deployed ML/Generative AI models via Docker/REST APIs with CI/CD, enabling automated monitoring, drift detection, and retraining, reducing deployment time by 35%.
Data Scientist at Accenture
January 1, 2022 - March 1, 2023Built learning-to-rank based e-commerce search system using TensorFlow and Scikit-learn (LambdaMART, XGBoost Ranker), improving ranking relevance across 80K+ daily queries in production systems. Engineered large-scale feature pipelines using PySpark and SQL incorporating clickstream behavior, user intent signals, and engagement metrics across distributed datasets. Developed deep learning ranking models using embedding-based architectures, improving semantic understanding of queries and increasing CTR by 18–25% in production systems. Implemented A/B testing framework for ranking models enabling controlled evaluation of model variants on live traffic and optimization. Designed MLOps pipelines using MLflow and Kubernetes-based workflows on AWS enabling automated training, experiment tracking, and deployment of ranking models. Containerized ML inference services using Docker and Kubernetes to deploy low-latency ranking APIs for real-time product search under traffic. Built real-time ETL
Junior Data Scientist at Accenture
January 1, 2022 - March 1, 2023Performed EDA on large-scale healthcare datasets (EHR, claims) using Python and SQL, improving data quality and consistency. Built predictive pipelines (Logistic Regression, Random Forest, XGBoost) to predict readmission with 21% accuracy gains and supported hospital risk stratification. Implemented NLP on physician notes (NLTK, TF-IDF, Word2Vec; LSTM-based models) achieving 87% F1. Developed CNN-based medical image classification models (pneumonia detection) with 92% validation accuracy. Enabled production-ready ML deployment via REST APIs (Flask/FastAPI) with monitoring; ensured HIPAA-aligned governance in AWS cloud environments.
Education
Master of Science in Data Analytics at University of Illinois Springfield
January 11, 2030 - March 27, 2026Master of Science in Data Analytics at University of Illinois Springfield
January 11, 2030 - March 27, 2026Master of Science in Data Analytics at University of Illinois Springfield
January 11, 2030 - June 29, 2026Qualifications
Industry Experience
Financial Services, Retail, Healthcare, Software & Internet, Professional Services
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Hire a Data Scientist
We have the best data scientist experts on Twine. Hire a data scientist in Chicago today.