Senior Data Scientist with 4.5+ years of experience delivering AI/ML solutions in fintech and customer analytics. Expertise includes predictive modeling, NLP, Generative AI, RAG pipelines, and end-to-end data engineering. Skilled in Python, PyTorch, SQL, feature engineering, LLM-based systems, and cloud-native ML workflows. Strong focus on MLOps, model interpretability, business intelligence (Power BI/Tableau), and building scalable analytics platforms with measurable impact.

Alekhya Dawat

Senior Data Scientist with 4.5+ years of experience delivering AI/ML solutions in fintech and customer analytics. Expertise includes predictive modeling, NLP, Generative AI, RAG pipelines, and end-to-end data engineering. Skilled in Python, PyTorch, SQL, feature engineering, LLM-based systems, and cloud-native ML workflows. Strong focus on MLOps, model interpretability, business intelligence (Power BI/Tableau), and building scalable analytics platforms with measurable impact.

Available to hire

Senior Data Scientist with 4.5+ years of experience delivering AI/ML solutions in fintech and customer analytics. Expertise includes predictive modeling, NLP, Generative AI, RAG pipelines, and end-to-end data engineering.

Skilled in Python, PyTorch, SQL, feature engineering, LLM-based systems, and cloud-native ML workflows. Strong focus on MLOps, model interpretability, business intelligence (Power BI/Tableau), and building scalable analytics platforms with measurable impact.

See more

Work Experience

Data Scientist at Goldman Sachs
May 1, 2025 - Present
Led development of a firmwide GS AI Assistant generative AI platform to summarize documents, draft reports, and perform data analysis, reducing manual workload by 25%. Designed and implemented RAG pipelines with LangChain and FAISS/Pinecone integrating structured and unstructured financial data for context-aware responses. Conducted LoRA/PEFT fine-tuning on proprietary financial corpora, improving domain Q&A accuracy by 15% while reducing GPU usage and inference latency. Built automated evaluation frameworks for model performance, robustness, fairness, and compliance, including custom regulatory risk and bias metrics. Developed containerized LLM-based microservices with Docker/Kubernetes, enabling version control, auto-scaling, and experiment tracking via MLflow/W&B. Engineered embedding pipelines with vector databases to accelerate retrieval from legacy reports and market datasets. Partnered with governance/compliance to implement explainable AI workflows using SHAP/LIME, enforce anon
Data Scientist at Capital One
May 1, 2021 - December 31, 2023
Architected end-to-end credit risk and underwriting models using Python, SQL, and Spark across transactional, bureau, behavioral, and alternative data; improved default prediction accuracy by 20% over legacy scorecards and supported credit line optimization. Programmed and optimized 150+ features (e.g., transaction velocity, revolving utilization trends, delinquency signals) using Pandas and statistical techniques (WOE/IV scoring) to improve KS and ROC-AUC. Built scalable real-time decisioning on AWS with Spark Structured Streaming, feature stores, and REST microservices for sub-200ms inference latency for approvals and personalized offers. Designed experimentation strategy using A/B testing, uplift modeling, and Bayesian inference, delivering 12% incremental revenue lift and 18% improved response rates. Built a LangChain + Pinecone GenAI internal assistant for semantic search and RAG over documents/underwriting cases, reducing analyst query resolution time by 35% and improving accessi
Junior Data Scientist at LTIMindtree
January 1, 2020 - April 30, 2021
Built end-to-end Customer 360 pipelines integrating EHR, CRM, and marketing data using Python for cleaning/transformation and feature engineering; reduced data retrieval latency by 40% and enabled near real-time analytics. Performed statistical analysis and EDA (hypothesis testing, correlation modeling, cohort analysis) to identify clinical/behavioral risk factors and created baseline ML models (logistic regression, random forest, gradient boosting). Orchestrated PyTorch-based segmentation and risk prediction, using CNNs for structured feature learning and LSTMs for longitudinal time-series modeling to improve targeted intervention success by 18%. Created NLP pipelines for unstructured clinical notes and engagement logs using TF-IDF, embeddings, and LSTM sequence models to enhance risk scoring and engagement forecasting. Developed a personalized recommendation engine (collaborative filtering + hybrid ML) with scikit-learn and deep embeddings, increasing cross-service adoption by 15% wh

Education

Master of Science in Computer Science at California State University, Fullerton
January 11, 2030 - July 8, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Software & Internet, Healthcare