Senior Data Scientist who architects scalable AI/ML solutions and production-grade pipelines, with a focus on secure, audit-ready data engineering and reliable enterprise deployments. Builds predictive and generative AI systems (RAG/NLP/LLMs), leads MLOps on Kubernetes, and ensures HIPAA/compliance-aligned workflows across healthcare, financial services, and retail analytics.

DEEPAK SWAMINATHAN

Senior Data Scientist who architects scalable AI/ML solutions and production-grade pipelines, with a focus on secure, audit-ready data engineering and reliable enterprise deployments. Builds predictive and generative AI systems (RAG/NLP/LLMs), leads MLOps on Kubernetes, and ensures HIPAA/compliance-aligned workflows across healthcare, financial services, and retail analytics.

Available to hire

Senior Data Scientist who architects scalable AI/ML solutions and production-grade pipelines, with a focus on secure, audit-ready data engineering and reliable enterprise deployments.

Builds predictive and generative AI systems (RAG/NLP/LLMs), leads MLOps on Kubernetes, and ensures HIPAA/compliance-aligned workflows across healthcare, financial services, and retail analytics.

See more

Experience Level

Work Experience

Senior Data Scientist at Morgan Stanley
April 1, 2023 - Present
Built enterprise data pipelines using Snowflake and Python for secure ingestion and financial forecasting, improving throughput by 40% while preserving auditability. Developed predictive risk/fraud modeling frameworks using Python and SQL to improve fraud detection accuracy by 25%. Optimized distributed NoSQL storage for real-time market data, reducing ETL/ELT latency by 30%. Led MLOps for inference services using Docker and Kubernetes to achieve 99.9% uptime and eliminate environment-related failures. Modernized time-series ML workflows by migrating monolithic tasks to cloud-native platforms, improving deployment velocity by 50%. Ensured reproducible experimentation with logging for compliance tracking, and monitored deep learning training via TensorBoard to improve production convergence and churn prediction by 20%. Deployed Keras/TensorFlow models via TensorFlow Serving and REST APIs, and implemented PyTorch architectures and boosting models (XGBoost/LightGBM) to improve accuracy/pr
Senior Data Scientist at Morgan Stanley, New York, NY
April 1, 2023 - Present
Built enterprise-grade secure data ingestion pipelines using Snowflake and Python to improve throughput while maintaining audit standards. Developed predictive risk and fraud detection models using Python and SQL, improving detection accuracy and strengthening risk management. Optimized real-time NoSQL data handling (MongoDB/Cassandra) to reduce latency for market analytics. Orchestrated production MLOps services with Docker and Kubernetes to achieve 99.9% uptime for critical inference APIs, and accelerated ML workflow deployment by migrating monolithic tasks to cloud-native platforms. Implemented compliance-oriented experiment tracking and governance, and deployed TensorFlow/Keras and PyTorch models for scalable REST inference. Enhanced forecasting/ranking models (including XGBoost/LightGBM), and delivered executive reporting via visualization and Power BI automation. Strengthened security governance using Azure AD/RBAC and supported large-scale analytics with Spark/Hadoop integrati
Data Scientist at Best Buy
January 1, 2021 - March 31, 2023
Automated Marketing Mix Modeling reporting using RMarkdown and secure data summaries, saving 10 hours/week. Created SQL-based ETL workflows for multi-source retail marketing time-series forecasting, reducing preparation time by 25%. Built customer acquisition segmentation models using SVM/Random Forest to improve validation outcomes and targeting efficiency by 15%. Prototyped PyTorch deep learning forecasting layers to accelerate research-to-prototype cycles by 35%. Migrated legacy pipelines to serverless workflows on AWS Lambda and Step Functions, reducing operational overhead by 40%. Implemented ensemble learning for stability (20% boost), tracked deep-learning loss and hyperparameters in TensorBoard, and optimized PyTorch serialization for lower inference latency by 50%. Integrated Spark pipelines with data catalogs for lineage/auditability. Deployed Kubernetes manifests for environment parity and production reliability, and automated reporting refresh via Power BI/SQL connections.
Data Scientist at Best Buy, Richfield, Minnesota
January 1, 2021 - March 31, 2023
Automated marketing reporting and Marketing Mix Modeling insights using RMarkdown, reducing manual reporting time. Built SQL-based ETL workflows to clean multi-source retail marketing data for time series forecasting, improving reliability and reducing preparation effort. Developed ML models for customer segmentation and acquisition ROI using SVM/Random Forests. Prototyped and deployed deep learning approaches for retail demand forecasting using PyTorch/TensorFlow and ensemble methods. Migrated legacy pipelines to serverless architectures with AWS Lambda and Step Functions, reducing operational overhead. Improved data lineage/auditability for Spark pipelines, and delivered scalable analytics configurations using Redshift/EMR. Used Kubernetes for environment parity, established CI/CD validation with Jenkins, and optimized cluster memory/tuning for batch analytics. Maintained auditable codebases with Git and automated refresh/reporting for Power BI.
Data Engineer at Fortis Healthcare
October 1, 2017 - November 30, 2020
Developed secure ingestion pipelines using Azure Data Factory and SQL for clinical/patient data migration with 100% data fidelity. Optimized SQL queries, stored procedures, and transformations to reduce clinical reporting latency by 40%. Orchestrated daily healthcare workflows with Apache Airflow to ensure timely delivery of operational datasets and reduce manual job management. Implemented statistical validation checks in pipelines to maintain clinical dataset fidelity and regulatory compliance. Built Power BI dashboards for clinical trends and operational performance gaps to support hospital leadership decisions. Managed version control and audit trails with Git. Optimized storage for multi-terabyte datasets using Hadoop/Hive to reduce storage footprint by 20%. Automated data quality reporting and discrepancy detection using Python/SQL to improve patient data reliability by 35%, and collaborated on predictive models for patient intake trends improving resource allocation efficiency b
Data Engineer at Fortis Healthcare, Gurgaon, India
October 1, 2017 - November 30, 2020
Developed secure clinical data ingestion pipelines using Azure Data Factory and SQL to ensure 100% data fidelity into centralized analytics platforms. Optimized SQL queries, stored procedures, and transformations to reduce clinical reporting latency. Orchestrated hospital workflows with Apache Airflow for daily production job automation. Implemented statistical validation and automated data quality checks using Python/SQL to resolve patient record discrepancies and improve reliability. Created Power BI dashboards to highlight clinical trends and operational gaps. Reduced clinical dataset storage footprint using Hadoop/Hive optimization strategies and ensured traceable, auditable development using Git.
Data Analyst at Tata AIA Life Insurance
August 1, 2014 - September 30, 2017
Authored complex SQL extraction/transformations for large-scale insurance datasets to ensure 100% fidelity in performance reporting and enable deeper trend analysis. Used R (dplyr/ggplot2) for statistical trend analysis and visualization of policy performance metrics. Produced leadership-ready insights with Excel modeling (formulas/pivot tables) and improved Tableau performance by optimizing extracts and query structures. Documented Python scripts/modules to reduce manual weekly reporting effort by 30% and improved reproducibility. Validated sensitive datasets in SAS with rigorous quality control to ensure correctness of outgoing financial reports. Compiled methodology/recommendation reports in Microsoft Word for executive stakeholders. Built automated R models for policy renewal forecasting to identify at-risk customers, improving retention by 12%, and monitored sales KPIs using dynamic Tableau dashboards to increase regional sales productivity by 20%.
Data Analyst at Tata AIA Life Insurance, Mumbai, India
August 1, 2014 - September 30, 2017
Authored complex SQL queries/functions to extract and transform large-scale insurance datasets for predictive reporting with 100% data fidelity. Performed statistical trend analysis and visualization using R (dplyr/ggplot2). Built leadership-ready insights using Excel modeling (formulas/pivots) and optimized Tableau dashboards by improving query/extract efficiency. Validated sensitive insurance datasets and ensured data integrity through QA processes. Forecasted policy renewals using automated R models to identify at-risk customers and improve retention. Monitored KPIs and agent performance with dynamic Tableau dashboards, improving productivity and accountability.

Education

Bachelor's in Computer Science and Engineering at JNTUH
June 1, 2010 - May 1, 2014
Bachelor's in Computer Science and Engineering at JNTUH
June 1, 2010 - May 1, 2014
Bachelor's in Computer Science and Engineering at JNTUH
June 1, 2010 - May 1, 2014

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Financial Services, Retail, Professional Services, Software & Internet