I am a Data Scientist with over four years of experience building ML and AI systems across semiconductor manufacturing, healthcare fraud detection, and e-commerce. I am proficient in Python, SQL, PySpark, TensorFlow, and PyTorch, and I specialize in machine learning, deep learning, NLP, computer vision, and graph analytics. I have hands-on experience with Generative AI and LLMs, including Retrieval Augmented Generation, prompt engineering, embeddings, vector databases, and LangChain. I have deployed production ML systems on AWS and GCP using Spark, Databricks, and MLOps tools like MLflow and CI/CD, and I'm focused on delivering high-impact, production-grade AI solutions at scale.

Sai Prasanth Gavvalaraju

I am a Data Scientist with over four years of experience building ML and AI systems across semiconductor manufacturing, healthcare fraud detection, and e-commerce. I am proficient in Python, SQL, PySpark, TensorFlow, and PyTorch, and I specialize in machine learning, deep learning, NLP, computer vision, and graph analytics. I have hands-on experience with Generative AI and LLMs, including Retrieval Augmented Generation, prompt engineering, embeddings, vector databases, and LangChain. I have deployed production ML systems on AWS and GCP using Spark, Databricks, and MLOps tools like MLflow and CI/CD, and I'm focused on delivering high-impact, production-grade AI solutions at scale.

Available to hire

I am a Data Scientist with over four years of experience building ML and AI systems across semiconductor manufacturing, healthcare fraud detection, and e-commerce. I am proficient in Python, SQL, PySpark, TensorFlow, and PyTorch, and I specialize in machine learning, deep learning, NLP, computer vision, and graph analytics.

I have hands-on experience with Generative AI and LLMs, including Retrieval Augmented Generation, prompt engineering, embeddings, vector databases, and LangChain. I have deployed production ML systems on AWS and GCP using Spark, Databricks, and MLOps tools like MLflow and CI/CD, and I’m focused on delivering high-impact, production-grade AI solutions at scale.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Language

English
Fluent

Work Experience

Data Scientist at NVIDIA
July 1, 2024 - Present
Developed AI-driven semiconductor defect prediction models using PyTorch and distributed PySpark pipelines on AWS SageMaker, analyzing 20M+ wafer inspection and telemetry records, improving defect detection accuracy by 34% and reducing manufacturing yield loss by 18%. Built CNN and Transformer-based computer vision models for automated wafer defect classification using PyTorch and Hugging Face, increasing inspection throughput by 3.5× and eliminating manual quality review for 60K+ wafers per production cycle. Implemented real-time predictive maintenance pipelines using XGBoost, Spark Streaming, and Kafka to detect thermal anomalies and hardware degradation patterns, preventing 1,200+ hours of annual GPU testing downtime and improving infrastructure reliability. Designed scalable data ingestion and feature engineering pipelines using Databricks, Spark, AWS Glue, and Delta Lake to process multi-terabyte manufacturing telemetry datasets, reducing batch processing latency by 45%. Built LL
Data Scientist at CVS Health
January 1, 2024 - July 1, 2024
Built scalable healthcare fraud detection pipelines using Python, Spark, and SQL to analyze multi-million insurance claims datasets, identifying abnormal billing behavior and exposing $2M+ fraudulent claim risk. Developed supervised fraud classification models using Logistic Regression, Random Forest, and XGBoost with advanced ICD and CPT feature engineering, processing 1M+ claims per batch and significantly reducing manual fraud investigations. Designed graph-based fraud detection models using Neo4j and NetworkX to analyze provider–patient–pharmacy networks, identifying 120+ organized fraud rings across 70K+ healthcare relationships. Implemented deep anomaly detection models using PyTorch autoencoders for high-dimensional claims data, improving early fraud detection and reducing suspicious claim identification time from days to hours. Built scalable ML training and deployment pipelines using Databricks, AWS SageMaker, MLflow, and Airflow, enabling automated model retraining and im
Data Scientist at Accenture
February 1, 2021 - July 1, 2022
Designed large-scale e-commerce recommendation systems using collaborative filtering, matrix factorization, and Spark pipelines on multi-million user-product interaction datasets to predict next-purchase behavior. Developed deep learning-based recommendation models using TensorFlow, XGBoost, and Scikit-learn, improving product ranking relevance across 3M+ catalog items and increasing recommendation conversion rates. Engineered real-time recommendation pipelines using Kafka and Spark Streaming to process clickstream and behavioral event data, enabling low-latency personalized recommendations for 30K+ daily active users. Built NLP-based product understanding pipelines using BERT embeddings and Hugging Face Transformers to analyze product titles, descriptions, and reviews, improving cold-start recommendations for newly launched products. Developed scalable REST APIs for model inference using FastAPI, Docker, and Kubernetes, enabling high-throughput recommendation serving for integrated we

Education

Master of Science – Computer & Information Sciences at Christian Brothers University
August 1, 2022 - May 1, 2024

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Manufacturing, Software & Internet, Professional Services, Other