Hello, I’m Satwik Kasturi, an AI/ML Engineer based in San Jose, CA, with 4+ years of experience building and deploying scalable Machine Learning, Generative AI, and LLM-powered solutions across enterprise environments. I specialize in production-grade RAG architectures, LLM fine-tuning, MLOps pipelines, and distributed AI systems using PyTorch, Python, AWS SageMaker, Kubernetes, and Apache Spark. I focus on creating high-performance AI platforms, optimizing inference infrastructure, and implementing automated model monitoring to boost scalability, reduce latency, and accelerate deployment cycles. I collaborate with cross-functional teams—data scientists, ML researchers, and product engineers—to deliver enterprise AI solutions that drive measurable business impact, reliability, and efficiency. I also have hands-on experience with RLHF, DPO, vector databases, semantic search, NLP, and cloud-native AI infrastructure.

Satwik Kasturi

Hello, I’m Satwik Kasturi, an AI/ML Engineer based in San Jose, CA, with 4+ years of experience building and deploying scalable Machine Learning, Generative AI, and LLM-powered solutions across enterprise environments. I specialize in production-grade RAG architectures, LLM fine-tuning, MLOps pipelines, and distributed AI systems using PyTorch, Python, AWS SageMaker, Kubernetes, and Apache Spark. I focus on creating high-performance AI platforms, optimizing inference infrastructure, and implementing automated model monitoring to boost scalability, reduce latency, and accelerate deployment cycles. I collaborate with cross-functional teams—data scientists, ML researchers, and product engineers—to deliver enterprise AI solutions that drive measurable business impact, reliability, and efficiency. I also have hands-on experience with RLHF, DPO, vector databases, semantic search, NLP, and cloud-native AI infrastructure.

Available to hire

Hello, I’m Satwik Kasturi, an AI/ML Engineer based in San Jose, CA, with 4+ years of experience building and deploying scalable Machine Learning, Generative AI, and LLM-powered solutions across enterprise environments. I specialize in production-grade RAG architectures, LLM fine-tuning, MLOps pipelines, and distributed AI systems using PyTorch, Python, AWS SageMaker, Kubernetes, and Apache Spark. I focus on creating high-performance AI platforms, optimizing inference infrastructure, and implementing automated model monitoring to boost scalability, reduce latency, and accelerate deployment cycles.

I collaborate with cross-functional teams—data scientists, ML researchers, and product engineers—to deliver enterprise AI solutions that drive measurable business impact, reliability, and efficiency. I also have hands-on experience with RLHF, DPO, vector databases, semantic search, NLP, and cloud-native AI infrastructure.

See more

Experience Level

Language

English
Fluent

Work Experience

AI/ML Engineer at Scale AI
May 1, 2025 - Present
Engineered enterprise-scale Generative AI and LLM-powered data validation pipelines using Python, PySpark, and Apache Airflow, processing 1M+ weekly payloads while reducing data curation costs by 35%. Built advanced RAG architectures leveraging FAISS Vector Databases, Semantic Chunking, Cross-Encoder Re-ranking, and Sentence Transformers, improving retrieval precision by 32% and answer faithfulness by 25%. Contributed to LLM fine-tuning and post-training alignment workflows for 10B+ parameter models including LLaMA 3 and Mistral using RLHF, DPO, PyTorch, and DeepSpeed, reducing hallucination rates by 18% across production use cases. Developed automated LLM evaluation pipelines on AWS SageMaker integrated with MLflow, increasing evaluation accuracy by 23% and accelerating model validation cycles by 4 days. Optimized high-performance LLM inference infrastructure using vLLM, TensorRT-LLM, and Kubernetes, reducing Time-to-First Token (TTFT) by 40% and improving scalable AI serving performa
Machine Learning Engineer at KPMG
September 1, 2020 - November 1, 2023
Architected scalable data engineering and real-time ML pipelines using PySpark, Apache Kafka, and AWS Glue, processing 5M+ daily transactions while reducing fraud detection false positives by 28%. Developed production-grade ML and deep learning models for customer churn prediction and credit risk scoring using XGBoost, LightGBM, PyTorch, and advanced feature engineering techniques, achieving 0.89 AUC and driving $200K+ annual business impact. Built scalable MLOps infrastructure using Docker, AWS SageMaker, Kubernetes (EKS), and GitHub Actions CI/CD, reducing inference latency by 40% and enabling automated model deployment pipelines. Developed automated Model Monitoring and Drift Detection frameworks using Evidently AI, Amazon CloudWatch, and observability tooling, maintaining 98% production model accuracy across 10+ live ML deployments. Built advanced Time-Series Forecasting Models using LSTM Networks, Prophet, and statistical forecasting techniques to automate enterprise asset risk re

Education

Master’s in Computer Science at University of Central Missouri, Warrensburg, MO, USA
January 11, 2030 - July 3, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Professional Services