Software Development Engineer with 4 years of experience building backend microservices and enterprise-grade Generative AI solutions across cloud-native AWS environments (EKS, Lambda, SageMaker). Skilled in designing production LLM/RAG pipelines and ML inference APIs with strong MLOps automation using FastAPI, LangChain, Kafka, and Airflow. Proven impact includes improving context retrieval accuracy by 25%, reducing MTTR by 15%, and cutting model deployment cycles from weeks to hours. Experienced end-to-end ML lifecycle management, including data pipeline engineering, model versioning, monitoring, and low-latency, high-throughput inference deployments.

Mohammad Gouse Ali Shaik

Software Development Engineer with 4 years of experience building backend microservices and enterprise-grade Generative AI solutions across cloud-native AWS environments (EKS, Lambda, SageMaker). Skilled in designing production LLM/RAG pipelines and ML inference APIs with strong MLOps automation using FastAPI, LangChain, Kafka, and Airflow. Proven impact includes improving context retrieval accuracy by 25%, reducing MTTR by 15%, and cutting model deployment cycles from weeks to hours. Experienced end-to-end ML lifecycle management, including data pipeline engineering, model versioning, monitoring, and low-latency, high-throughput inference deployments.

Available to hire

Software Development Engineer with 4 years of experience building backend microservices and enterprise-grade Generative AI solutions across cloud-native AWS environments (EKS, Lambda, SageMaker). Skilled in designing production LLM/RAG pipelines and ML inference APIs with strong MLOps automation using FastAPI, LangChain, Kafka, and Airflow.

Proven impact includes improving context retrieval accuracy by 25%, reducing MTTR by 15%, and cutting model deployment cycles from weeks to hours. Experienced end-to-end ML lifecycle management, including data pipeline engineering, model versioning, monitoring, and low-latency, high-throughput inference deployments.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more

Language

Work Experience

Software Development Engineer (Contract) at ServiceNow, CA
February 1, 2025 - Present
Designed and developed backend microservices and API interfaces using FastAPI, Node.js, and gRPC on AWS (EKS, EC2, Lambda), implementing Redis caching and asynchronous request handling to improve API performance. Engineered Generative AI features (Now Assist) using LLMs and RAG architectures for workflow summarization and incident resolution, improving context retrieval accuracy by 25% and reducing MTTR by 15%. Implemented Kafka-based event streaming and AWS Lambda consumers to reduce inter-service synchronization latency. Automated deployment validation and release workflows using GitHub Actions, Jenkins, Docker, and Kubernetes. Built Python model inference APIs with FastAPI, TensorFlow, and PyTorch, adding input validation, structured logging, and latency monitoring. Optimized LLM inference latency and throughput via dynamic request batching, model quantization, and evaluation guardrails on AWS EKS. Integrated MLflow for experiment tracking and model versioning; contributed to RAG pi
Software Engineer - Machine Learning at Orion Technolab, India
January 1, 2021 - July 1, 2023
Developed and deployed a predictive supply chain forecasting platform processing 15M+ daily transaction records with up to 175 TPS throughput. Built production forecasting pipelines using XGBoost and LightGBM with rolling-window feature generation, improving forecast accuracy by 18% over legacy statistical models. Architected an automated MLOps framework using Apache Airflow and AWS SageMaker, reducing model deployment cycles from weeks to hours while using MLflow for tracking and lineage across 100+ training runs. Built an NLP compliance engine by fine-tuning BERT to extract risk parameters from unstructured legal vendor contracts, automating a previously manual 30-day auditing workflow. Designed a low-latency inference layer with Docker and high-throughput FastAPI endpoints achieving sub-120ms response times. Implemented data validation and drift detection to monitor inference telemetry, reducing data anomalies by 35%.

Education

Master of Science, Information Systems at California State University, Long Beach (CSULB)
January 11, 2030 - May 1, 2025

Qualifications

AWS Academy Graduate - Machine Learning Foundations - Amazon Web Services
January 11, 2030 - July 30, 2026
AWS Academy Graduate - Cloud Foundations - Amazon Web Services
January 11, 2030 - July 30, 2026
Google AI Essentials - Coursera / Google
January 11, 2030 - July 30, 2026

Industry Experience

Software & Internet, Computers & Electronics, Professional Services, Financial Services, Government