I’m an AI/ML Engineer in Atlanta with 3+ years of experience building and shipping production Generative AI, LLM, and MLOps systems across finance and healthcare. I specialize in agentic AI workflows, retrieval-augmented generation (RAG), and optimizing LLM inference on AWS using tools like Bedrock and SageMaker, while focusing on measurable outcomes such as improved forecast accuracy, better clinical risk prediction, and faster, safer model deployments. In my recent work, I’ve designed LLM routing to reduce inference cost, engineered enterprise RAG with hybrid retrieval and verifiable citations, and built guarded agent workflows with least-privilege tool access. I also bring strong MLOps discipline—evaluation, monitoring, and automated release pipelines—so models move smoothly from prototype to secure, scalable production for stakeholders like finance and clinical teams.

Sai Krishnam Raju Mudunuri

I’m an AI/ML Engineer in Atlanta with 3+ years of experience building and shipping production Generative AI, LLM, and MLOps systems across finance and healthcare. I specialize in agentic AI workflows, retrieval-augmented generation (RAG), and optimizing LLM inference on AWS using tools like Bedrock and SageMaker, while focusing on measurable outcomes such as improved forecast accuracy, better clinical risk prediction, and faster, safer model deployments. In my recent work, I’ve designed LLM routing to reduce inference cost, engineered enterprise RAG with hybrid retrieval and verifiable citations, and built guarded agent workflows with least-privilege tool access. I also bring strong MLOps discipline—evaluation, monitoring, and automated release pipelines—so models move smoothly from prototype to secure, scalable production for stakeholders like finance and clinical teams.

Available to hire

I’m an AI/ML Engineer in Atlanta with 3+ years of experience building and shipping production Generative AI, LLM, and MLOps systems across finance and healthcare. I specialize in agentic AI workflows, retrieval-augmented generation (RAG), and optimizing LLM inference on AWS using tools like Bedrock and SageMaker, while focusing on measurable outcomes such as improved forecast accuracy, better clinical risk prediction, and faster, safer model deployments.

In my recent work, I’ve designed LLM routing to reduce inference cost, engineered enterprise RAG with hybrid retrieval and verifiable citations, and built guarded agent workflows with least-privilege tool access. I also bring strong MLOps discipline—evaluation, monitoring, and automated release pipelines—so models move smoothly from prototype to secure, scalable production for stakeholders like finance and clinical teams.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

English
Advanced

Work Experience

AI/ML Engineer at Honeywell
October 1, 2024 - Present
Designed a hybrid LLM routing architecture that classifies finance queries by complexity and sensitivity, sending sensitive requests to Claude via AWS Bedrock and routing routine workloads to fine-tuned open-source models on vLLM. Built enterprise RAG pipelines using AWS Bedrock, Titan embeddings, and OpenSearch hybrid retrieval with metadata filtering, reranking, and inline source citations. Implemented LangGraph-based agent workflows with tool-access guardrails and developed MCP/FastMCP servers with authenticated, permission-scoped access to SAP ERP, Oracle Financials, and finance knowledge bases. Deployed and scaled GPU-enabled LLM inference on AWS EKS using continuous batching and parallelism to reduce latency and improve throughput. Fine-tuned domain LLMs using LoRA/QLoRA/PEFT and quantization (AWQ/GPTQ), implemented LLM evaluation/guardrails (RAGAS/DeepEval/LangSmith), and set up monitoring and canary release/rollback using SageMaker Pipelines and GitHub Actions. Also delivered r
ML Engineer at Infosys
January 1, 2022 - July 1, 2023
Built end-to-end ML pipelines processing 15M+ EHR, lab, and patient encounter records to enable predictive analytics across healthcare facilities. Trained XGBoost/LightGBM/Random Forest models to flag high-risk patients for readmission and disease progression, improving prediction performance by 19%. Automated ETL and feature engineering using AWS Glue, EMR, and S3 to integrate EHR, lab, and claims data, improving data availability by 35%. Developed NLP extraction pipelines with Hugging Face Transformers (including BERT/ClinicalBERT) for NER of diagnoses, medications, symptoms, procedures, and clinical observations from unstructured notes, reducing manual chart review effort by 60%. Delivered Power BI dashboards with risk scores and AI-generated insights, and implemented MLOps with SageMaker Pipelines, MLflow, Docker, and automated deployment. Added drift/quality monitoring using SageMaker Model Monitor and CloudWatch, and supported HIPAA-compliant, explainable modeling using SHAP.

Education

Master of Science in Information System Technology at Wilmington University, Delaware
January 11, 2030 - September 2, 2026
Bachelor of Technology in Computer Science and Engineering at Rajalakshmi Engineering College, Chennai, India
January 11, 2030 - September 2, 2026
Master of Science in Information System Technology at Wilmington University, Delaware
January 11, 2030 - September 2, 2026
Bachelor of Technology in Computer Science and Engineering at Rajalakshmi Engineering College, Chennai, India
January 11, 2030 - September 2, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Software & Internet, Professional Services