Hi, I’m Sai, an AI/ML Engineer who designs and builds scalable, production‑ready AI platforms. I specialize in large language models, multi‑agent systems, and cloud‑native MLOps, delivering real‑time, reliable intelligent workflows in enterprise settings. I enjoy optimizing inference, reducing costs, and creating observable, secure AI pipelines that scale with business needs. I collaborate across engineering, data science, and DevOps teams to turn research into impactful products. I’m passionate about continuous experimentation, robust monitoring, and designing systems that stay resilient under load while maintaining strong governance and security.

Sai Nikhil Samudrala

Hi, I’m Sai, an AI/ML Engineer who designs and builds scalable, production‑ready AI platforms. I specialize in large language models, multi‑agent systems, and cloud‑native MLOps, delivering real‑time, reliable intelligent workflows in enterprise settings. I enjoy optimizing inference, reducing costs, and creating observable, secure AI pipelines that scale with business needs. I collaborate across engineering, data science, and DevOps teams to turn research into impactful products. I’m passionate about continuous experimentation, robust monitoring, and designing systems that stay resilient under load while maintaining strong governance and security.

Available to hire

Hi, I’m Sai, an AI/ML Engineer who designs and builds scalable, production‑ready AI platforms. I specialize in large language models, multi‑agent systems, and cloud‑native MLOps, delivering real‑time, reliable intelligent workflows in enterprise settings. I enjoy optimizing inference, reducing costs, and creating observable, secure AI pipelines that scale with business needs.

I collaborate across engineering, data science, and DevOps teams to turn research into impactful products. I’m passionate about continuous experimentation, robust monitoring, and designing systems that stay resilient under load while maintaining strong governance and security.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Language

English
Fluent

Work Experience

AI/ML Engineer at Perplexity
November 1, 2025 - Present
Led Python‑based multi‑agent orchestration pipelines for autonomous browser workflows, enabling contextual task execution for 3,000+ active users and improving response accuracy by 34% in production. Built retrieval‑augmented generation pipelines with LangGraph and vector databases to reduce hallucinations by 29% and cut contextual latency by 41%. Optimized distributed LLM inference with TensorRT‑LLM and vLLM, orchestrated via Kubernetes, achieving a 28% cost reduction while supporting real‑time browser automation and agentic reasoning tasks. Implemented microservices for browser state management, agent communication, and asynchronous workflow orchestration using Kafka, Redis, FastAPI, and event‑driven architectures. Designed scalable memory and embedding infrastructures (Pinecone, Elasticsearch, PostgreSQL) for persistent contextual understanding and long‑context conversations. Deployed Playwright‑based browser automation with CDP for autonomous form handling and dynam
AI/ML Engineer at NVIDIA
April 1, 2024 - November 1, 2025
Engineered Python‑based distributed training pipelines across Kubernetes and Slurm clusters, improving multi‑node throughput by 37% for large‑scale LLM training and inference. Optimized TensorRT‑LLM and Triton inference services using CUDA, NCCL, and PyTorch Distributed, reducing enterprise inference latency by 31%. Developed utilization prediction and anomaly detection models using Prometheus telemetry pipelines to increase cluster resource efficiency by 29%. Built scalable AI orchestration services with Kubernetes, Docker, Helm, and Argo Workflows to automate distributed training, checkpoint management, and inference workflows. Implemented high‑performance distributed learning with PyTorch, DeepSpeed, Megatron‑LM, and Ray across thousands of NVIDIA nodes. Designed observability and monitoring platforms (Grafana, Prometheus, OpenTelemetry, ELK, DCGM) to proactively detect bottlenecks and failures. Architected microservices‑based AI inference platforms (FastAPI, gRPC, Tri
Machine Learning Engineer at Accenture
July 1, 2020 - July 1, 2023
Engineered Python‑based MLOps pipelines (MLflow, PySpark, Databricks) to automate model training, feature engineering, deployment, and lifecycle management across enterprise AI environments serving 5,000+ users. Improved model inference latency by 37% and reduced cloud infrastructure costs by 28% via Kubernetes autoscaling, container optimization, resource tuning, and distributed batch processing. Built enterprise CI/CD workflows (Jenkins, GitHub Actions, Kubeflow, Airflow) for automated retraining, validation, deployment approvals, rollback, and monitoring. Created scalable Python microservices and REST APIs (FastAPI, Docker, Kubernetes, AWS EKS) for secure low‑latency model serving. Designed cloud‑native ML infrastructure on AWS (SageMaker, S3, Lambda, CloudWatch, Terraform, Databricks) to support scalable experimentation, model governance, monitoring, and production deployments. Implemented centralized feature stores, model registries, and drift monitoring (MLflow, Delta Lake,

Education

Master of Science in Computer Science at Florida Atlantic University
January 11, 2030 - June 29, 2026
Bachelor of Science in Computer Science at CMR Engineering College
January 11, 2030 - June 29, 2026

Qualifications

AWS Certified Machine Learning – Specialty
January 11, 2030 - June 29, 2026
AWS Certified Solutions Architect – Professional
January 11, 2030 - June 29, 2026
Generative AI with Large Language Models (LLMs)
January 11, 2030 - June 29, 2026

Industry Experience

Software & Internet, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Hire a AI Engineer

We have the best ai engineer experts on Twine. Hire a ai engineer in San Francisco today.