Available to hire
Hi, I’m Sai, an AI/ML Engineer who designs and builds scalable, production‑ready AI platforms. I specialize in large language models, multi‑agent systems, and cloud‑native MLOps, delivering real‑time, reliable intelligent workflows in enterprise settings. I enjoy optimizing inference, reducing costs, and creating observable, secure AI pipelines that scale with business needs.
I collaborate across engineering, data science, and DevOps teams to turn research into impactful products. I’m passionate about continuous experimentation, robust monitoring, and designing systems that stay resilient under load while maintaining strong governance and security.
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Language
English
Fluent
Work Experience
AI/ML Engineer at Perplexity
November 1, 2025 - PresentLed Python‑based multi‑agent orchestration pipelines for autonomous browser workflows, enabling contextual task execution for 3,000+ active users and improving response accuracy by 34% in production. Built retrieval‑augmented generation pipelines with LangGraph and vector databases to reduce hallucinations by 29% and cut contextual latency by 41%. Optimized distributed LLM inference with TensorRT‑LLM and vLLM, orchestrated via Kubernetes, achieving a 28% cost reduction while supporting real‑time browser automation and agentic reasoning tasks. Implemented microservices for browser state management, agent communication, and asynchronous workflow orchestration using Kafka, Redis, FastAPI, and event‑driven architectures. Designed scalable memory and embedding infrastructures (Pinecone, Elasticsearch, PostgreSQL) for persistent contextual understanding and long‑context conversations. Deployed Playwright‑based browser automation with CDP for autonomous form handling and dynam
AI/ML Engineer at NVIDIA
April 1, 2024 - November 1, 2025Engineered Python‑based distributed training pipelines across Kubernetes and Slurm clusters, improving multi‑node throughput by 37% for large‑scale LLM training and inference. Optimized TensorRT‑LLM and Triton inference services using CUDA, NCCL, and PyTorch Distributed, reducing enterprise inference latency by 31%. Developed utilization prediction and anomaly detection models using Prometheus telemetry pipelines to increase cluster resource efficiency by 29%. Built scalable AI orchestration services with Kubernetes, Docker, Helm, and Argo Workflows to automate distributed training, checkpoint management, and inference workflows. Implemented high‑performance distributed learning with PyTorch, DeepSpeed, Megatron‑LM, and Ray across thousands of NVIDIA nodes. Designed observability and monitoring platforms (Grafana, Prometheus, OpenTelemetry, ELK, DCGM) to proactively detect bottlenecks and failures. Architected microservices‑based AI inference platforms (FastAPI, gRPC, Tri
Machine Learning Engineer at Accenture
July 1, 2020 - July 1, 2023Engineered Python‑based MLOps pipelines (MLflow, PySpark, Databricks) to automate model training, feature engineering, deployment, and lifecycle management across enterprise AI environments serving 5,000+ users. Improved model inference latency by 37% and reduced cloud infrastructure costs by 28% via Kubernetes autoscaling, container optimization, resource tuning, and distributed batch processing. Built enterprise CI/CD workflows (Jenkins, GitHub Actions, Kubeflow, Airflow) for automated retraining, validation, deployment approvals, rollback, and monitoring. Created scalable Python microservices and REST APIs (FastAPI, Docker, Kubernetes, AWS EKS) for secure low‑latency model serving. Designed cloud‑native ML infrastructure on AWS (SageMaker, S3, Lambda, CloudWatch, Terraform, Databricks) to support scalable experimentation, model governance, monitoring, and production deployments. Implemented centralized feature stores, model registries, and drift monitoring (MLflow, Delta Lake,
Education
Master of Science in Computer Science at Florida Atlantic University
January 11, 2030 - June 29, 2026Bachelor of Science in Computer Science at CMR Engineering College
January 11, 2030 - June 29, 2026Qualifications
AWS Certified Machine Learning – Specialty
January 11, 2030 - June 29, 2026AWS Certified Solutions Architect – Professional
January 11, 2030 - June 29, 2026Generative AI with Large Language Models (LLMs)
January 11, 2030 - June 29, 2026Industry Experience
Software & Internet, Professional Services
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Hire a AI Engineer
We have the best ai engineer experts on Twine. Hire a ai engineer in San Francisco today.