DevOps / Platform Engineer with 5+ years of experience designing and operating cloud infrastructure and Kubernetes platforms (AKS/EKS), building secure, high-availability systems through Infrastructure as Code, CI/CD, GitOps, and observability. I focus on improving platform reliability and site reliability, accelerating delivery with automated pipelines and canary deployments, reducing MTTR with distributed tracing and centralized logging, and optimizing cloud costs using FinOps practices.

Kalyani Musuluri

DevOps / Platform Engineer with 5+ years of experience designing and operating cloud infrastructure and Kubernetes platforms (AKS/EKS), building secure, high-availability systems through Infrastructure as Code, CI/CD, GitOps, and observability. I focus on improving platform reliability and site reliability, accelerating delivery with automated pipelines and canary deployments, reducing MTTR with distributed tracing and centralized logging, and optimizing cloud costs using FinOps practices.

Available to hire

DevOps / Platform Engineer with 5+ years of experience designing and operating cloud infrastructure and Kubernetes platforms (AKS/EKS), building secure, high-availability systems through Infrastructure as Code, CI/CD, GitOps, and observability.

I focus on improving platform reliability and site reliability, accelerating delivery with automated pipelines and canary deployments, reducing MTTR with distributed tracing and centralized logging, and optimizing cloud costs using FinOps practices.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert

Work Experience

DevOps /Platform Engineer at Wise
September 1, 2025 - Present
Provisioned compliant Azure landing zones using Terraform (VNet peering, Azure Database for PostgreSQL) to reduce environment setup from 4 hours to 45 minutes for PSD2-regulated payment infrastructure. Maintained and scaled AKS clusters supporting 150+ transaction-processing microservices at 15K+ TPS with 99.95% uptime using automated node health checks and proactive capacity scaling. Built GitOps delivery pipelines with GitHub Actions and ArgoCD, adding canary deployments and automated rollbacks to cut release lead time by 75% and achieve 97% deployment success. Implemented OpenTelemetry and Azure Monitor/Application Insights for distributed tracing, reducing MTTR by 55% via correlated alerting and clearer service dependency mapping. Applied FinOps using Azure Cost Management and custom tagging policies to improve chargeback visibility across 10+ teams and identify 120K+ annual savings opportunities, including rightsizing and reserved instances for a 25% cost reduction with zero SLA i
DevOps / Platform Engineer at Wise
September 1, 2025 - Present
Provisioned compliant Azure landing zones using Terraform (VNet peering, Azure Database for PostgreSQL) reducing environment setup time from 4 hours to 45 minutes for PSD2-regulated payment infrastructure. Maintained and scaled AKS clusters supporting 150+ transaction-processing microservices at 15K+ TPS with 99.95% uptime via automated node health checks and proactive capacity scaling. Built GitOps delivery pipelines using GitHub Actions and ArgoCD with canary deployments and automated rollback, cutting release lead time by 75% and improving deployment success to 97%. Integrated OpenTelemetry with Azure Monitor/Application Insights to reduce MTTR by 55% through correlated alerting and service dependency mapping. Implemented FinOps using Azure Cost Management and tagging policies to enable chargeback visibility across 10+ teams and identify 120K+ annual savings opportunities. Optimized spend using rightsizing and Reserved Instances achieving 25% cost reduction with no SLA impact.
DevOps Engineer at Flipkart
May 1, 2020 - September 1, 2023
Maintained 400+ node EKS clusters supporting 200+ microservices using node affinity and pod disruption budgets, sustaining 99.95% availability during flash sales with up to 1.2M concurrent users. Modularized AWS infrastructure across 30+ environments using Terraform and Sentinel policy-as-code checks, cutting rebuild time by 70%. Implemented event-driven autoscaling with KEDA on Kafka (MSK) topics scaling pods from 40 to 320 within 4 minutes to handle 1.2M events/sec during peak sales with zero latency drop. Standardized GitLab CI templates with Trivy vulnerability scanning, reducing onboarding time from 2 weeks to under 2 days and reducing CVE exposure by 92%. Built a multi-cluster observability stack with Prometheus and Thanos and shifted batch workloads to EC2 Spot Instances, reducing P1 MTTR from 42 to 11 minutes and annual cloud spend by 28% (150K).
DevOps Engineer at Cisco
January 1, 2019 - May 1, 2020
Automated network-device configuration and server provisioning using Ansible playbooks with custom Python modules, reducing manual errors by 85% and shortening onboarding from 5 days to 4 hours across 200+ nodes. Maintained Jenkins shared-library CI/CD pipelines for 20+ internal enterprise tools, cutting build-to-deploy time from 2.5 hours to 18 minutes via parallel execution. Deployed centralized logging with the ELK stack (Filebeat, Metricbeat) enabling proactive alerting that reduced unplanned downtime by 40%. Wrote Bash and Python scripts for routine server health checks and patch deployments across 50+ Linux hosts, reducing manual ticket volume by 30%.

Education

Master of Science in International Business with Advanced Research at University of Hertfordshire, Havilland
October 1, 2023 - August 1, 2025
Master of Science in International Business with Advanced Research at University of Hertfordshire
October 1, 2023 - August 1, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Software & Internet, Professional Services, Computers & Electronics

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert