Senior DevOps Engineer and SRE with 8+ years of experience delivering cloud-native infrastructure, CI/CD automation, and enterprise observability in large-scale financial services environments (Barclays, Walmart Tech). Strong focus on building scalable Kubernetes platforms on Azure AKS and AWS EKS using GitLab CI/CD, GitOps, and Terraform IaC across Azure and HPE GreenLake. Experienced with SRE practices including SLI/SLO management, OpenTelemetry distributed tracing, Prometheus/Grafana monitoring, and chaos engineering (LitmusChaos) to improve resilience. Hands-on support for GenAI/RAG infrastructure using Azure OpenAI and Azure AI Search, along with DevSecOps using Vault, OPA Gatekeeper, and container security practices.

Sairam Madichetty

Senior DevOps Engineer and SRE with 8+ years of experience delivering cloud-native infrastructure, CI/CD automation, and enterprise observability in large-scale financial services environments (Barclays, Walmart Tech). Strong focus on building scalable Kubernetes platforms on Azure AKS and AWS EKS using GitLab CI/CD, GitOps, and Terraform IaC across Azure and HPE GreenLake. Experienced with SRE practices including SLI/SLO management, OpenTelemetry distributed tracing, Prometheus/Grafana monitoring, and chaos engineering (LitmusChaos) to improve resilience. Hands-on support for GenAI/RAG infrastructure using Azure OpenAI and Azure AI Search, along with DevSecOps using Vault, OPA Gatekeeper, and container security practices.

Available to hire

Senior DevOps Engineer and SRE with 8+ years of experience delivering cloud-native infrastructure, CI/CD automation, and enterprise observability in large-scale financial services environments (Barclays, Walmart Tech). Strong focus on building scalable Kubernetes platforms on Azure AKS and AWS EKS using GitLab CI/CD, GitOps, and Terraform IaC across Azure and HPE GreenLake.

Experienced with SRE practices including SLI/SLO management, OpenTelemetry distributed tracing, Prometheus/Grafana monitoring, and chaos engineering (LitmusChaos) to improve resilience. Hands-on support for GenAI/RAG infrastructure using Azure OpenAI and Azure AI Search, along with DevSecOps using Vault, OPA Gatekeeper, and container security practices.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Work Experience

Senior DevOps Engineer at Barclays
October 1, 2023 - Present
Built and managed GitLab CI/CD pipelines, Azure Kubernetes Service (AKS), and Terraform IaC for an enterprise GenAI/RAG platform, provisioning Azure OpenAI, Azure AI Search, and Cosmos DB while maintaining 99.9% availability. Re-architected monolithic ingestion into event-driven Kafka architecture with scalable Kubernetes worker pods, reducing document processing time from 14 hours to 90 minutes. Troubleshot Azure AI Search vector search production issues using OpenTelemetry tracing, implemented circuit breakers and automated alerting, restoring reliability within 3 hours. Managed Helm-based Kubernetes deployments for Camunda 8 (Zeebe StatefulSets) and microservices, using Prometheus external metrics autoscaling to scale workloads (5 to 40 pods). Contributed to SRE/Observability initiatives with OpenTelemetry, Prometheus, Grafana, ArgoCD, and LitmusChaos, improving SLO monitoring and reducing alerting effort by 35%. Developed Python automation and Kubernetes CronJobs for telemetry moni
DevOps Engineer at Barclays
December 1, 2022 - October 1, 2023
Engineered and optimized Jenkins and GitLab CI/CD pipelines for 4 microservices using reusable YAML templates, reducing deployment time from 45 to 18 minutes. Implemented canary deployments and progressive delivery using Argo Rollouts, Kubernetes, and Helm, enabling automated rollback and reducing production failures and ML model regression incidents by 80%. Managed scalable Kubernetes infrastructure across Dev/SIT/UAT/Production with Terraform, Helm charts, HPA, and PodDisruptionBudgets, maintaining 99.95% availability during high-volume periods. Strengthened security with HashiCorp Vault, Kubernetes secrets management, RBAC, and dynamic credential rotation for PostgreSQL. Built monitoring and observability with Prometheus, Grafana, Alertmanager, AppDynamics, and PagerDuty, reducing detection time from 22 minutes to under 6 minutes.
DevOps Engineer at Walmart Tech
March 1, 2017 - July 1, 2022
Managed Kubernetes platform migrations on WCNP, provisioning RBAC, namespaces, and cluster access for 20+ application teams migrating from OneOps VMs to containers, including Kubernetes upgrades (1.16–1.21) with zero downtime. Developed and maintained Helm charts, implemented OPA Gatekeeper security policies across 50+ microservices, and integrated CI/CD pipelines with Kubernetes to reduce deployment configuration issues by 30%. Provisioned AWS infrastructure (EC2, EKS, S3, IAM, VPC, CloudWatch) for hybrid workloads, automating deployments with Terraform and Python/Boto3 and enforcing least-privilege IAM across 10+ AWS accounts. Administered high-availability Apache Kafka clusters (scaling, rebalancing, rolling upgrades, consumer lag monitoring) while maintaining zero data loss and sub-10ms latency. Migrated Apache Spark from YARN to Kubernetes using Docker and Ansible automation, improving scheduling efficiency by 25%. Supported chaos engineering testing with blue-green Helm deploym
Cloud Engineer at Walmart Tech
February 1, 2015 - March 1, 2017
Engineered and maintained large-scale OpenStack private cloud infrastructure (Nova, Neutron, Cinder, Keystone, Ceph, RabbitMQ, MariaDB Galera) across 200+ compute nodes, scaling capacity from 10,000 to 30,000+ cores for eCommerce workloads. Designed and executed enterprise OpenStack Juno-to-Kilo migration across multi-region environments using blue-green deployments, migrating 8,000+ VMs within a 6-week sprint and reducing P1 incidents by 35% post-migration. Configured Azure disaster recovery and overflow workloads, implementing cross-cloud failover between on-prem OpenStack and Azure and validating quarterly DR tests to meet RTO/RPO. Automated provisioning and operations using Python and Bash (Nova/Glance), reducing manual effort by 40%. Provided 24x7 production support during peak events (Black Friday/Cyber Monday), monitoring HAProxy, RabbitMQ, MariaDB Galera, and OpenStack APIs to maintain zero downtime.
System Administrator / Infra Engineer at Capgemini
October 1, 2012 - January 1, 2015
Migrated 80 VMware vSphere 4.1 VMs to vSphere 5.5 and upgraded 22 ESXi hosts using VMware Update Manager to enable HA/DRS. Migrated SCCM 2007 SP2 to SCCM 2012 R2 across 18,000+ endpoints and 400 servers, improving patch compliance from 54% to 97%. Implemented VMware Site Recovery Manager (SRM) 5.5 with vSphere Replication for DR of 120 VMs, meeting RTO/RPO targets during DR testing. Optimized SRM failover plans and application startup sequencing to reduce recovery execution time and eliminate recurring RPO breach incidents via MPLS traffic prioritization. Redesigned SCCM boundary groups and deployed 12 regional distribution points to improve remote-site deployments and reduce WAN utilization. Automated Active Directory provisioning/deprovisioning with PowerShell and created 35+ knowledge base articles/SOPs/documentation for support teams.
Infrastructure Support Engineer at Capgemini
December 1, 2009 - October 1, 2012
Assisted migration of 60+ Windows Server 2003 systems to Windows Server 2008 R2 across dual data centers with zero production outages; improved patch compliance using SCCM 2007. Supported large-scale P2V migrations and VMware initiatives, reducing physical footprint by 55% and improving VM provisioning time from 3 days to 4 hours via template-based deployments. Migrated 5,800 Exchange Server 2003 mailboxes to Exchange Server 2010 with zero data loss, reducing post-migration incidents by 40%. Resolved AD replication and DNS issues affecting 400+ users, restoring services within 2 hours and preventing SLA escalation. Managed 20–30 weekly incidents/support tickets for Windows Server, VMware, and Exchange, reducing average resolution time from 6.2 to 3.8 hours through standardization and runbooks. Improved backup/recovery success from 78% to 96% by resolving Symantec NetBackup configuration issues in a 38-host environment while monitoring using HP OpenView, Nagios, and vSphere Client in

Education

Post Graduate Diploma in Data Science at International Institute of Information Technology, Bangalore
January 11, 2030 - August 20, 2026

Qualifications

AWS Certified Solutions Architect – Associate
January 11, 2030 - August 20, 2026
Microsoft Certified: Azure Administrator Associate
January 11, 2030 - August 20, 2026
Certified Kubernetes Administrator (CKA) - Linux Foundation
January 11, 2030 - August 20, 2026

Industry Experience

Financial Services, Software & Internet, Computers & Electronics