I’m a Senior Site Reliability Engineer and platform engineer with 11 years of experience designing, automating, and operating large-scale cloud infrastructure on AWS and Azure. I focus on making operations faster and safer—reducing investigation time, improving reliability, and lowering cloud costs through repeatable infrastructure, strong change governance, and observability that leads to action. Recently, I led an Agentic AI for Ops initiative delivering an SRE Agent POC using LLM agents and MCP integrations to triage incidents end-to-end across PagerDuty, Grafana/Prometheus metrics, Kubernetes health signals, and MongoDB/Elasticsearch logs. I also drive incident response and reliability improvements by owning the full lifecycle (detect, mitigate, communicate, and follow-through), mentoring engineers, and building secure CI/CD and cloud-native platforms that scale across multi-region environments.

Kalyan Omkaram

I’m a Senior Site Reliability Engineer and platform engineer with 11 years of experience designing, automating, and operating large-scale cloud infrastructure on AWS and Azure. I focus on making operations faster and safer—reducing investigation time, improving reliability, and lowering cloud costs through repeatable infrastructure, strong change governance, and observability that leads to action. Recently, I led an Agentic AI for Ops initiative delivering an SRE Agent POC using LLM agents and MCP integrations to triage incidents end-to-end across PagerDuty, Grafana/Prometheus metrics, Kubernetes health signals, and MongoDB/Elasticsearch logs. I also drive incident response and reliability improvements by owning the full lifecycle (detect, mitigate, communicate, and follow-through), mentoring engineers, and building secure CI/CD and cloud-native platforms that scale across multi-region environments.

Available to hire

I’m a Senior Site Reliability Engineer and platform engineer with 11 years of experience designing, automating, and operating large-scale cloud infrastructure on AWS and Azure. I focus on making operations faster and safer—reducing investigation time, improving reliability, and lowering cloud costs through repeatable infrastructure, strong change governance, and observability that leads to action.

Recently, I led an Agentic AI for Ops initiative delivering an SRE Agent POC using LLM agents and MCP integrations to triage incidents end-to-end across PagerDuty, Grafana/Prometheus metrics, Kubernetes health signals, and MongoDB/Elasticsearch logs. I also drive incident response and reliability improvements by owning the full lifecycle (detect, mitigate, communicate, and follow-through), mentoring engineers, and building secure CI/CD and cloud-native platforms that scale across multi-region environments.

See more

Experience Level

Work Experience

Senior Site Reliability Engineer at Swimlane
January 1, 2025 - Present
Led an Agentic AI initiative delivering an SRE Agent POC that automates incident triage end-to-end across PagerDuty, Grafana/Prometheus, Kubernetes health, Loki/ELK logs, and MongoDB/Elasticsearch signals. Defined agent tool schemas and guardrails, created read-only diagnostic workflows, and recommended root cause and remediation steps to accelerate on-call first response. Designed MongoDB performance and cloud cost optimizations using scheduled and on-demand compaction strategies across multi-tenant production environments, and drove creation of a new production environment in an in-use-as environments strategy using Terraform and Kustomize with reusable Kubernetes deployment patterns. Owned P1 incident handling via PagerDuty and automated Elasticsearch reindexing across multiple production clusters; also contributed to on-prem to SaaS migrations with zero critical incidents.
Senior Site Reliability Engineer at CyberArk
January 1, 2022 - January 1, 2025
Architected an incident response framework to reduce MTTR by ~40% using systematic RCA, automated runbooks, and cross-functional coordination. Designed and built greenfield production environments in Zurich and Milan across AWS regions to establish multi-region DR and reduce RTO from hours to minutes. Established a security champion program with threat modeling and automated security scanning in CI/CD, reducing vulnerabilities by ~60%. Led infrastructure modernization to cloud-native architectures, delivering ~35% cost reduction and improved scalability; mentored junior SREs and standardized SRE best practices, documentation, and on-call processes. Contributed to multi-region HA/DR, capacity planning, and reliability improvements alongside performance tuning and FinOps alignment.
Site Reliability Engineer at ADP
January 1, 2019 - January 1, 2021
Led migration from Puppet to Ansible and Trophosphere to Terraform across 200+ servers, reducing configuration drift by ~80%. Architected hybrid Terraform/Ansible provisioning and automated Kong API Gateway using Python, cutting provisioning time by ~60% and configuration errors by ~70%. Established Helm chart standards and self-service automation to accelerate application onboarding by ~50%. Improved deployment governance via standardized Helm modules and repeatable automation practices.
DevOps Engineer at iNfoTech
January 1, 2017 - January 1, 2019
Implemented multi-cloud Terraform for AWS and Azure covering 100+ resources with ~99.95% availability and ~30% cost reduction. Built CI/CD pipelines and Prometheus/Grafana monitoring, reducing release cycles and MTTR by ~50%. Focused on automation, infrastructure-as-code governance, and monitoring-driven operational improvements across environments.
DevOps Engineer at Mindtree
January 1, 2015 - January 1, 2017
Engineered build/release automation and created 20+ Chef cookbooks, improving release frequency by ~3x and reducing configuration errors by ~65%. Established CI/CD best practices with automated testing, quality gates, and code-driven deployments to improve reliability and reduce operational friction.

Education

Bachelor of Technology at JNTU Anantapur
January 1, 2008 - January 1, 2012

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Computers & Electronics

Experience Level

Hire a DevOps Developer

We have the best devops developer experts on Twine. Hire a devops developer in Hyderabad today.