Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Work Experience
Staff Site Reliability Engineer, AI Platform & Infrastructure at Together.ai
October 1, 2025 - PresentDesigned Kubernetes multi-cluster production architecture for GPU-backed inference and training workloads, optimizing for fault-domain isolation, workload segmentation, and cost-aware autoscaling strategies. Evaluated build-vs-buy decisions across provisioning, automation, and configuration management layers, defining long-term platform sustainability for very large node counts. Co-designed hardware/software acceptance and validation tooling with NVIDIA for critical inference and training reliability, and improved high-performance fabric management using NVIDIA InfiniBand, UFM, and Wekaclusters. Owned multi-year technical roadmap focused on hybrid bare-metal performance and cloud-native developer workflows, and established reliability governance (SLOs/SLIs, error budgets, observability standards) with cross-team alignment. Led high-impact incident response during complex scaling events and system outages, influencing CI/CD and dependency standards to prevent recurrence.
Senior Site Reliability Engineer at Booking.com
October 1, 2022 - September 1, 2025Introduced and implemented a Site Reliability Framework for a large and distributed Kafka farm. Contributed to platform-level architectural decisions impacting distributed Kafka infrastructure across multiple production environments while balancing resiliency, throughput, and operational complexity. Developed and managed platforms for developer experience on AWS and Kubernetes. Led development of a Go-based orchestration system governing lifecycle management of Kafka bare-metal nodes, establishing architectural patterns adopted across platform teams and reducing restart windows at scale. Automated infrastructure provisioning via a custom Terraform module to reduce provisioning time, led incident response for high-severity issues with strong post-mortems and root-cause analysis, and mentored engineers. Led integration of automated CI/CD pipelines for microservices, reducing deployment time and achieving near-zero downtime during releases. Championed comprehensive monitoring/observabilit
Engineering Manager - Site Reliability Engineering at Booking.com
September 1, 2021 - October 1, 2022Inherted and expanded an engineering team of five engineers, mixing developers and SREs, and built a complete platform engineering team for bare-metal provisioning and asset management. Owned architectural direction for multi-datacenter bare-metal provisioning platforms supporting large server counts, defined lifecycle management standards, failure recovery patterns, and long-term scalability roadmaps. Set SLI/SLOs and SRE practices, established production readiness standards and SRE governance models influencing platform reliability across multiple engineering organizations, and defined an end-to-end incident management process using SLI/error budget practice. Spearheaded robustness improvements that improved system uptime, leveraging real-time monitoring and automated alerting. Spearheaded observability engineering initiatives improving incident response time using data from real-time monitoring tools to ensure smooth service continuity across multi-datacenter environments.
Senior Engineering Manager - Site Reliability Engineering at Walmart Labs
May 1, 2019 - August 1, 2021Interfaced and collaborated with multi-disciplinary, multi-division organizations across initiatives. Led project work to build a centralized DevOps lifecycle framework for multi-platform releases. Dockerized pharmacy microservices to adhere to security compliance. Developed and implemented an engineering team-based design to run initiatives. Redefined an existing CI/CD pipeline using Docker to achieve faster and more resilient deployments, and coordinated with NOC/support to onboard and support ongoing initiatives including incident runbooks. Supported hiring process improvements via talent acquisition collaboration and setup of two DevOps teams.
DevOps Manager at Paytm Money
April 1, 2018 - August 1, 2019Built and managed infrastructure, DevOps practices, automation, and security end-to-end on AWS backed by Kubernetes and EKS. Owned existing products and solutions used in the Paytm Money cloud infrastructure at application and system levels. Automated configuration management tools including Ansible and Terraform, installed and managed containerization technologies such as Docker, and set up monitoring with Grafana, Prometheus, Alertmanager, and PagerDuty. Built infrastructure dashboards and reports, automated CI code builds using GitLab CI, and supported OS hardening and security audits (Lyinnis). Worked extensively on core AWS services (EC2, VPC, RDS, Route53, CloudWatch, CloudTrail, CloudFront, S3, Aurora), set up Lambda for monitoring and other purposes, and integrated tools like SOLR, Cassandra, ZooKeeper, Exhibitor, and MongoDB. Set up Elasticsearch/Filebeat/Kafka/Logstash for application logging.
Principal Cloud DevOps Engineer at Oracle
March 1, 2016 - March 1, 2018Migrated an entire infrastructure from a managed hosting environment in a private cloud to Oracle Public Cloud or Oracle Bare Metal Cloud using Docker and Terraform. Performed load and performance analysis and improvements for Oracle IaaS. Designed systems and architecture to meet capacity and throughput demands and performance requirements, and modified software to accommodate changes in networks and systems. Wrote Bash scripts to automate Mesosphere DC/OS installation on Oracle Public Cloud for partners, and built scripts to automate partner workshop environments for Bare Metal Cloud. Engineered distributed system architecture using Docker and Terraform, reducing deployment time and improving network scalability across Oracle Public Cloud environments.
Team Lead – Site Reliability Engineering and Cloud Ops at Radiant Info Systems Ltd.
April 1, 2006 - March 1, 2016Built and configured host and network security scans in production and internal networks. Created custom logging/reporting/graphing tools to analyze bottlenecks, enable problem notifications, and improve tuning across hardware, VMs, databases, and JVM tuning. Deployed, maintained, troubleshot, and tuned multi-tier distributed cloud-based application components. Designed, integrated, and managed AWS cloud solutions and provided operational support for monitoring and availability services (e.g., Nagios and similar tools). Managed VM operations and support for server farms running in virtualized environments, supported SSL certificate management for enterprises across multiple SSL providers, and integrated certificates into products like Nginx, Apache, Tomcat, and AWS ELB. Supported day-to-day rack and networking infrastructure operations, provided operational and systems engineering support for web applications, generated and analyzed performance/availability metrics, developed Bash scri
Education
Masters in Information Technology at SSM College of Engineering-Periyar University Erode
June 1, 2000 - March 1, 2005Masters in Information Technology at SSM College of Engineering-Periyar University Erode
June 1, 2000 - March 1, 2005Qualifications
Industry Experience
Software & Internet, Computers & Electronics, Financial Services, Other, Professional Services
Skills
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Hire a Architect
We have the best architect experts on Twine. Hire a architect in Amsterdam today.