I am a Senior Reliability & Operations Leader with 10+ years of experience designing, governing, and operating mission-critical systems in regulated environments. I stabilize production platforms, manage release risk, and improve service reliability without compromising security or compliance. I’m a trusted advisor to engineering teams and stakeholders, known for calm decision-making under operational pressure and for building systems that behave predictably in production. As a leader, I drive cross-functional collaboration, define observability strategies, and design disaster recovery and resilience plans that keep critical services available. I thrive in regulated contexts, mentoring teams to ship confidently while maintaining rigorous change governance and risk management.

Vugar Aghayev

I am a Senior Reliability & Operations Leader with 10+ years of experience designing, governing, and operating mission-critical systems in regulated environments. I stabilize production platforms, manage release risk, and improve service reliability without compromising security or compliance. I’m a trusted advisor to engineering teams and stakeholders, known for calm decision-making under operational pressure and for building systems that behave predictably in production. As a leader, I drive cross-functional collaboration, define observability strategies, and design disaster recovery and resilience plans that keep critical services available. I thrive in regulated contexts, mentoring teams to ship confidently while maintaining rigorous change governance and risk management.

Available to hire

I am a Senior Reliability & Operations Leader with 10+ years of experience designing, governing, and operating mission-critical systems in regulated environments. I stabilize production platforms, manage release risk, and improve service reliability without compromising security or compliance. I’m a trusted advisor to engineering teams and stakeholders, known for calm decision-making under operational pressure and for building systems that behave predictably in production.

As a leader, I drive cross-functional collaboration, define observability strategies, and design disaster recovery and resilience plans that keep critical services available. I thrive in regulated contexts, mentoring teams to ship confidently while maintaining rigorous change governance and risk management.

See more

Experience Level

Expert
Expert
Expert
Expert
Intermediate

Language

Azerbaijani
Fluent
Turkish
Advanced
Russian
Advanced
English
Fluent

Work Experience

Senior Site Reliability Engineer at Commonwealth Bank of Australia
June 1, 2024 - Present
Own and govern the Fast Track Release (FTR) pipeline, enabling low-risk production deployments while maintaining strict change, quality, and risk controls. Reduced Lead Time to Change (LTTC) by approximately 50%. Defined and operationalised Critical User Journeys (CUJs) for Card Pin Management, identifying key failure modes across APIs and dependent systems. Introduced targeted reliability metrics to ensure critical user work flows remained stable during releases and under load. Acted as a production guardrail and release authority, reviewing API releases, change records, rollback readiness, and cyber approvals. Ensured teams met operational readiness standards before production deployment. Established reliability metrics and operational baselines, including MTTR, deployment frequency, change failure rate, and service availability. Supported squads in building observability capabilities, defining SLIs/SLOs, and creating dashboards and alerts.
Senior DevOps Engineer at Calven
August 1, 2022 - February 1, 2024
Set technical and operational direction for cloud infrastructure, focusing on availability, performance, cost efficiency, and operational consistency. Introduced Infrastructure as Code as a control mechanism to standardise environments, reduce configuration drift, and improve auditability across the infrastructure lifecycle. Integrated infrastructure and application security into daily operations, improving visibility, detection, and response to security events. Led security assessments, audits, and remediation activities to strengthen the organisation’s overall risk posture. Designed and governed disaster recovery and business continuity strategies, aligning RTO and RPO targets with business impact and risk tolerance. Oversaw implementation of backup, failover, and high-availability mechanisms to minimise downtime and data loss. Actively contributed to a security-first and operations-aware engineering culture through knowledge sharing and practical guidance.
DevOps Engineer at Socotra
December 1, 2021 - March 1, 2022
Designed and led delivery of scalable, secure, and resilient infrastructure for a marketplace platform, aligning technical solutions with business growth and performance requirements. Facilitated agile ceremonies as an accountability and alignment mechanism, ensuring predictable delivery, clear ownership, and continuous improvement. Mentored and onboarded junior DevOps engineers, accelerating their operational readiness and embedding good production practices early.
DevOps Engineer at Cocoon Data
February 1, 2020 - November 1, 2021
Cloud & Infrastructure: AWS, GCP, Azure, Terraform, Terragrunt, CloudFormation. Platforms & Delivery: Kubernetes, Docker, GitHub Actions, ArgoCD, Jenkins. Security & Compliance: SIEM (Elastic/Kibana), SAST/DAST, SOC 2. Networking & Edge: API Gateways, Load Balancers, WAF, Akamai. Programming & Scripting: Python, Go, Bash, JavaScript. Led cloud & infrastructure initiatives to improve availability, scalability, and security.
Cloud Engineer at Deloitte
July 1, 2019 - January 1, 2020
Contributed to cloud-focused projects, supporting infrastructure automation, security integration, and disaster recovery planning to minimise downtime and risk.
DevSecOps Engineer at CarsGuide
June 1, 2018 - November 1, 2018
Implemented security-integrated DevOps practices, performing security assessments and remediation activities to strengthen the organisation’s risk posture.
DevOps Engineer at Dynamic4
February 1, 2018 - May 1, 2018
Built secure and scalable infrastructure and supported rapid delivery cycles to meet business goals.
Software Engineer at SympleKYC
September 1, 2016 - February 1, 2018
Developed software solutions and contributed to project delivery with a focus on reliability and performance.

Education

Bachelor’s Degree in Computer Science & Technology at University of Sydney
January 1, 2014 - January 1, 2016

Qualifications

AWS Certified Machine Learning Engineer - Associate
January 1, 2025 - January 23, 2026
AWS Certified Developer – Associate
January 1, 2020 - January 23, 2026
Oracle Certified Associate, Java Programmer
January 1, 2013 - January 23, 2026

Industry Experience

Computers & Electronics, Software & Internet, Professional Services