Job Purpose
The Senior DevOps & Site Reliability Engineer is responsible for designing, implementing, and maintaining highly available, scalable, secure, and automated technology platforms that support mission-critical business applications. The role combines software engineering, platform engineering, cloud infrastructure, automation, observability, and operational excellence to improve system reliability, deployment velocity, platform resilience, and customer experience.
The incumbent will drive DevOps and SRE best practices across delivery teams, ensuring that systems are built, deployed, monitored, and operated efficiently while maintaining stringent availability, security, and performance standards.
Key Responsibilities
Platform Engineering & Automation
Design, build, and maintain cloud-native infrastructure and platform services.
Develop Infrastructure as Code (IaC) solutions using modern automation frameworks.
Build reusable deployment templates, pipelines, and automation tooling.
Standardize platform engineering practices across teams.
Automate provisioning, configuration management, and operational processes.
DevOps & CI/CD
Design and maintain CI/CD pipelines supporting both application and infrastructure deployments.
Implement automated testing, security scanning, code quality controls, and release automation.
Drive continuous improvement of release management processes.
Enable fully automated deployment and rollback capabilities.
Improve deployment frequency while reducing deployment risk.
Site Reliability Engineering (SRE)
Establish reliability engineering practices and operational standards.
Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
Improve system availability, performance, resilience, and scalability.
Lead incident response, problem management, and root cause analysis activities.
Drive proactive reliability improvements and technical debt reduction.
Cloud Operations
Design and manage cloud infrastructure environments.
Optimize performance, availability, security, and cost management.
Implement high-availability and disaster recovery solutions.
Support hybrid-cloud and multi-cloud environments where applicable.
Monitoring & Observability
Implement enterprise monitoring and observability platforms.
Establish logging, metrics, tracing, and alerting standards.
Build operational dashboards and platform insights.
Reduce mean time to detect (MTTD) and mean time to recover (MTTR).
Drive predictive monitoring and proactive issue detection.
Security & Compliance
Implement DevSecOps practices throughout the software delivery lifecycle.
Integrate security controls into CI/CD pipelines.
Support vulnerability management and remediation activities.
Ensure compliance with organizational security and regulatory requirements.
Collaborate with security teams to improve platform security posture.
Leadership & Mentoring
Provide technical leadership across engineering teams.
Mentor DevOps, platform, cloud, and reliability engineers.
Promote engineering excellence and operational best practices.
Contribute to architectural decisions and technology roadmaps.
Lead cross-functional initiatives to improve engineering productivity.
Minimum Qualifications
Preferred:
Minimum Experience
8+ years of software engineering, infrastructure, cloud, DevOps, or platform engineering experience.
5+ years of hands-on DevOps engineering experience.
3+ years in Site Reliability Engineering (SRE) or production operations environments.
Proven experience managing mission-critical production systems.
Experience operating large-scale enterprise platforms.
Technical Skills
Cloud Platforms
Strong experience with:
Exposure to AWS and Google Cloud is advantageous.
DevOps Tooling
Azure DevOps
GitHub Enterprise
Git
Jenkins
SonarQube
Artifactory
Nexus
Infrastructure as Code
Terraform
Bicep
ARM Templates
Ansible
Containerisation & Orchestration
Docker
Kubernetes
Helm
OpenShift (advantageous)
Observability
Dynatrace
Grafana
Prometheus
Elastic Stack
Splunk
Azure Monitor
OpenTelemetry
Programming & Scripting
Python
PowerShell
Bash
C#
Java
Go (advantageous)
Technical Competencies
Behavioural Competencies
Strategic Thinking
Problem Solving
Collaboration
Decision Making
Stakeholder Management
Innovation
Continuous Improvement
Coaching and Mentoring
Customer Centricity
Accountability
Key Performance Indicators (KPIs)
The role will be measured against:
Platform availability targets
SLO/SLA compliance
Deployment success rate
Change failure rate
Mean Time to Detect (MTTD)
Mean Time to Recover (MTTR)
Automation coverage
Security and compliance adherence
Cost optimization targets
Engineering productivity improvements
Preferred Certifications
Microsoft Certified: Azure DevOps Engineer Expert
Microsoft Certified: Azure Solutions Architect Expert
Certified Kubernetes Administrator (CKA)
Certified Kubernetes Application Developer (CKAD)
HashiCorp Terraform Associate
AWS Certified DevOps Engineer
ITIL Foundation
SRE Foundation Certification
Ideal Candidate Profile
An experienced engineering professional who can bridge software development, cloud infrastructure, platform engineering, and operations. The successful candidate will possess deep technical expertise, a strong automation mindset, and a passion for building reliable, secure, scalable platforms while enabling engineering teams to deliver business value rapidly and safely