This is a versatile role that bridges cloud infrastructure, automated delivery pipelines, production observability and first-line support operations across our multi-tenant client ecosystem. It suits an engineer who enjoys both sides of platform work: building the automation and monitoring that keeps systems healthy, and being the person who responds when something goes wrong. You will provision environments and improve pipelines, and you will also triage live alerts, work through runbooks and escalate with clean diagnostic data.
Key Responsibilities:
- Provision and maintain cloud environments; build and optimise CI/CD pipelines using GitHub Actions and AWS DevOps tooling.
- Configure uptime monitoring, automated alert routing and operational dashboards in Better Stack and Datadog.
- Act as first-line support for system alerts — diagnosing issues, executing runbooks and escalating with clear triage data.
- Implement AI-assisted monitoring tooling to summarise logs and isolate anomalies during initial service investigations.
- Troubleshoot container health, auto-scaling behaviour and load-balancer configuration.