Site Reliability Engineer (SRE) - Cloud Kubernetes Platform
Site Reliability Engineer (SRE) - Cloud Kubernetes Platform
Location: Glasgow/Leeds (Hybrid)
Experience: 4-8 Years
Contract: 6 months
Role Overview
We are looking for an experienced Site Reliability Engineer (SRE) to join our Cloud Platform team. The successful candidate will have strong hands-on experience managing Kubernetes environments on public cloud platforms such as AKS, EKS, or GKE, along with solid knowledge of SRE, DevOps, infrastructure automation, and observability practices.
You will be responsible for maintaining the availability, reliability, scalability, and performance of cloud-native infrastructure and CI/CD platforms.
Key Responsibilities
- Manage, monitor, and optimize production Kubernetes clusters across AKS, EKS, or GKE.
- Implement and maintain Infrastructure as Code (IaC) using Terraform.
- Build and maintain CI/CD pipelines using Jenkins or similar tools.
- Work closely with development and operations teams to improve system reliability and deployment automation.
- Troubleshoot production incidents and perform Root Cause Analysis (RCA).
- Implement preventive and corrective measures to improve platform reliability.
- Automate operational activities using Python or other Scripting languages.
- Improve platform observability, monitoring, alerting, and performance.
- Support incident management and on-call rotations.
- Contribute to GitOps-based deployment practices using tools such as Flux.
Required Skills & Experience
- 4-8 years of experience in SRE, DevOps, Cloud Infrastructure, or related roles.
- Strong hands-on experience managing Kubernetes clusters in production.
- Experience with at least one major cloud Kubernetes platform: Azure AKS, AWS EKS, or Google GKE.
- Strong knowledge of Terraform and cloud infrastructure automation.
- Practical experience with Jenkins and CI/CD pipelines.
- Good understanding of SRE principles, including:
- Incident management
- Blameless postmortems
- Capacity planning
- Error budgets
- Reliability engineering
- Strong Scripting/programming skills in Python or similar languages.
- Experience with observability tools such as Prometheus, Grafana, or OpenTelemetry.
- Exposure to GitOps practices, preferably with Flux or similar tools.
- Strong troubleshooting, analytical, and problem-solving skills.
- Good written and verbal communication skills.