Infrastructure Site Reliability Engineer
- Location
- Gloucester, England, United Kingdom
infrastructure as code to support platform lifecycle, as well as simplifying troubleshooting for Incident resolution and provision of tooling for our support organisation Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. Maintain and enhance ’s observability stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services … Bash, Python, Ansible) Deep understanding of observability principles and tools (Prometheus, Grafana) Hands-on experience operating orchestration platforms (Kubernetes, MAAS, Tinkerbell) Strong grasp of ITSM and service operation best practices Excellent communication and mentorship skills Comfortable interfacing with internal stakeholders and external customers Bonus: Knowledge of HPC workloads ...