7 of 7 Prometheus Jobs in Gloucester

Developer Experience (DevEx) Engineer — Pipeline Squad

Hiring Organisation
PTC
Location
Gloucester, Gloucestershire, United Kingdom
Salary
£ 70 K
engineers you serve, and treat internal tooling as a real product.Security-conscious by default: least privilege, secrets hygiene, supply-chain awareness.Observability tooling (Grafana, Prometheus).Nice to haveExperience creating agentic development workflows or writing and maintaining AI skills.npm workspaces/shared-package build systems.Experience supporting SaaS platforms with strict data-privacy ...

Lead DevOps Engineer — Secure Cloud & On-Prem Platform

Location
Gloucester, England, United Kingdom
driving secure, scalable services. You'll mentor engineers, implement CI/CD, GitOps and IaC using ArgoCD, Terraform and Helm, and ensure observability with Prometheus and Grafana in demanding environments. #J-18808-Ljbffr ...

Lead AI Infra SRE: Scale, Reliability & Mentorship

Location
Gloucester, England, United Kingdom
support AI/HPC workloads. You will configure and operate resilient Linux systems (Ubuntu), refine performance, and contribute to the observability stack with Prometheus and Grafana. #J-18808-Ljbffr ...

AI Infra SRE: Scale Kubernetes, Linux & 24/7 Ops

Location
Gloucester, England, United Kingdom
infrastructure. You will own Kubernetes clusters, tune Linux and I/O, and drive automation across the platform. You will champion ITSM practices, maintain Prometheus/Grafana monitoring, and participate in 24x7 on-call support. Mentoring and cross-training with Platform SRE and HPC teams are key parts ...

DevOps & Infrastructure Engineer

Location
Gloucester, England, United Kingdom
solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver repeatable, assured …/CD and GitOps tooling experience, ideally including ArgoCD and automated deployment pipelines. Good understanding of observability, monitoring and alerting using tools such as Prometheus and Grafana, alongside security, networking, logging, secrets management and operational assurance. Able to learn new technologies quickly and help others adopt them safely and effectively. ...

Platform Site Reliability Engineer

Location
Gloucester, England, United Kingdom
tooling for our support organisation Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. Maintain and enhance Radiant’s observability stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis, document … routing, switching Strong experience with API interrogation Strong experience with infrastructure scripting and automation (Bash, Python, Ansible) Deep understanding of observability principles and tools (Prometheus, Grafana preferred) Strong grasp of ITSM and service operation best practices Excellent communication and mentorship skills Comfortable interfacing with internal stakeholders and external customers Bonus ...

Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
provision of tooling for our support organisation Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. Maintain and enhance ’s observability stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis …/IP, DNS, DHCP, VLANs, routing, switching Strong experience with infrastructure scripting and automation (Bash, Python, Ansible) Deep understanding of observability principles and tools (Prometheus, Grafana) Hands-on experience operating orchestration platforms (Kubernetes, MAAS, Tinkerbell) Strong grasp of ITSM and service operation best practices Excellent communication and mentorship skills Comfortable ...