16 of 16 Prometheus Jobs in the East of England

Staff Engineer - Devops

Location
Peterborough, England, United Kingdom
plus; experience with other cloud providers is good to have. Configuration Management with Ansible or similar tools. Monitoring and Metrics using Prometheus, Grafana, or equivalent. Bash Scripting expertise. CICD Pipeline Development using GitHub/GitLab, Jenkins, and related tools (e.g., YAML, Groovy). Basic Linux Proficiency for underlying infrastructure management. ...

Staff DevOps Engineer

Location
Cambridge, England, United Kingdom
understanding of Software Development Lifecycle. Good understanding of microservice architecture. Good experience in Monitoring k8s clusters, APM, services and infrastructure located in AWS (NewRelic,Prometheus). Ability to create automation tools in Go, Python, Bash, Powershell. Familiarity with Continuous Delivery tools used with microservices. Coding Experience. Additional Information Benefits Private ...

Head of Cloud

Hiring Organisation
Epos Now
Location
Norwich, Norfolk, UK
Employment Type
Full-time
teams. Our StackCloud: AWS (expert), Kubernetes (EKS)IaC & Orchestration: Terraform, Helm, TerragruntLanguages: Go, Node.js, Python (automation/tooling)CI/CD: GitHub Actions, ArgoCDObservability: Prometheus, Grafana, OpenTelemetryWhat We're Looking ForProven Leadership: Experience managing and scaling high-performing engineering teams. Cloud Expertise: Deep hands-on experience architecting and operating cloud ...

Head of Cloud

Location
Norwich, England, United Kingdom
Cloud: AWS (expert), Kubernetes (EKS) IaC & Orchestration: Terraform, Helm, Terragrunt Languages: Go, Node.js, Python (automation/tooling) CI/CD: GitHub Actions, ArgoCD Observability: Prometheus, Grafana, OpenTelemetry What We’re Looking For Proven Leadership: Experience managing and scaling high-performing engineering teams. Cloud Expertise: Deep hands-on experience architecting ...

Developer Experience (DevEx) Engineer — Pipeline Squad

Location
Cambridge, England, United Kingdom
serve, and treat internal tooling as a real product.* Security-conscious by default: least privilege, secrets hygiene, supply-chain awareness.* Observability tooling (Grafana, Prometheus).Nice to have* Experience creating agentic development workflows or writing and maintaining AI skills.* npm workspaces/shared-package build systems.* Experience supporting SaaS platforms with ...

Senior Platform Engineer

Location
Cambridge, England, United Kingdom
production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real-world, not greenfield-perfect ...

Senior Storage Architect

Location
Cambridge, England, United Kingdom
analytical, troubleshooting, communication and documentation skills. We also value: Knowledge of GPU compute environments or AI training infrastructure. Experience with monitoring and observability tools (Prometheus, Grafana, etc.). Contributions to open-source storage, data management, or infrastructure projects. Familiarity with object storage systems (S3, RADOS Gateway, MinIO, etc.). ...

Senior Software Infrastructure Engineer

Location
Cambridge, England, United Kingdom
Code (IaC) tools (e.g. Terraform/OpenTofu) Experience with GitHub Actions Experience with build tools (e.g. CMake) Experience with modern observability tooling (e.g. prometheus) Experience with Grafana Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and health cash plan ...

Senior Private Cloud Engineer

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
debug compute, networking, storage in large production environments."Nice To Have" Skills and Experience: Exposure to modern cloud native principles. Familiarity with observability tools (Prometheus, Grafana)!Exposure to large-scale or multi-tenant environments. Exposure to GitOps driven and CI/CD pipelines (Jenkins, ArgoCD. FluxCD, etc)Contribution to Open ...

Platform Technical Lead (OpenShift)

Location
Welwyn Garden City, England, United Kingdom
design: authoring and versioning APIs consumed by other product teams, managing breaking changes and deprecation, Kubernetes API extension patterns (CRDs, operators). Observability engineering: Prometheus, Alertmanager, Grafana or equivalent; defining SLIs/SLOs and designing alerting that reflects service health rather than component noise. Designing and operating fault‐tolerant, highly ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
Experience with FastAPI, Pydantic, SQLAlchemy, Alembic, pytest, mypy, or Ruff Familiarity with Postgres, Redis, event-driven systems, queues, or streaming architectures Experience with OpenTelemetry, Prometheus, Grafana, Loki, Tempo, or structured logging Comfort with Docker, Kubernetes, Helm, Terraform, GitHub Actions, or release automation Exposure to LLM infrastructure, model routing, embeddings ...

Senior Software Engineer - Live & VOD Video Infrastructure

Hiring Organisation
Roku
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
similar technologiesExperience with GPU-accelerated encoding or hardware media pipelinesFamiliarity with Kubernetes, ECS, Nomad, or other orchestration platformsExperience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or DatadogExperience building fault-tolerant ingest or transcoding platforms operating across multiple regions#LI-JC5What's Roku's approach to hybrid working? Roku fosters ...

Infrastructure Monitoring Engineer

Hiring Organisation
BT Group
Location
Ipswich, Suffolk, UK
Employment Type
Full-time
CVMandatoryFamiliarity with using Linux operating systemsExperience with one or more of the following areas Administering or deploying monitoring and observability tools such as CheckMK, Prometheus or ZabbixAdministering or deploying SIEM tools such as Elastic Security, Splunk or Trellix ESMWorking with dashboarding and visualisation tools such as Kibana or GrafanaPreferredAnsible automation ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
models execute on that hardware (inference vs. training, matrix multiplication, KV-caching, etc.) Proficiency with profiling tools (NVIDIA Nsight, PyTorch Profiler) and monitoring stacks (Prometheus, Grafana) Capability to work in Python for data analysis (Pandas, NumPy) and scripting The following are also highly valued: Post-graduate degrees and research experience ...

Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) |

Hiring Organisation
Pure Resourcing Solutions
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £120,000 per annum
training versus inference, matrix multiplication, KV-caching, that level of detail Comfortable in profiling tools like Nsight or PyTorch Profiler, and monitoring stacks like Prometheus and Grafana Python for data work, Pandas and NumPy, plus general scripting Nice to have rather than essential: a postgraduate degree and research background (publications ...

Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)

Hiring Organisation
Pure Resourcing Solutions Limited
Location
Linton, Dry Drayton, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£55000 - £70000/annum
work with GPU or accelerator code, CUDA or similar Familiarity with profiling tools (Nsight, PyTorch Profiler) and ideally some exposure to monitoring stacks (Prometheus, Grafana) Strong Python for data work, Pandas and NumPy, genuine scripting ability Nice to have: exposure to inference serving frameworks like vLLM, published research, or open ...