126 to 150 of 209 Prometheus Jobs in the UK excluding London

Consulting Engineer, Infrastructure & AI Architecture

Location
City of Edinburgh, Scotland, United Kingdom
private and hybrid cloud platforms (e.g., VMware, OpenShift, Nutanix, Azure Stack, AWS Outposts). Expertise with observability and FinOps tooling across hybrid estates (e.g., Prometheus, Grafana, ELK/OpenSearch, Datadog, cloud cost management platforms). Experience architecting under data sovereignty, regulated industry, or public sector constraints. Contributions to open-source ...

Platform Technical Lead (OpenShift)

Location
Welwyn Garden City, England, United Kingdom
design: authoring and versioning APIs consumed by other product teams, managing breaking changes and deprecation, Kubernetes API extension patterns (CRDs, operators). Observability engineering: Prometheus, Alertmanager, Grafana or equivalent; defining SLIs/SLOs and designing alerting that reflects service health rather than component noise. Designing and operating fault‐tolerant, highly ...

MongoDB Site Reliability Engineer

Location
Knutsford, England, United Kingdom
Some other highly valued skills may include: Using Percona, ClusterControl, CI/CD tools, and automation platforms like Ansible or Chef. Monitoring systems with Prometheus, Grafana, ELK stack, and running containers with Kubernetes. Building APIs with FastAPI and supporting scalable, high-performance systems. You may be assessed ...

Observability SME/Architect/Consultant

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
Making Cross-functional Collaboration Technical ExpertiseCandidates should demonstrate experience with one or more of the following technologies and platforms:Observability Platforms Dynatrace Splunk Grafana Prometheus Elastic/ELK Stack AppDynamics New Relic Service Management & IT Operations ServiceNow ServiceNow Event Management ServiceNow ITOM CMDB and Dependency Mapping Solutions Cloud Monitoring Azure ...

Data Engineer

Hiring Organisation
Searchability NS&D
Location
Gloucester, England, United Kingdom
Apache NiFI SQL and noSQL databases (e.g. MongoDB) ETL processing languages such as Groovy, Python or Java Desirable skills: Java Docker Kubernetes Grafana/Prometheus Integration/debugging Understanding complex system architectures To be Considered: Please either apply by clicking online or emailing me directly to henry.clay-davies@searchability.com. ...

Senior Software Engineer - Space Reliability

Hiring Organisation
Spire Global
Location
Glasgow, UK
Employment Type
Full-time
Postgres, Redis, Elasticsearch, or S3Hands-on experience with Databricks or modern data Lakehouse toolingExperience building or operating monitoring and alerting systems such as Grafana, Prometheus, or NagiosInfrastructure as Code experience with tools such as Terraform or AnsibleML or AI applications in an operational or monitoring contextExperience with Python data visualization ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

Site Reliability Engineer III TLNT1 NI

Hiring Organisation
CME Technology Support Services Ltd
Location
Belfast, UK
Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection. Incident Response & Operations: Engage with urgency … with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development lifecycles. Certifications: GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection. Incident Response & Operations: Engage with urgency … with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development lifecycles. Certifications: GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
incident commanders across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net‐new procurement. Define the instrumentation standards all new services must meet. Partner with … Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd). Hands‐on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like. Familiarity with integrating AI/ML‐based ...

Senior Platform Engineer: AI-Ready Infra & Security

Location
Slough, England, United Kingdom
features, automate operations, and build self-service workflows. The role emphasizes strong Python, Linux, Terraform/Ansible, Docker and Kubernetes proficiency, plus observability with Prometheus, Grafana and OpenTelemetry. #J-18808-Ljbffr ...

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working. ...

Database Engineer

Location
Crawley, England, United Kingdom
configurations for high traffic, data intensive workloads Own backup, restore, PITR and HA strategy across globally distributed systems Monitor platform health using Ops Manager, Prometheus and Grafana Work alongside software and infrastructure teams on schema design, automation and application performance What they are looking for A degree in Computer Science ...

Software Engineer - Backend Developer

Location
Manchester, England, United Kingdom
distributed transaction handling, and canonical data models. solid grasp of TDD/automated testing, CI/CD pipelines, and active operational telemetry (using OpenTelemetry, Prometheus, or Grafana). Manchester - 2 days in the office | 6 Months Contract + Extension | £72.00 per hour Inside IR35 #J-18808-Ljbffr ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, United Kingdom
Employment Type
Full-Time
Salary
£65,000 - £75,000 per annum
with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working. ...

Software Engineer - Backend Developer

Hiring Organisation
Randstad Digital
Location
Manchester, North West, United Kingdom
Employment Type
Contract
Contract Rate
£70 - £72 per hour
distributed transaction handling, and canonical data models. solid grasp of TDD/automated testing, CI/CD pipelines, and active operational telemetry (using OpenTelemetry, Prometheus, or Grafana). Manchester - 2 days in the office | 6 Months Contract + Extension | £72.00 per hour Inside IR35 If you enjoy tackling complex integration ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
Experience with FastAPI, Pydantic, SQLAlchemy, Alembic, pytest, mypy, or Ruff Familiarity with Postgres, Redis, event-driven systems, queues, or streaming architectures Experience with OpenTelemetry, Prometheus, Grafana, Loki, Tempo, or structured logging Comfort with Docker, Kubernetes, Helm, Terraform, GitHub Actions, or release automation Exposure to LLM infrastructure, model routing, embeddings ...

Senior Infrastructure Engineer

Location
Blantyre, Scotland, United Kingdom
recovery software Server estates – HPE/Dell/Cisco/Lenovo hardware management. Observability — hands-on experience across the monitoring stack: Metrics collection (Zabbix, Prometheus, Nagios or similar) Log collection and aggregation (Graylog, Loki, Logstash or similar) Visualisation and dashboarding (Grafana, Kibana or similar) Working with technologies you have never ...

Platform Engineer

Location
Blantyre, Scotland, United Kingdom
recovery software Server estates – HPE/Dell/Cisco/Lenovo hardware management. Observability — hands-on experience across the monitoring stack: Metrics collection (Zabbix, Prometheus, Nagios or similar) Log collection and aggregation (Graylog, Loki, Logstash or similar) Visualisation and dashboarding (Grafana, Kibana or similar) Working with technologies you have never ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance ...

Lead SRE - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance ...

Software Engineering Specialist

Hiring Organisation
BT Group
Location
Cheltenham, Gloucestershire, UK
Employment Type
Full-time
with the Elastic stack technologies and Kibana plugin developmentHave used DevOps tools and principles like Git, Jenkins, GitOpsUsed system monitoring tools such as CheckMK, Prometheus, Grafana or LokiOur PackageTailored benefits make a real difference. That's why we offer a comprehensive range to support your growth, wellbeing, and everyday life. ...

Senior Software Engineer - Live & VOD Video Infrastructure

Hiring Organisation
Roku
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
similar technologiesExperience with GPU-accelerated encoding or hardware media pipelinesFamiliarity with Kubernetes, ECS, Nomad, or other orchestration platformsExperience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or DatadogExperience building fault-tolerant ingest or transcoding platforms operating across multiple regions#LI-JC5What's Roku's approach to hybrid working? Roku fosters ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
multi-cloud—while ensuring rigorous data integrity and mobility A Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual intervention Engineering via Code: While ...