401 to 425 of 608 Prometheus Jobs in the UK

Senior Platform Engineer: AI-Ready Infra & Security

Location
Slough, England, United Kingdom
features, automate operations, and build self-service workflows. The role emphasizes strong Python, Linux, Terraform/Ansible, Docker and Kubernetes proficiency, plus observability with Prometheus, Grafana and OpenTelemetry. #J-18808-Ljbffr ...

EKS Engineer

Location
Greater London, England, United Kingdom
Implement and manage CI/CD pipelines for containerized workloads • Ensure security, compliance, and governance across EKS environments • Monitor cluster performance using tools like Prometheus, Grafana, CloudWatch • Manage networking components (VPC, load balancers, ingress controllers) • Optimize cost, performance, and resource utilization • Troubleshoot cluster, networking, and application issues • Collaborate with development ...

Senior Devops/Infrastructure Engineer

Hiring Organisation
Intellectual Capital Resources
Location
London, UK
Employment Type
Full-time
alongside engineers and telco engineering teams. Key skills/experience required: AWS Kubernetes IaC (Terraform) GitOps (Helm, ArgoCD) Monitoring and alerting for production systems (Prometheus/Grafana or similar) Azure (desirable) MLOps in k8s (Kubeflow etc) Running GPU workloads on K8s (drivers, scheduling, utilisation) Familiarity with AI engineering Experience deploying ...

Lead Cloud Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
Administration and Configuration Management experience, as well as networking experience and an understanding of DevOps principles and practices. Experience with observability tools such as Prometheus and Grafana and working with Kubernetes in production environments at scale is a plus. If you're open to hearing further details about this opportunity ...

DevOps and Infrastructure Engineer

Hiring Organisation
Sanderson Government and Defence
Location
Gloucestershire, South West, United Kingdom
Employment Type
Permanent
solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver repeatable, assured ...

System Engineer: £120k + Bonus/benefits (AI Trading)

Hiring Organisation
Hunter Bond
Location
London, United Kingdom
management tools (Chef, Puppet, or Ansible) Exposure to distributed storage systems and related protocols Experience with observability and monitoring tools (Elasticsearch, Logstash, Kibana, Datadog, Prometheus, Grafana) Strong written and verbal communication skills Demonstrated ability to learn quickly and adapt to evolving technologies Ability to work effectively in a fast-paced ...

Principal DevOps Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
DevOps tools Kubernetes, Jenkins, Gitlab, Terraform, and more. Optimising automation and performance. Champion containerisation and high performance base images. Elevate monitoring systems Zabbix, Prometheus, Thanos ensuring 24/7 operational excellence. Secure infrastructure access management, balancing innovation with ironclad security. They're offering a career defining role in a company ...

SRE Technical Lead - SC Cleared

Hiring Organisation
F5 consultants
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
engineering experience within large-scale production environments Experience defining or working with SLAs, SLOs and error budgets Strong observability experience with tools such as Prometheus, Grafana, Loki, Tempo or OpenTelemetry Strong Infrastructure as Code and GitOps experience - ideally Helm, Kustomize, ArgoCD and/or Tekton Experience across multi-cloud ...

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working. ...

Software Engineer - Backend Developer

Hiring Organisation
Randstad Digital
Location
Manchester, North West, United Kingdom
Employment Type
Contract
Contract Rate
£70 - £72 per hour
distributed transaction handling, and canonical data models. solid grasp of TDD/automated testing, CI/CD pipelines, and active operational telemetry (using OpenTelemetry, Prometheus, or Grafana). Manchester - 2 days in the office | 6 Months Contract + Extension | £72.00 per hour Inside IR35 If you enjoy tackling complex integration ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
Experience with FastAPI, Pydantic, SQLAlchemy, Alembic, pytest, mypy, or Ruff Familiarity with Postgres, Redis, event-driven systems, queues, or streaming architectures Experience with OpenTelemetry, Prometheus, Grafana, Loki, Tempo, or structured logging Comfort with Docker, Kubernetes, Helm, Terraform, GitHub Actions, or release automation Exposure to LLM infrastructure, model routing, embeddings ...

Software Engineer — Observability Instrumentation

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
such as Terraform, ArgoCD, Helm or JenkinsInterest in AI engineering and SRE practices to improve incident response and RCADesirable but not essential experience includes: Prometheus/PromQL, VictoriaMetrics, OpenSearch, Grafana or similar observability backendsAuto-instrumentation, distributed tracing, structured logging or trace/metric correlationKafka or telemetry pipeline architecturesWhy should ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Site Reliability Engineer - Private Cloud Compute

Location
Greater London, England, United Kingdom
high-level programming language like: Java, Go, Python, or Perl Proclivity towards efficient programming emphasizing improvement via complexity analysis. Experience with Kubernetes, Nginx, Envoy, Prometheus, and/or Docker. Preferred Qualifications Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model ...

eTrading Platform Product Owner

Hiring Organisation
ING Banking
Location
London, UK
Employment Type
Full-time
supporting root-cause analysis and measurable improvement actions. Improve observability by strengthening meaningful monitoring, alerting, service-level measures and reporting using tools such as Prometheus and Grafana. Drive automation through Azure DevOps, pipelines, YAML and Ansible, reducing manual activity and improving repeatability, control and delivery confidence. Support robust change, release ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
Autonomous delivery and constructive collaboration with application engineers. Experience supporting business-critical services through an on-call rota. Added Bonus: Experience with Honeycomb, OpenTelemetry, Prometheus, Splunk, including SLO-led practices. Pragmatic use of AI-assisted engineering tools to improve quality and productivity. At Zopa we value flexible ways of working. ...

Software Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
with large GPU clusters or distributed training environments. Familiarity with distributed training techniques such as DDP or FSDP. Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry. Experience with data pipeline orchestration tools such as Airflow, Flyte, Ray, Metaflow, or Argo Workflows. Experience with containerisation and infrastructure tooling ...

Neo4j Platform Consultant

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
scaling)Security (RBAC, authentication, data protection)Preferred SkillsExperience with Graph RAG workloads and traversal optimizationDevOps tools (Docker, Kubernetes, CI/CD pipelines)Monitoring tools (Prometheus, Grafana, etc.)Experience with other graph databases (Neptune, TigerGraph)Job ExpectationsEnsure stable, secure, and high-performing Neo4j platform operationsEnable efficient graph query execution ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
extensibility, such as Bazel, Buck, Pants, Please, etc. Experience with infrastructure as code (Terraform, OpenTofu, or Pulumi) Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on‐call practice Strong proficiency in at least ...

Platform Engineer

Location
Blantyre, Scotland, United Kingdom
recovery software Server estates – HPE/Dell/Cisco/Lenovo hardware management. Observability — hands-on experience across the monitoring stack: Metrics collection (Zabbix, Prometheus, Nagios or similar) Log collection and aggregation (Graylog, Loki, Logstash or similar) Visualisation and dashboarding (Grafana, Kibana or similar) Working with technologies you have never ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance ...

Lead SRE

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Lead SRE - Chase UK

Location
London, United Kingdom
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Platform Engineer

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
good grounding in Linux, networking and security fundamentalsExperience troubleshooting production systems and working through operational issuesUseful experienceTerragrunt or HelmGitHub Actions, ArgoCD or FluxDatadog, Prometheus or GrafanaPostgreSQL or MySQLSupporting developer-experience or internal-platform improvementsWorking with security tooling or policy-as-codeWe don't expect every box to be ticked. ...