76 to 100 of 240 Prometheus Jobs in London

Database Platform Engineer

Location
Greater London, England, United Kingdom
environments Understanding of AWS services relevant to data platforms such as RDS, Aurora, S3, EC2 or EKS Familiarity with modern observability stacks such as Prometheus, Grafana, Elk or OTel Desirable: experience with cloud‐native and distributed SQL databases such as Aurora, YugabyteDB or TiDB Desirable: knowledge of data streaming ...

Senior Cloud Engineer (K8S)

Location
Greater London, England, United Kingdom
focus on end-user availability. Desirable but notrequired: Experience with Openstack cloud platform s. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. ...

AI Engineer

Location
Greater London, England, United Kingdom
Azure) and infrastructure‐as‐code (Terraform etc). Hands‐on with DevOps/Infra tooling (CI/CD, Docker, K8s) and observability (Prometheus, Grafana, Datadog etc) Experience building distributed systems Knowledge and hands‐on experience with multiple datastores (both SQL and NoSQL) Desired experience in building agents and workflows (e.g ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity with infrastructure-as-code (e.g. ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
observability, APM and infrastructure monitoring, and application‐specific logging Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, and cloud native tools Knowledge/experience with cloud management tools such as Ansible, Terraform, Vault, and Vagrant Works independently, with guidance in only ...

AI Engineer

Hiring Organisation
Ten Group
Location
London, United Kingdom
Salary
£ 70 K
Azure) and infrastructure-as-code (Terraform etc). Hands-on with DevOps/Infra tooling (CI/CD, Docker, K8s) and observability (Prometheus, Grafana, Datadog etc) Experience building distributed systemsKnowledge and hands-on experience with multiple datastores (both SQL and NoSQL) Desired experience in building agents and workflows (e.g autonomous ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
modern platform and engineering tooling such as GitHub Actions, Apigee, Airflow, and related cloud‐native technologies. Expertise with observability platforms such as Data Dog, Prometheus, Grafana, ELK, Splunk or equivalent, including monitoring, logging, tracing, reliability engineering, and incident management. Sound technical judgement with the ability to balance innovation, risk, operational ...

Senior SRE - Networks

Location
Greater London, England, United Kingdom
analyze internet traffic patterns across multiple dimensions using flow-based tools. Experience working with alerting, monitoring and visibility tools (such as Graphite/Grafana, Prometheus, or Splunk). Knowledge across cloud hosting solutions (i.e., GCP, AWS and Azure). Knowledge of DevOps practices and CI/CD pipelines (ie. ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
model testing, validation gates, and promotion pipelines (continuous training/continuous delivery) that move models safely from experimentation to production.Infrastructure & GPU observability: NVIDIA DCGM, Prometheus/Grafana, and related telemetry stacks for GPU utilization, thermal, and cluster health monitoring.Model & LLM observability: production model performance monitoring, data/concept drift detection ...

Senior DevOps & Cloud SRE Lead: AWS & Kubernetes

Location
Greater London, England, United Kingdom
implement infrastructure-as-code with Terraform, and ensure 99.99% uptime. Lead automated CI/CD with GitHub Actions, Docker, and Helm; build observability with Prometheus, Grafana and Datadog; and design incident response processes to keep systems resilient. #J-18808-Ljbffr ...

Site Reliability Engineer (Chinese speaking, £90k, Financial Services,Technology, London)

Location
Greater London, England, United Kingdom
VMware, and OpenStack. Understanding of CI/CD systems, GitLab & Git, Jenkins, and other similar tools. Expertise in alerting and monitoring technologies such as Prometheus, Slack, Zabbix, Firebase, Grafana, and the ELK stack. Understanding of the TCP/IP protocol and Nginx administration. English language abilities, both written and spoken. ...

Junior Platform Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 70 K
Ansible and Terraform,Experience with HPC and orchestration technologies, such as Slurm or Kubernetes,Experience with Databases and Observability systems, such as Elasticsearch, Datadog, Prometheus, PostgreSQL. ...

DevOps Engineer

Hiring Organisation
Anson Mccade
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£80,000
secure, highly available infrastructure Good communication skills and the ability to work with technical and senior stakeholders Experience with GitLab, Jenkins, Packer, Vault, Prometheus, Grafana, Elastic Stack, Artifactory or Nexus would also be beneficial. Whats in it for you? Work on complex, technically challenging projects within National Security/Government ...

Senior AWS DevOps Architect

Hiring Organisation
IBU CONSULTING LTD
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£400 per day
/CD (GitHub Actions, Jenkins, Azure DevOps) Networking (VPC, Transit Gateway, VPN, Direct Connect) Security (IAM, Security Hub, GuardDuty, WAF) Monitoring (CloudWatch, Grafana, Prometheus) Preferred AWS Services EC2, VPC, IAM, Route53 ECS, EKS, Lambda S3, EFS, FSx RDS, Aurora, DynamoDB CloudWatch, CloudTrail Systems Manager AWS Organizations Control Tower Certifications ...

Senior Dev Ops Engineer

Location
Greater London, England, United Kingdom
Kubernetes, etc.) Extensive experience in designing and building high availability, resilient infrastructure Experience in setting up and maintaining monitoring, alerting, log collection (rsyslog, Grafana, Prometheus, elk) Automation and scripting experience (terraform, packer, ansible, bash) Extensive Linux administration experience (Ubuntu LTS) Bachelor's Degree in Computer Science or equivalent experience Medium ...

Software Engineer II- Global Banking Platform

Location
Greater London, England, United Kingdom
Familiar with databases (SQL or NoSQL). Experience with client/server software architectures & networking, or microservice architectures. Experience with observability tools like Grafana, Prometheus, Open Telemetry and others. Experience with streaming architectures and tools (e.g. Kafka) #J-18808-Ljbffr ...

Lead, Platform Engineer

Location
Greater London, England, United Kingdom
engineers can ship and run their own services with confidence. Reliability and security Set up monitoring, logging and alerting using Cloud Monitoring, Cloud Logging, Prometheus and Grafana. Lead root cause analysis after incidents and see the fixes through. Build security into the delivery process early, and track and address vulnerabilities ...

Site Reliability Engineer (SRE), London

Location
Greater London, England, United Kingdom
high-level programming language like: Java, Swift, Python, or TypeScript Proclivity towards efficient programming emphasizing improvement via complexity analysis. Experience with Nginx, Envoy, Prometheus, and/or Docker. Preferred Qualifications Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model, Subnetting ...

Software Engineering Manager (Test & Devops)

Location
Greater London, England, United Kingdom
Experience managing complex build environments (CMake, Conan, Ceedling) or cloud-native infrastructure-as-code (Terraform, CDK) Familiarity with observability tooling such as Grafana and Prometheus Benefits: Company equity plan so all employees share in the success of the company Salary-sacrifice pension scheme Private medical, dental and vision insurance (medical ...

Deployed Architect, Professional Services (London)

Location
Greater London, England, United Kingdom
strategies, and sizing Experience designing high-availability and disaster recovery solutions Strong understanding of networking, security (SSO/RBAC, TLS, secrets management), and observability (Prometheus, Grafana, Datadog) Experience with CI/CD pipelines for infrastructure and applications Agent Engineering & Development: 1+ years of experience building production AI/ML applications ...

Professional Services Consultant - AI Security

Hiring Organisation
Cato Networks
Location
London, United Kingdom
Salary
£ 70 K
plusFamiliarity with container security, runtime protection, and service mesh architectures (Istio, App Mesh)Practical experience with observability and monitoring stacks (e.g., CloudWatch, Prometheus, Grafana, Datadog)Knowledge of data sovereignty, residency requirements and compliance frameworks relevant to AI workloads (e.g., FedRAMP, SOC 2, ISO 27001, NIST 800-53, HIPPA, PCI, HITRUST ...

Technical Support Engineer

Location
Greater London, England, United Kingdom
etc.). Hands-on understanding of major DevOps tools: Docker, Kubernetes (GCP GKE), CI/CD tools, Terraform, Ansible, Monitoring & Logging Tools (Zipkin, Sentry, Prometheus, ELK stack). Experience in Cloud Based Services (e.g. AWS, GCP). Experience working on Linux based infrastructure. Experience working with scripting languages such ...

Platform Engineer – Monitoring, Observability & SIEM (MONSO)

Location
Greater London, England, United Kingdom
service enablement. Modern Observability & Telemetry: Strong background in Splunk (SPL, dashboards, alerts, data ingestion, forwarders) and also configuring OpenTelemetry collectors and pipelines; knowledge of Prometheus and Grafana or similar tools. Kubernetes (Power User): Strong, hands-on experience deploying and operating workloads, stateful appli-cations, Helm charts, and manifests ...

Professional Services Consultant - AI Security

Location
Greater London, England, United Kingdom
plus Familiarity with container security, runtime protection, and service mesh architectures (Istio, App Mesh) Practical experience with observability and monitoring stacks (e.g., CloudWatch, Prometheus, Grafana, Datadog) Knowledge of data sovereignty, residency requirements and compliance frameworks relevant to AI workloads (e.g., FedRAMP, SOC 2, ISO 27001, NIST 800‐53, HIPPA ...

Software Engineer II- Global Banking Platform

Location
Greater London, England, United Kingdom
Familiar with databases (SQL or NoSQL). Experience with client/server software architectures & networking, or microservice architectures. Experience with observability tools like Grafana, Prometheus, Open Telemetry and others. Experience with streaming architectures and tools (e.g. Kafka) J.P. Morgan is a global leader in financial services, providing strategic advice ...