151 to 175 of 242 Prometheus Jobs in London

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
sense of ownership Bonus Points: Experience with modern Unix/Linux development and runtime environments Experience with monitoring and logging tools like Prometheus and Grafana. Experience with containerization and orchestration technologies, such as Docker and Kubernetes. Compensation For Portugal based hires: Estimated annual salary is between €54,000 – €75,000. ...

Test Environment Manager

Location
Greater London, England, United Kingdom
Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning time, and stability metrics. Monitor environment health using observability tools (Prometheus, Grafana, Splunk, etc.) and proactively identify and resolve performance issues or bottlenecks. Incident & Problem Management Lead incident response for environment-related issues, driving quick resolution … Management teams to ensure test data is consistent, compliant, refreshed automatically, and aligned with environment provisioning needs. Technical Skills & Experience Monitoring & Observability: Expertise with Prometheus, Grafana, Splunk, ELK/EFK, or similar platforms. CI/CD & Automation Tools: Strong experience with Jenkins, GitLab CI, GitHub Actions, and configuration management tools ...

Platform Engineer

Location
Greater London, England, United Kingdom
Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest … supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on‐call rotation, responding to incidents and helping restore service. Contribute to blameless post‐incident reviews and implement follow ...

EKS Engineer

Location
Greater London, England, United Kingdom
Implement and manage CI/CD pipelines for containerized workloads • Ensure security, compliance, and governance across EKS environments • Monitor cluster performance using tools like Prometheus, Grafana, CloudWatch • Manage networking components (VPC, load balancers, ingress controllers) • Optimize cost, performance, and resource utilization • Troubleshoot cluster, networking, and application issues • Collaborate with development ...

System Engineer: £120k + Bonus/benefits (AI Trading)

Hiring Organisation
Hunter Bond
Location
London, United Kingdom
management tools (Chef, Puppet, or Ansible) Exposure to distributed storage systems and related protocols Experience with observability and monitoring tools (Elasticsearch, Logstash, Kibana, Datadog, Prometheus, Grafana) Strong written and verbal communication skills Demonstrated ability to learn quickly and adapt to evolving technologies Ability to work effectively in a fast-paced ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Site Reliability Engineer - Private Cloud Compute

Location
Greater London, England, United Kingdom
high-level programming language like: Java, Go, Python, or Perl Proclivity towards efficient programming emphasizing improvement via complexity analysis. Experience with Kubernetes, Nginx, Envoy, Prometheus, and/or Docker. Preferred Qualifications Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
Autonomous delivery and constructive collaboration with application engineers. Experience supporting business-critical services through an on-call rota. Added Bonus: Experience with Honeycomb, OpenTelemetry, Prometheus, Splunk, including SLO-led practices. Pragmatic use of AI-assisted engineering tools to improve quality and productivity. At Zopa we value flexible ways of working. ...

Software Engineer

Location
Greater London, England, United Kingdom
with large GPU clusters or distributed training environments. Familiarity with distributed training techniques such as DDP or FSDP. Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry. Experience with data pipeline orchestration tools such as Airflow, Flyte, Ray, Metaflow, or Argo Workflows. Experience with containerisation and infrastructure tooling ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
extensibility, such as Bazel, Buck, Pants, Please, etc. Experience with infrastructure as code (Terraform, OpenTofu, or Pulumi) Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on‐call practice Strong proficiency in at least ...

Lead SRE

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Lead SRE - Chase UK

Location
London, United Kingdom
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Lead SRE

Location
Westminster, West End, United Kingdom
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Platform Engineer

Location
Greater London, England, United Kingdom
hands-on with Kubernetes, and enough Terraform to have opinions about how to structure it. Familiarity with CI/CD pipelines and observability tooling (Prometheus, Grafana, Datadog, or similar). Strong technical foundation: Demonstrated ability to write production-quality code and solve hard technical problems (experience with Python, TypeScript/ ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Equity options – share in our success and growth. 10% employer pension contribution – invest in your future. Free office lunches – great food ...

Lead SRE - Chase UK

Location
Greater London, England, United Kingdom
discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information. ...

Software Engineer (ML Projects)

Location
Greater London, England, United Kingdom
cloud‐native TeamCity for CI/CD (lots of teams are releasing code 15-20 times per day!) Terraform Prometheus and Grafana If you have built and deployed complex Python applications or have hands‐on experience with generative AI and LLMs, we would be especially keen to talk. ...

Platform Engineer (10x Openings)

Location
Greater London, England, United Kingdom
Ceph or similar) at an engineering level. Background building Kubernetes operators using frameworks such as Kopf, controller‐runtime, or similar. Experience with observability tooling: Prometheus, Grafana, OpenTelemetry, or structured logging in distributed systems. Experience building SaaS or PaaS layers on top of an IaaS platform. Exposure to serverless or inference ...

Test Environment Manager (10105)

Location
Greater London, England, United Kingdom
Improvement Monitor environment availability, health, performance and utilisation. Develop appropriate metrics and dashboards for environment reporting. Work with monitoring and logging technologies such as Prometheus, Grafana and Splunk. Identify opportunities to improve environment reliability, automation, scalability and cost efficiency. Drive continuous improvement across Test Environment Management processes. Essential Experience 5+ … Operations teams. Highly Desirable Experience Azure/Azure DevOps Jenkins/GitLab Terraform or other IaC technologies HP NonStop infrastructure Java-based environments Prometheus/Grafana/Splunk UFT/Selenium/Cucumber Linux shell scripting Large-scale financial services or similarly complex regulated environments *Rates depend on experience ...

Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)

Hiring Organisation
National Health Service
Location
London, United Kingdom
Salary
£ 70 K
programming/scripting languages such as Python, PowerShell or BashUnderstanding of Linux/Unix & Windows systems, networking, and distributed systemsExperience with observability tools (e.g., Prometheus, Grafana, Datadog) and alerting systemsUnderstanding of infrastructure automation (e.g., Terraform, Ansible, PowerShell, Helm)Excellent communication and collaboration skillsPossesses problem solving skills and the ability … programming/scripting languages such as Python, PowerShell or BashUnderstanding of Linux/Unix & Windows systems, networking, and distributed systemsExperience with observability tools (e.g., Prometheus, Grafana, Datadog) and alerting systemsUnderstanding of infrastructure automation (e.g., Terraform, Ansible, PowerShell, Helm)Excellent communication and collaboration skillsPossesses problem solving skills and the ability ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
Job Title: Java Kafka EngineerLocation: NorthamptonAbout the Job you are considering:Join the Financial Services Business Unit within Capgemini’s Cloud & Custom Applications (C&CA) practice, where we Consult with Purpose, Continuously Evolve, and Architect ...

Senior AI Platform Engineer — Multi-Cloud Infra & Reliability

Location
Greater London, England, United Kingdom
Platform/DevOps engineer to own our infrastructure as code, CI/CD pipelines, and multi-cloud reliability. You’ll work across Terraform, Prometheus, Grafana, and Python to keep services fast, auditable, and secure for enterprise customers. You’ll join a small, fast-moving team building the V7 Go platform ...

Platform / System Engineer

Location
City Of London, England, United Kingdom
Demonstrable experience of external vendor relationship management Nice to haves Containerization (Docker/Kubernetes) in a production environment Monitoring tools in a production environment (Prometheus/Grafana/ELK stack/Splunk) If you are interested, submit your application now. #J-18808-Ljbffr ...

Senior Backend Engineer (Python | AI | 3D Environments | £130,000)

Hiring Organisation
Paradigm Talent
Location
City of London, Greater London, UK
Familiarity with auth, billing, or subscription systems . Background in 3D graphics, creative tooling, or ML pipelines . Knowledge of observability tools like Grafana, Prometheus, or OpenTelemetry. This is a rare opportunity to join an early-stage team backed by leading deep-tech investors, building the foundation of a platform ...