301 to 325 of 344 Prometheus Jobs in London

Senior Manager Consulting ( Cloud Architect)

Location
Greater London, England, United Kingdom
rollback and controlled artefact management using tools such as Jenkins, GitLab and AWS CodePipeline. Establish effective monitoring, logging and incident-management capabilities using CloudWatch, Prometheus, Grafana and the ELK stack, while advising stakeholders on AWS container best practices. Work model We believe hybrid work is the way forward … Experience implementing secure secrets management, Kubernetes security policies, network policies and controls for multi-tenant platforms. Knowledge of observability and centralized logging using CloudWatch, Prometheus, Grafana and ELK in high-security environments. Strong analytical and problem-solving skills, with the ability to create resilient solutions and manage technical ambiguity with ...

Data Engineer

Location
Greater London, England, United Kingdom
with distributed systems technologies such as Kafka and Redis Strong understanding of performance optimisation and debugging Nice to have: ClickHouse, Snowflake or similar technologies Prometheus, Grafana or Sentry You’ll have the opportunity to work on genuinely large-scale data engineering problems where performance and reliability matter, while building infrastructure ...

Performance Engineering Manager

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
high-memory systemsExcellent communication and collaboration skills across research, infrastructure and engineering domainsFamiliarity with profiling and monitoring tools, including perf, eBPF, VTune, Flamegraphs, Prometheus and GrafanaWhy should you apply Highly competitive compensation plus annual discretionary bonusLunch provided (via Just Eat for Business) and dedicated barista bar35 days’ annual leave9% company ...

Senior Product Manager, Observability

Location
Greater London, England, United Kingdom
product that captures and surfaces logs, metrics, and traces at scale, and you understand the architectural and UX tradeoffs involved. Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry. Experience with deployment tooling in a data centre or infrastructure context, including provisioning workflows, networking automation, or zero-touch ...

Principal Engineer - Edge Delivery & Observability

Hiring Organisation
The Financial Times
Location
London, United Kingdom
Salary
£ 100 K
FT. Examples of the kind of work this team tackles are:Managing and improving our central solution for observability tools like Graphite, Grafana, Splunk, Prometheus and Cloudflare.Providing self service APIs and tools that enable other delivery teams to utilise the monitoring solutions.Providing support to other delivery teams ...

Senior Response Engineer - Cloudflare Managed Defense Center (CMDC)

Location
Greater London, England, United Kingdom
leveraging AI/ML models or LLM APIs to build automated troubleshooting diagnostics and dynamic threat mitigations. System Literacy: Experience with monitoring platforms (e.g., Prometheus/Grafana) and querying large network datasets to operationalize contextual routing and security data. Certifications: Advanced networking and network security credentials such as Cisco CCNP ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade … this role includes on-call responsibilities. Key Responsibilities Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
cloud networking Desirable: relevant Cisco, Juniper, or AWS certifications Desirable: experience in the Media or Broadcast technology sector Desirable: familiarity with Nautobot, Netbox, Batfish, Prometheus, Grafana, ELK stack, or Solarwinds Core Competencies Demonstrates extensive expertise in designing, deploying, and troubleshooting large-scale enterprise-grade networks, with a strong focus … Qualifications Cisco Certifications Juniper Certifications AWS Certifications Industry Keywords Media Technology Broadcast Technology Tools & Technologies GitLab CI Jenkins AWS Azure GCP Nautobot Netbox Batfish Prometheus Grafana #J-18808-Ljbffr ...

DevOps and Automation Engineer (Contract)

Location
Greater London, England, United Kingdom
towards self-service environment provisioning using Infrastructure-as-Code (Terraform) and pipeline-driven automation.* Implement advanced observability and monitoring: Use platforms such as Datadog, Prometheus, Grafana, and OpenTelemetry to provide real-time insights into system health, deployments, and business metrics.* Embed security and compliance by design: Integrate security into every … generate new automation ideas and create user stories for rapid prototyping.* Use technologies like Terraform, Ansible, Azure DevOps, Github, OctopusDeploy, Kubernetes, OpenTelemetry, Datadog, Grafana, Prometheus, low-code automation platforms (e.g., Power Automate, UiPath) to help evolve the team's capabilities.* Coordinate planned outage and environment refreshes in collaboration with project ...

Cloud Infra Engineer (Go/Kubernetes) for AI Platform

Location
Greater London, England, United Kingdom
seeks a Confluent Cloud Infrastructure Software Engineer to design, implement and operate cloud foundation services. You will work on Kubernetes operators, Terraform, Datadog and Prometheus, focusing on high availability and scalable infrastructure within the Confluent Cloud Platform. Experience with large-scale systems and public cloud environments is essential. You will ...

Senior Network Engineer – Low latency trading

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 100 K
Linux, covering boot sequence, service management, file, process, network, and kernel management.Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack.Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule, including weekends.This ...

Senior Network Architect – HFT

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 80 K
Linux, covering boot sequence, service management, file, process, network, and kernel management.Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack.Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule, including weekends.This ...

Senior Network Architect - Low latency trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Linux, covering boot sequence, service management, file, process, network, and kernel management. Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack. Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule ...

Senior Network Engineer - Low latency trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Linux, covering boot sequence, service management, file, process, network, and kernel management. Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack. Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule ...

Senior Network Architect - HFT

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Linux, covering boot sequence, service management, file, process, network, and kernel management. Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack. Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule ...

Software Engineer II - Backend (Ruby)

Location
Greater London, England, United Kingdom
grade systems (e.g. you know your way around logs, traces, metrics, feature flags, and can debug live traffic issues)* Familiarity with observability tooling (e.g. Prometheus, Grafana, Sentry and Lightstep)* A good understanding of Domain Driven Design* experience integrating with third-party APIs and services, particularly those with nuanced state transitions ...

Architect/Staff Systems Software Engineer

Location
Greater London, England, United Kingdom
dataflow or non‐GPU accelerator architectures; pre/post‐silicon bring‐up on custom hardware (ASIC/FPGA); production observability at scale (hardware counters, Prometheus/Grafana‐style export, device and cluster views). Compensation & Equity Competitive Salary: Commensurate with your experience, skills, and location Equity & Ownership: Meaningful stock options. ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
infrastructure — monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues and ensuring service uptime. Conduct post‐incident reviews and help improve system … Strong experience with AWS services and containerised applications Strong experience operating operational data stores (Aurora MySQL, DynamoDB). Expertise in using monitoring tools(e.g. Prometheus, Grafana, CloudWatch) for real‐time platform performance insights. Strong understanding of network security and Cloudflare, VPC and networking fundamentals, with a clear grasp ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop … Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital ...

Senior MLOps Engineer

Hiring Organisation
Searchability NS&D
Location
Brentford, England, United Kingdom
registry and ML lifecycle experience Experience with TensorRT, ONNX, quantisation and model deployment Strong Docker, CI/CD and GitHub Actions experience Experience with Prometheus/Grafana monitoring and production observability Cloud/edge AI deployment, with NVIDIA Jetson/DeepStream highly desirable Knowledge of Kubernetes, Kubeflow, Metaflow and Kafka …/DeepSpeed/PyTorch Lightning/MLflow/TensorRT/ONNX/Docker/Kubernetes/CI/CD/GitHub Actions/Prometheus/Grafana/Metaflow/Kafka/NVIDIA Jetson/DeepStream/Edge AI/Computer Vision/Deep Learning ...

Network Automation & OSS Designer

Location
Greater London, England, United Kingdom
pipelines. Architect AIOps capabilities including closed‐loop automation, anomaly detection, and predictive analytics for network operations. Integrate OSS observability with cloud‐native monitoring stacks (Prometheus, Grafana, OpenTelemetry, Elasticsearch). Lead design of intent‐based networking and policy‐driven automation frameworks. Collaborate with product managers, network engineers, platform teams, and DevOps … MANO, VNF/CNF lifecycle management). Understanding of AIOps platforms and closed‐loop automation design for network operations. Experience with observability tooling: OpenTelemetry, Prometheus, Grafana, Jaeger, Loki, ELK Stack. Knowledge of ML/AI model integration for anomaly detection, root‐cause analysis, and predictive network management. Strong grasp ...

Network Automation & OSS Designer

Hiring Organisation
Bounteous
Location
London, United Kingdom
Salary
£ 70 K
pipelines. Architect AIOps capabilities including closed-loop automation, anomaly detection, and predictive analytics for network operations. Integrate OSS observability with cloud-native monitoring stacks (Prometheus, Grafana, OpenTelemetry, Elasticsearch). Lead design of intent-based networking and policy-driven automation frameworks. Collaborate with product managers, network engineers, platform teams, and DevOps … MANO, VNF/CNF lifecycle management). Understanding of AIOps platforms and closed-loop automation design for network operations. Experience with observability tooling: OpenTelemetry, Prometheus, Grafana, Jaeger, Loki, ELK Stack. Knowledge of ML/AI model integration for anomaly detection, root-cause analysis, and predictive network management. Strong grasp ...

Low-Latency Market Data Engineer

Location
Greater London, England, United Kingdom
gaps, data quality checks) Liaise with Market Data capture vendor on issues, upgrades, and performance enhancements Implement a strategic monitoring platform with streaming infrastructure (Prometheus/Grafana) Develop and expand upon AI based tools that link to key data sources for greater instrumentation, visibility, and MTTR Automate both data gathering … Terraform) Experience with a scripting language (Bash, Python) Experience with cloud services and provisioning tools Experience with monitoring and instrumentation of infrastructure/applications (Prometheus/Grafana) Excellent understanding of Market Data and/or third-party feed handlers Ability to be an independent problem solver with troubleshooting, decision making ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
integrations with third-party custodians. Build out Kubernetes and Docker deployments as we containerize more of the custody stack. Set up monitoring and alerting (Prometheus, Grafana, or equivalent) so we know about problems before customers do. Apply security controls and standards throughout - access control, key management, incident response. Provide … skills, with real production experience. A strong security focus - you've worked on systems where key management and access control are critical. Experience with Prometheus, Grafana, or an equivalent monitoring stack. Experience integrating with third-party custodians like BitGo or Fireblocks - this is a key differentiator for this role. What ...

Senior Software Engineer, Custody Services

Hiring Organisation
Robinhood Financial
Location
London, United Kingdom
Salary
£ 80 K
operations.Own the integrations with third-party custodians.Build out Kubernetes and Docker deployments as we containerize more of the custody stack.Set up monitoring and alerting (Prometheus, Grafana, or equivalent) so we know about problems before customers do.Apply security controls and standards throughout — access control, key management, incident response.Provide on-call support … skills, with real production experience.A strong security focus — you've worked on systems where key management and access control are critical.Experience with Prometheus, Grafana, or an equivalent monitoring stack.Bonus pointsExperience integrating with third-party custodians like BitGo or Fireblocks — this is a key differentiator for this role.What we offerChallenging, high ...