376 to 400 of 608 Prometheus Jobs in the UK

NOC Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£50,000
Terraform and Docker Supporting live production systems and handling incidents Scripting in Python, Bash or Go Using observability and monitoring tools like Grafana, Prometheus, Datadog, Splunk or CloudWatch A solid grasp of core networking concepts, including DNS, TCP/IP and load balancing A genuine drive to automate, improve ...

Senior Site Reliability Engineer (SRE)

Hiring Organisation
fortice
Location
London, UK
Employment Type
Full-time
services using monitoring and alerting solutionsAn understanding of Infrastructure as Code, CI/CD pipelines & associated automation frameworksDemonstrable knowledge of AWS Console/CLI, Prometheus & GrafanaExperience of privileged access management processes & technologiesGood understanding of container orchestration technologies and observability services for Docker & KubernetesTroubleshooting and debug issues across cloud environmentsBenefits ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
assurance, continuous integration and deployment Experience supporting production systems Experience with any of the following: gRPC microservices, Postgres, Pandas, Golang, R, Git, Jenkins, Bazel, Prometheus, Grafana, Airflow, Kubernetes Equal Opportunity EmployerThe the company Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national ...

Core AI Engineer

Location
Greater London, England, United Kingdom
workload isolation technologies Experience in quantitative finance or low‐latency systems AWS experience particularly in hybrid environments Experience with observability tooling such as Prometheus, Grafana or OpenTelemetry Contributions to open‐source projects in relevant domains Why join us? Highly competitive compensation plus annual discretionary bonus Lunch provided (via Just ...

Python Backend Developer

Location
Greater London, England, United Kingdom
frontend work, and Go for select infrastructure Tools: RabbitMQ and Kafka for messaging, PostgreSQL and Redis for data storage Environment: Linux servers Observability: OpenTelemetry, Prometheus, Grafana and Zabbix Must-Haves: Strong background in software development, with strong experience with Python. A degree in Computer Science or a numerical subject from ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
Fargate clusters in AWS, creating common tooling to aid in development tasks, and running shared services such as Opensearch, Envoy, Vault and Prometheus to name a few. We are committed to Open Source software and give back to the community by open sourcing interesting projects. What you’ll be doing ...

Feature Lead - Technology

Location
Greater London, England, United Kingdom
/Linux based environments including shell basics as well as process and network diagnostics. Exposure to monitoring, metrics and tracing tooling - ELK stack, Splunk, Prometheus, Grafana, Graphite, OpenTSDB, OpenTrace, Jaeger Exposure to message-oriented architecture - ZeroMQ, JMS, AMPS, RabbitMQ, Kafka, Google pub/sub Exposure to process/container orchestration ...

Senior Backend Engineer | AI Platform

Location
Greater London, England, United Kingdom
organization. Tech Stack: Backend Python FastAPI Agent Development Kit (ADK) Datastores PostgreSQL BigQuery Firestore Infrastructure Google Cloud Platform (GCP) RabbitMQ Terraform Monitoring & Observability Grafana Prometheus Langfuse incident.io Sentry What to Expect from Our Hiring Process At Plum, we value a lot the time you devote to the hiring process, this ...

Site Reliability Engineer - NS London

Hiring Organisation
BAE SYSTEMS
Location
London, United Kingdom
Salary
£ 70 K
Mongo, Postgreso Know your way around Linux and Windows command lines, e.g. Bash and PowerShello Monitoring large systems using technologies such as Grafana, Prometheus, ELK, Splunko Experience of working in Agile teams, and the tooling that supports it, e.g. Atlassiano Diagnosing and troubleshooting application issues resulting in service outageso Troubleshooting ...

Senior Network Engineer - Low Latency Trading

Hiring Organisation
Hudson River Trading
Location
London, UK
Employment Type
Full-time
Familiarity with firewalls (including Cisco, Palo Alto, Fortinet), VPNs and NATFamiliarity with configuration management tools, such as Ansible or Salt is desirableFamiliarity with Python, Prometheus, Grafana, ELK, GitHub is desirableFamiliarity with public cloud networks, such as AWS, GCP, Azure, is desirableThe estimated base salary range for this position ...

Senior Software Engineer, Video Encoding

Location
United Kingdom
Experience with GPU‐accelerated encoding or hardware media pipelines Familiarity with Kubernetes, ECS, Nomad, or other orchestration platforms Experience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or Datadog Experience building fault‐tolerant ingest or transcoding platforms operating across multiple regions Our Hybrid Work Approach Roku fosters an inclusive ...

Senior Infrastructure Engineer

Location
Greater London, England, United Kingdom
Python or Golang. Deep understanding of container and orchestration internals (e.g., Kubernetes, Docker). Hands-on experience extending and integrating open-source tools like Prometheus, OpenTelemetry, and Grafana. Proven ability to package and scale developer-facing infrastructure utilities with a focus on improving system reliability. Desirable Experience building cloud-agnostic ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer – Associate - London

Location
City Of London, England, United Kingdom
distributed systems, data structures, algorithms and software design fundamentals. Hands‐on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes and deliver durable engineering fixes that improve system behaviour and reduce repeat ...

Senior Software Development Engineer

Location
Reading, England, United Kingdom
Stack .NET (latest versions) Kubernetes & Docker Azure (SQL, CosmosDB, cloud services) PostgreSQL TypeScript/modern web frameworks (e.g. Vue.js) Observability tooling (e.g. Azure Monitor, Prometheus, Grafana) Azure DevOps/CI-CD pipelines Holidays: 25 days per annum + 8 days bank holidays (options to buy/sell days) 37.5 hour ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Location
Westminster, West End, United Kingdom
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
skillsAdvanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps).Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL).Expertise in performance optimisation of distributed systems (e.g., caching, network latency).Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium).Demonstrated success ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
Fargate clusters in AWS, creating common tooling to aid in development tasks, and running shared services such as Opensearch, Envoy, Vault and Prometheus to name a few. The team has also expanded its scope to simplify Data engineering in the organisation using the same techniques we used to ease creating ...

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
sense of ownership Bonus Points Experience with modern Unix/Linux development and runtime environments Experience with monitoring and logging tools like Prometheus and Grafana. Experience with containerization and orchestration technologies, such as Docker and Kubernetes. Compensation For Portugal based hires: Estimated annual salary is between €54,000 - €75,000. ...

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
sense of ownership Bonus Points: Experience with modern Unix/Linux development and runtime environments Experience with monitoring and logging tools like Prometheus and Grafana. Experience with containerization and orchestration technologies, such as Docker and Kubernetes. Compensation For Portugal based hires: Estimated annual salary is between €54,000 – €75,000. ...

Test Environment Manager

Location
Greater London, England, United Kingdom
Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning time, and stability metrics. Monitor environment health using observability tools (Prometheus, Grafana, Splunk, etc.) and proactively identify and resolve performance issues or bottlenecks. Incident & Problem Management Lead incident response for environment-related issues, driving quick resolution … Management teams to ensure test data is consistent, compliant, refreshed automatically, and aligned with environment provisioning needs. Technical Skills & Experience Monitoring & Observability: Expertise with Prometheus, Grafana, Splunk, ELK/EFK, or similar platforms. CI/CD & Automation Tools: Strong experience with Jenkins, GitLab CI, GitHub Actions, and configuration management tools ...

Senior Platform/Dev Ops Engineer

Location
Manchester, England, United Kingdom
Terraform Enterprise. Lead Kubernetes platform architecture, Helm-based deployments, and container orchestration best practices. Establish observability standards and operational excellence using Grafana, Loki, Prometheus, OpenTelemetry, and related tooling. Collaborate with engineering, security, cloud, and architecture teams to improve platform capabilities and developer workflows. Mentor engineers and contribute to platform engineering … Enterprise Kubernetes & Helm Docker Strong experience with multi-cloud infrastructure: AWS, Azure, and/or GCP Strong understanding of observability ecosystems including: Grafana Loki Prometheus OpenTelemetry tooling Strong scripting/programming experience with Python, Kotlin, Bash, or similar languages. Experience with build and dependency tooling such as Gradle, Poetry ...

Platform Engineer

Location
Greater London, England, United Kingdom
Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest … supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on‐call rotation, responding to incidents and helping restore service. Contribute to blameless post‐incident reviews and implement follow ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection. Incident Response & Operations: Engage with urgency … with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development lifecycles. Certifications: GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
incident commanders across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet. Partner with … Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd). Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like. Familiarity with integrating AI/ML-based ...