276 to 300 of 434 Prometheus Jobs in England

Site Reliability Engineer - NS London

Hiring Organisation
BAE SYSTEMS
Location
London, United Kingdom
Salary
£ 70 K
Mongo, Postgreso Know your way around Linux and Windows command lines, e.g. Bash and PowerShello Monitoring large systems using technologies such as Grafana, Prometheus, ELK, Splunko Experience of working in Agile teams, and the tooling that supports it, e.g. Atlassiano Diagnosing and troubleshooting application issues resulting in service outageso Troubleshooting ...

Senior Infrastructure Engineer

Location
Greater London, England, United Kingdom
Python or Golang. Deep understanding of container and orchestration internals (e.g., Kubernetes, Docker). Hands-on experience extending and integrating open-source tools like Prometheus, OpenTelemetry, and Grafana. Proven ability to package and scale developer-facing infrastructure utilities with a focus on improving system reliability. Desirable Experience building cloud-agnostic ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer – Associate - London

Location
City Of London, England, United Kingdom
distributed systems, data structures, algorithms and software design fundamentals. Hands‐on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes and deliver durable engineering fixes that improve system behaviour and reduce repeat ...

Senior Software Development Engineer

Location
Reading, England, United Kingdom
Stack .NET (latest versions) Kubernetes & Docker Azure (SQL, CosmosDB, cloud services) PostgreSQL TypeScript/modern web frameworks (e.g. Vue.js) Observability tooling (e.g. Azure Monitor, Prometheus, Grafana) Azure DevOps/CI-CD pipelines Holidays: 25 days per annum + 8 days bank holidays (options to buy/sell days) 37.5 hour ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Location
Westminster, West End, United Kingdom
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
Fargate clusters in AWS, creating common tooling to aid in development tasks, and running shared services such as Opensearch, Envoy, Vault and Prometheus to name a few. The team has also expanded its scope to simplify Data engineering in the organisation using the same techniques we used to ease creating ...

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
sense of ownership Bonus Points Experience with modern Unix/Linux development and runtime environments Experience with monitoring and logging tools like Prometheus and Grafana. Experience with containerization and orchestration technologies, such as Docker and Kubernetes. Compensation For Portugal based hires: Estimated annual salary is between €54,000 - €75,000. ...

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
sense of ownership Bonus Points: Experience with modern Unix/Linux development and runtime environments Experience with monitoring and logging tools like Prometheus and Grafana. Experience with containerization and orchestration technologies, such as Docker and Kubernetes. Compensation For Portugal based hires: Estimated annual salary is between €54,000 – €75,000. ...

Test Environment Manager

Location
Greater London, England, United Kingdom
Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning time, and stability metrics. Monitor environment health using observability tools (Prometheus, Grafana, Splunk, etc.) and proactively identify and resolve performance issues or bottlenecks. Incident & Problem Management Lead incident response for environment-related issues, driving quick resolution … Management teams to ensure test data is consistent, compliant, refreshed automatically, and aligned with environment provisioning needs. Technical Skills & Experience Monitoring & Observability: Expertise with Prometheus, Grafana, Splunk, ELK/EFK, or similar platforms. CI/CD & Automation Tools: Strong experience with Jenkins, GitLab CI, GitHub Actions, and configuration management tools ...

Senior Platform/Dev Ops Engineer

Location
Manchester, England, United Kingdom
Terraform Enterprise. Lead Kubernetes platform architecture, Helm-based deployments, and container orchestration best practices. Establish observability standards and operational excellence using Grafana, Loki, Prometheus, OpenTelemetry, and related tooling. Collaborate with engineering, security, cloud, and architecture teams to improve platform capabilities and developer workflows. Mentor engineers and contribute to platform engineering … Enterprise Kubernetes & Helm Docker Strong experience with multi-cloud infrastructure: AWS, Azure, and/or GCP Strong understanding of observability ecosystems including: Grafana Loki Prometheus OpenTelemetry tooling Strong scripting/programming experience with Python, Kotlin, Bash, or similar languages. Experience with build and dependency tooling such as Gradle, Poetry ...

Platform Engineer

Location
Greater London, England, United Kingdom
Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest … supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on‐call rotation, responding to incidents and helping restore service. Contribute to blameless post‐incident reviews and implement follow ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
incident commanders across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet. Partner with … Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd). Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like. Familiarity with integrating AI/ML-based ...

Senior Platform Engineer: AI-Ready Infra & Security

Location
Slough, England, United Kingdom
features, automate operations, and build self-service workflows. The role emphasizes strong Python, Linux, Terraform/Ansible, Docker and Kubernetes proficiency, plus observability with Prometheus, Grafana and OpenTelemetry. #J-18808-Ljbffr ...

EKS Engineer

Location
Greater London, England, United Kingdom
Implement and manage CI/CD pipelines for containerized workloads • Ensure security, compliance, and governance across EKS environments • Monitor cluster performance using tools like Prometheus, Grafana, CloudWatch • Manage networking components (VPC, load balancers, ingress controllers) • Optimize cost, performance, and resource utilization • Troubleshoot cluster, networking, and application issues • Collaborate with development ...

DevOps and Infrastructure Engineer

Hiring Organisation
Sanderson Government and Defence
Location
Gloucestershire, South West, United Kingdom
Employment Type
Permanent
solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver repeatable, assured ...

System Engineer: £120k + Bonus/benefits (AI Trading)

Hiring Organisation
Hunter Bond
Location
London, United Kingdom
management tools (Chef, Puppet, or Ansible) Exposure to distributed storage systems and related protocols Experience with observability and monitoring tools (Elasticsearch, Logstash, Kibana, Datadog, Prometheus, Grafana) Strong written and verbal communication skills Demonstrated ability to learn quickly and adapt to evolving technologies Ability to work effectively in a fast-paced ...

SRE Technical Lead - SC Cleared

Hiring Organisation
F5 consultants
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
engineering experience within large-scale production environments Experience defining or working with SLAs, SLOs and error budgets Strong observability experience with tools such as Prometheus, Grafana, Loki, Tempo or OpenTelemetry Strong Infrastructure as Code and GitOps experience - ideally Helm, Kustomize, ArgoCD and/or Tekton Experience across multi-cloud ...

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working. ...

Software Engineer - Backend Developer

Hiring Organisation
Randstad Digital
Location
Manchester, North West, United Kingdom
Employment Type
Contract
Contract Rate
£70 - £72 per hour
distributed transaction handling, and canonical data models. solid grasp of TDD/automated testing, CI/CD pipelines, and active operational telemetry (using OpenTelemetry, Prometheus, or Grafana). Manchester - 2 days in the office | 6 Months Contract + Extension | £72.00 per hour Inside IR35 If you enjoy tackling complex integration ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
Experience with FastAPI, Pydantic, SQLAlchemy, Alembic, pytest, mypy, or Ruff Familiarity with Postgres, Redis, event-driven systems, queues, or streaming architectures Experience with OpenTelemetry, Prometheus, Grafana, Loki, Tempo, or structured logging Comfort with Docker, Kubernetes, Helm, Terraform, GitHub Actions, or release automation Exposure to LLM infrastructure, model routing, embeddings ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Site Reliability Engineer - Private Cloud Compute

Location
Greater London, England, United Kingdom
high-level programming language like: Java, Go, Python, or Perl Proclivity towards efficient programming emphasizing improvement via complexity analysis. Experience with Kubernetes, Nginx, Envoy, Prometheus, and/or Docker. Preferred Qualifications Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
Autonomous delivery and constructive collaboration with application engineers. Experience supporting business-critical services through an on-call rota. Added Bonus: Experience with Honeycomb, OpenTelemetry, Prometheus, Splunk, including SLO-led practices. Pragmatic use of AI-assisted engineering tools to improve quality and productivity. At Zopa we value flexible ways of working. ...