251 to 275 of 397 Prometheus Jobs in England

SRE Observability Technical Lead - Vice President

Location
Greater London, England, United Kingdom
solutions to improve the service reliability and/or increase productivity and efficiency* Hands-on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms.* Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high-availability environments.* Proven ability to troubleshoot ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
JP Morgan Chase
Location
London, United Kingdom
Salary
£ 100 K
skillsAdvanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps).Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL).Expertise in performance optimisation of distributed systems (e.g., caching, network latency).Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium).Demonstrated success ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Location
London, United Kingdom
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cillium ...

SRE Observability Technical Lead - Vice President

Hiring Organisation
Citigroup
Location
London, United Kingdom
Salary
£ 80 K
scalable solutions to improve the service reliability and/or increase productivity and efficiencyHands-on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms.Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high-availability environments.Proven ability to troubleshoot integration issues ...

SRE Observability Technical Lead - Vice President

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
scalable solutions to improve the service reliability and/or increase productivity and efficiencyHands-on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms. Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high-availability environments. Proven ability to troubleshoot ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
Fargate clusters in AWS, creating common tooling to aid in development tasks, and running shared services such as Opensearch, Envoy, Vault and Prometheus to name a few. The team has also expanded its scope to simplify Data engineering in the organisation using the same techniques we used to ease creating ...

Platform Engineer

Location
Greater London, England, United Kingdom
Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest … supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on‐call rotation, responding to incidents and helping restore service. Contribute to blameless post‐incident reviews and implement follow ...

Senior Devops/Infrastructure Engineer

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 80 K
alongside engineers and telco engineering teams. Key skills/experience required: AWS Kubernetes IaC (Terraform) GitOps (Helm, ArgoCD) Monitoring and alerting for production systems (Prometheus/Grafana or similar) Azure (desirable) MLOps in k8s (Kubeflow etc) Running GPU workloads on K8s (drivers, scheduling, utilisation) Familiarity with AI engineering Experience deploying ...

Lead Cloud Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 80 K
Administration and Configuration Management experience, as well as networking experience and an understanding of DevOps principles and practices. Experience with observability tools such as Prometheus and Grafana and working with Kubernetes in production environments at scale is a plus. If you're open to hearing further details about this opportunity ...

Principal DevOps Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 120 K
lead the evolution of DevOps tools Kubernetes, Jenkins, Gitlab, Terraform, and more.Optimising automation and performance.Champion containerisation and high performance base images.Elevate monitoring systems Zabbix, Prometheus, Thanos ensuring 24/7 operational excellence.Secure infrastructure access management, balancing innovation with ironclad security. They're offering a career defining role in a company ...

Platform Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
Central London, London, United Kingdom
Employment Type
Permanent
Salary
£75,000
narrow slice of the stack and told to stay in it, this isn'tthe one. Core stack: Azure, Kubernetes (AKS), Terraform, GitHub Actions, Prometheus, Grafana, PowerShell, etc. You don't need to be a senior engineer. The team already has senior people. What they need is someone solid who wants ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU Limited
Location
Milton Keynes, Buckinghamshire, United Kingdom
Salary
£ 80 K
experience with both Azure, and on-premise virtual machines.Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.Ability ...

Cloud Operations Engineer (remote – London)

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 80 K
Experienced and Certified in cloud computing with AWSExperience in a public cloud such as AWSKnowledge of monitoring and alerting technologies such as Grafana, Prometheus,Expereince of working with of Docker & Kubernetes and Container technology in productionWindows and Linux Operating System Management TechniquesSolid understanding of the OSI ModelExperience in database technology ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent
Salary
£75,000
with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working. ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working. ...

AWS Solutions Architect

Hiring Organisation
Tata Technologies
Location
Gaydon, England, United Kingdom
Code: Terraform and/or CloudFormation, including modular design, version control and environment promotion practices. Monitoring and operations: CloudWatch, CloudTrail, OpenSearch/ELK, Grafana, Prometheus or similar tooling. Operating systems and platform administration: Linux and Windows, scripting and automation using Shell, Python or equivalent. DevOps and CI/ ...

Software Engineer — Observability Instrumentation

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
such as Terraform, ArgoCD, Helm or JenkinsInterest in AI engineering and SRE practices to improve incident response and RCADesirable but not essential experience includes:Prometheus/PromQL, VictoriaMetrics, OpenSearch, Grafana or similar observability backendsAuto-instrumentation, distributed tracing, structured logging or trace/metric correlationKafka or telemetry pipeline architecturesWhy should ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Core AI Engineer

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
servicesFamiliarity with sandboxing and workload isolation technologiesExperience in quantitative finance or low-latency systemsAWS experience particularly in hybrid environmentsExperience with observability tooling such as Prometheus, Grafana or OpenTelemetryContributions to open-source projects in relevant domainsWhy join us Highly competitive compensation plus annual discretionary bonusLunch provided (via Just Eat for Business ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
speed and reducing deployment risk.* **Adaptable & Problem-Solver**: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.* **Ownership & Quality**: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Site Reliability Engineer - Private Cloud Compute

Location
Greater London, England, United Kingdom
high-level programming language like: Java, Go, Python, or Perl Proclivity towards efficient programming emphasizing improvement via complexity analysis. Experience with Kubernetes, Nginx, Envoy, Prometheus, and/or Docker. Preferred Qualifications Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model ...

Senior Java Software Engineer- Platform Engineering

Hiring Organisation
Wise
Location
London, United Kingdom
Salary
£ 80 K
/CD platform.Practical experience applying SRE principles, such as service-level objectives, error budgets, and automated incident prevention.Familiarity with observability tooling such as Prometheus, Grafana, Elastic, or distributed tracing.Knowledge of cloud networking and security.Experience working in a regulated environment or with standards such as PCI DSS.Experience with capacity planning, performance ...

Software Engineer

Location
Greater London, England, United Kingdom
with large GPU clusters or distributed training environments. Familiarity with distributed training techniques such as DDP or FSDP. Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry. Experience with data pipeline orchestration tools such as Airflow, Flyte, Ray, Metaflow, or Argo Workflows. Experience with containerisation and infrastructure tooling ...