551 to 575 of 4,890 Permanent Observability Jobs

Senior Platform Engineer Engineering · Manchester ·

Location
Manchester, England, United Kingdom
over £50 billion in loans annually, you will embed automated security scanning and compliance-as-code into delivery pipelines, spearhead cloud cost-optimization and observability initiatives, and provide technical guidance and mentorship to mid-level engineers across the team. About you: GCP Cloud Expertise: Proven track record managing, scaling … Paths," creating internal self-service tooling, writing developer documentation, and conducting code reviews to drive a strong "you build it, you run it" culture. Observability & Cost Optimization: Strong background in system resilience, latency reduction, observability implementation, and cloud cost-optimization initiatives (e.g., resource tagging and footprint reduction). What will ...

Senior DevOps Engineer

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
performance, and availability Investigate and resolve production incidents in a timely manner Perform root cause analysis and implement preventative measures Enhance logging, alerting, and observability across the platform Security & Governance Implement Azure security best practices and policies Manage identity and access controls in line with governance standards Ensure compliance with … modern development workflows Desirable Knowledge of scripting languages (PowerShell, Bash, or Python) Experience working with Azure Front Door, WAF, or CDN technologies Exposure to observability tooling (distributed tracing, metrics platforms) Experience with cost management and optimisation in Azure Understanding of DevOps principles and Agile delivery practices Personal Attributes Strong problem ...

DevOps / SRE Engineer (Sheffield)

Location
Sheffield, England, United Kingdom
DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud‐native platform. You’ll work across the business to improve how we build, deploy and monitor our systems , while playing a key role … platforms Solid experience with Terraform and IaC automation Experience participating in or managing production incidents and on‐call Strong grasp of monitoring, alerting, and observability principles Ability to diagnose and fix complex distributed systems issues Demonstrated use of GenAI tools (ChatGPT, GitHub Copilot, Claude) in engineering workflows Excellent communication ...

DevOps / SRE Engineer (London)

Location
Greater London, England, United Kingdom
DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud‐native platform. You’ll work across the business to improve how we build, deploy and monitor our systems , while playing a key role … platforms Solid experience with Terraform and IaC automation Experience participating in or managing production incidents and on‐call Strong grasp of monitoring, alerting, and observability principles Ability to diagnose and fix complex distributed systems issues Demonstrated use of GenAI tools (ChatGPT, GitHub Copilot, Claude) in engineering workflows Excellent communication ...

Lead Agentic AI Architect

Location
United Kingdom
memory, orchestration and tool integration Production-grade LLM and RAG architectures Intelligent search, knowledge assistants, workflow automation and decision-support systems Enterprise LLMOps, MLOps, observability and AI governance AI security, guardrails, evaluation and Responsible AI Cloud AI platforms and hybrid/multi-cloud architecture You'll also provide technical leadership … Hybrid/Multi-Cloud Engineering & Platform Python Java Scala TypeScript SQL Kubernetes Docker Terraform Bicep CI/CD Infrastructure as Code AI Governance & Observability LLMOps MLOps Model Governance AI Security & Guardrails LangSmith Arize Monitoring & Evaluation Frameworks Qualifications Relevant qualifications may include: Master's degree in Computer Science, AI, Data Science ...

Lead Engineer, Site Reliability Engineering

Location
Greater London, England, United Kingdom
Write automation to scale systems sustainably, prevent service issues, or when they occur, quickly recover service. Partner with development teams to improve system reliability, observability, and release velocity. Participate in on-call rotations, incident response, postmortems, and root cause analysis and resolution. Be a vocal advocate of strong/sound … Infrastructure as Code using Terraform.* Hands on Experience with one of the following cloud platforms: Azure, AWS, or GCP.* Knowledge on Docker and Kubernetes.* Observability tools like Datadog, Dynatrace or similar.* Implement and maintain CI/CD pipelines.* Incident response and running blameless post-mortems.* Proficient in Git workflows.## ## ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Falkirk, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Nottingham, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Yelverton, Devon, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Denbigh, Denbighshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
New Tredegar, Caerphilly, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Llangennech, Carmarthenshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Forfar, Angus, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Dunfermline, Fife, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Alloa, Clackmannanshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Heckmondwike, West Yorkshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Wadhurst, East Sussex, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Walsall, West Midlands, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Newport Pagnell, Buckinghamshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Potters Bar, Hertfordshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Bathgate, West Lothian, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Magherafelt, Co. Londonderry, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Lochgilphead, Argyll & Bute, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Downpatrick, Co. Down, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Kilmarnock, East Ayrshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...