1 to 25 of 159 Service-Level Objective Jobs in the UK

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting, and operational practices Experience providing technical leadership across multiple engineering teams or a wider engineering organisation … Soft Skills Technical Leadership Mentoring Influencing Stakeholders Knowledge Sharing Continuous Improvement Industry Keywords Platform Engineering Cloud-Native Architectures Engineering Workflows Service-Level Objectives Operational Practices Tools & Technologies Automated Cloud Provisioning Self-Service Tooling Engineering Standards Operational Processes ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
team members, provide timely and actionable feedback, and actively develop engineers' skills including proficiency with AI-augmented workflowsEstablish and monitor service-level objectives, reliability metrics, and engineering efficiency indicators to maintain high engineering craft and product qualityCommunicate production system health, incident … tooling or automation frameworks within an engineering team to expand operational scope and reduce toilExperience managing on-call rotations, defining service-level objectives, and implementing observability and alerting improvements at scaleFamiliarity with container orchestration, service mesh architectures ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
Puppet, Ansible). Advanced Kubernetes (Karpenter, KEDA, HPA/VPA, Service Mesh). Observability tooling (Datadog, OpenTelemetry) and SLI/SLO design. Security and identity tooling (SSO, IAM, PKI). Queue or streaming design patterns (SQS, Kafka). Commercial or low-level ...

Site Reliability Software Engineer (Hybrid)

Location
Greater London, England, United Kingdom
operations teams. Support healthcare integrations involving EHR/EMR platforms, APIs, data feeds, file transfers, or interface engines. Help define service-level objectives, uptime expectations, escalation procedures, and operational runbooks. Ensure systems are designed and maintained with healthcare security, privacy ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
alerting, and logging solutions to ensure clear system visibility and proactive issue identification.* Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call ...

Senior DevOps Engineer

Location
United Kingdom
requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application incidents. Coordinate emergency changes, hot fixes and rollbacks. Lead ...

Senior DevOps Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Permanent
requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application incidents. Coordinate emergency changes, hot fixes and rollbacks. Lead ...

Software Engineer III - SRE

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, UK
Employment Type
Full-time
applications and platforms. Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customersCollaborate with development teams and stakeholders to enhance system reliability, scalability, and performance. Utilize ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
error budgets, and blameless post-mortems) across a large engineering organization. Key Responsibilities Partner with engineering leadership to establish service level objectives (SLOs), service level indicators (SLIs), and error budgets. Collaborate with product ...

Service Manager, Site Reliability Engineering

Location
Belfast City District, Northern Ireland, United Kingdom
customer expectations.* Establish, monitor, and report on Service Level Agreements (SLAs), Service Level Objectives (SLOs), Error Budgets, availability, performance, and operational KPIs to drive service excellence.* Lead Major Incident ...

Senior Platform Engineer GCP & AWS

Location
Watford, England, United Kingdom
Google Cloud Operations suite). Ensure full telemetry visibility across GCP and AWS workloads to proactively manage platform health and service-level objectives (SLOs). Governance & Continuous Improvement Create and maintain comprehensive documentation for our infrastructure, processes, and tools in Confluence. ...

Director of DevOps & SRE

Location
Greater London, England, United Kingdom
changes, ensuring appropriate change management and CAB governance, with occasional off-hours support for the teams. Observability and service levels. SLO/SLA and monitoring strategy across Grafana and our wider observability tooling, building alert coverage that is meaningful rather than noisy. AI in engineering operations. ...

Senior / Lead Site Reliability Engineer

Location
Watford, England, United Kingdom
ownership & reporting Own and prioritise the SRE backlog, balancing: Reliability improvements Technical debt Automation opportunities to reduce/offload toil Produce structured reporting covering: SLO performance Incident trends and MTTR Platform health and risk areas Provide clear updates to engineering leadership and business stakeholders Collaboration & culture Embed SRE practices across ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core ...

Senior Site Reliability Engineer

Hiring Organisation
Morgan McKinley
Location
City of London, London, United Kingdom
Optimization & Reliability: Deep-dive into system internals to tune low-latency Windows infrastructure, manage resource distribution, and establish actionable Service Level Objectives (SLOs). L3 Escalation & RCA: Serve as an ultimate technical escalation point for production incidents, leading detailed Root Cause ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
cause support or service desk workflows. Knowledge of observability and reliability practices, including SRE principles, servicelevel objectives, alert tuning, capacity planning and production readiness reviews. Experience operating SaaS products for financial services, enterprise technology or other ...

Senior Data Platform Engineer - Data Enablement

Location
Greater London, England, United Kingdom
that we uphold data subject rights for our customers Experience introducing a data governance & observability stack enabling rich data lineage, data contracts, SLA/SLO, tagging and data quality monitoring capabilities both on our own platform but also for data owners Proven experience designing, building, and scaling data platforms ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
production automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents ...

Principal AI Platform Engineer

Location
Cambridge, England, United Kingdom
provisioning, configuration, upgrades and lifecycle management using infrastructure-as-code and GitOps patterns. Reliability, observability and support: Define and implement service-level indicators, service-level objectives, alerting, dashboards, runbooks and support workflows. Incident ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
similar assurance frameworks Preferred: AI, Agentic AI, AIOps or intelligent automation experience Preferred: Observability and reliability practices, including SRE principles, service-level objectives, alert tuning, capacity planning and production readiness reviews Preferred: SaaS experience for financial services, enterprise technology or regulated ...

SRE

Location
Hove, England, United Kingdom
Establish and enforce SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets. Define and implement an AIOps roadmap to enhance operational intelligence and automation. Automate repetitive operational ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Ansible and Kubernetes. Define, implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability ...

Test Environment Manager

Location
Greater London, England, United Kingdom
pipelines to enable on-demand, self-service environment delivery. Reliability & Observability Define and maintain Service Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning ...