1 to 25 of 79 Permanent Service-Level Objective Jobs in London

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting, and operational practices Experience providing technical leadership across multiple engineering teams or a wider engineering organisation … Soft Skills Technical Leadership Mentoring Influencing Stakeholders Knowledge Sharing Continuous Improvement Industry Keywords Platform Engineering Cloud-Native Architectures Engineering Workflows Service-Level Objectives Operational Practices Tools & Technologies Automated Cloud Provisioning Self-Service Tooling Engineering Standards Operational Processes ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
team members, provide timely and actionable feedback, and actively develop engineers' skills including proficiency with AI-augmented workflowsEstablish and monitor service-level objectives, reliability metrics, and engineering efficiency indicators to maintain high engineering craft and product qualityCommunicate production system health, incident … tooling or automation frameworks within an engineering team to expand operational scope and reduce toilExperience managing on-call rotations, defining service-level objectives, and implementing observability and alerting improvements at scaleFamiliarity with container orchestration, service mesh architectures ...

Site Reliability Software Engineer (Hybrid)

Location
Greater London, England, United Kingdom
operations teams. Support healthcare integrations involving EHR/EMR platforms, APIs, data feeds, file transfers, or interface engines. Help define service-level objectives, uptime expectations, escalation procedures, and operational runbooks. Ensure systems are designed and maintained with healthcare security, privacy ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
alerting, and logging solutions to ensure clear system visibility and proactive issue identification.* Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call ...

Director of DevOps & SRE

Location
Greater London, England, United Kingdom
changes, ensuring appropriate change management and CAB governance, with occasional off-hours support for the teams. Observability and service levels. SLO/SLA and monitoring strategy across Grafana and our wider observability tooling, building alert coverage that is meaningful rather than noisy. AI in engineering operations. ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core ...

Senior Site Reliability Engineer

Hiring Organisation
Morgan McKinley
Location
City of London, London, United Kingdom
Optimization & Reliability: Deep-dive into system internals to tune low-latency Windows infrastructure, manage resource distribution, and establish actionable Service Level Objectives (SLOs). L3 Escalation & RCA: Serve as an ultimate technical escalation point for production incidents, leading detailed Root Cause ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
cause support or service desk workflows. Knowledge of observability and reliability practices, including SRE principles, servicelevel objectives, alert tuning, capacity planning and production readiness reviews. Experience operating SaaS products for financial services, enterprise technology or other ...

Senior Data Platform Engineer - Data Enablement

Location
Greater London, England, United Kingdom
that we uphold data subject rights for our customers Experience introducing a data governance & observability stack enabling rich data lineage, data contracts, SLA/SLO, tagging and data quality monitoring capabilities both on our own platform but also for data owners Proven experience designing, building, and scaling data platforms ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
production automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
similar assurance frameworks Preferred: AI, Agentic AI, AIOps or intelligent automation experience Preferred: Observability and reliability practices, including SRE principles, service-level objectives, alert tuning, capacity planning and production readiness reviews Preferred: SaaS experience for financial services, enterprise technology or regulated ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Ansible and Kubernetes. Define, implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability ...

Test Environment Manager

Location
Greater London, England, United Kingdom
pipelines to enable on-demand, self-service environment delivery. Reliability & Observability Define and maintain Service Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
Ansible and Kubernetes. Define, implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability ...

Engineer - Site Reliability Engineering

Location
Greater London, England, United Kingdom
people with a passion to learn, and who bring a continuous improvement mentality to our team!* SREs maintain Service Level Objectives for the systems they own. Constantly measuring and improving availability, latency, and overall system health is at the core ...

Site Reliability Engineer

Hiring Organisation
Trainline
Location
London, UK
Employment Type
Full-time
business goalsOur Tech Stack AWSNew RelicELK stackGrafanaIncident.ioDocker, ECSTerraformGithub ActionsWe'd love to hear from you if you have...Experience of SRE concepts such as SLI, SLO and error budgets. Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similarExperience working with cloud providers (preferably ...

Lead SRE - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
building reliable infrastructure and tooling that expedites feature development. Develop meaningful service metrics, user journeys, service-level indicators, service-level objectives, error budgets, dashboards, and actionable alerts. Engage with ...

Lead SRE - Chase UK

Location
Westminster, West End, United Kingdom
building reliable infrastructure and tooling that expedites feature development. Develop meaningful service metrics, user journeys, service-level indicators, service-level objectives, error budgets, dashboards, and actionable alerts. Engage with ...

Lead SRE - Chase UK

Location
Greater London, England, United Kingdom
building reliable infrastructure and tooling that expedites feature development. Develop meaningful service metrics, user journeys, service-level indicators, service-level objectives, error budgets, dashboards, and actionable alerts. Engage with ...

Site Reliability Engineer

Hiring Organisation
Infinity Quest
Location
City of London, London, United Kingdom
leading blameless post-incident reviews Define and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across services Scope technical projects and break them down into user stories and tasks, driving them to completion ...

DevOps Team Manager

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
workloads, with the ability to govern standards and technical gates. A track record of operating production, multi-tenant SaaS at scale, including SLA/SLO ownership, major-incident command, root-cause analysis, disaster recovery and service improvement. Willingness and ability to participate personally ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
Proficiency with database development (SQL Server, PostgreSQL, MySQL, etc) Proficiency with defining, right-sizing, tracking, and reporting on Service Level Objectives (SLOs), Service Level Indicators (SLIs), system availability, and the progress and outcomes ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
Dynatrace into CI/CD pipelines (Jenkins, GitLab, GitHub Actions) and "GitOps" workflows (ArgoCD, Flux).* Help customers implement Service Level Objectives (SLOs) and Error Budgets within their cloud-native stacks to drive SRE maturity.4. Ecosystem Advocacy:* Stay at the forefront ...

Site Reliability Engineer / Senior Engineer

Location
Greater London, England, United Kingdom
support a wide ranging CSR programme + 2 days’ volunteering leave per year Your key responsibilities Defining and implementing Service Level Objectives (SLOs) and embed Site Reliability Engineering (SRE) principles to improve service reliability and operational excellence ...