24 of 24 Remote/Hybrid Service-Level Objective Jobs in London

Site Reliability Software Engineer (Hybrid)

Location
Greater London, England, United Kingdom
operations teams. Support healthcare integrations involving EHR/EMR platforms, APIs, data feeds, file transfers, or interface engines. Help define service-level objectives, uptime expectations, escalation procedures, and operational runbooks. Ensure systems are designed and maintained with healthcare security, privacy ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
cause support or service desk workflows. Knowledge of observability and reliability practices, including SRE principles, servicelevel objectives, alert tuning, capacity planning and production readiness reviews. Experience operating SaaS products for financial services, enterprise technology or other ...

Senior Data Platform Engineer - Data Enablement

Location
Greater London, England, United Kingdom
that we uphold data subject rights for our customers Experience introducing a data governance & observability stack enabling rich data lineage, data contracts, SLA/SLO, tagging and data quality monitoring capabilities both on our own platform but also for data owners Proven experience designing, building, and scaling data platforms ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
Proficiency with database development (SQL Server, PostgreSQL, MySQL, etc) Proficiency with defining, right-sizing, tracking, and reporting on Service Level Objectives (SLOs), Service Level Indicators (SLIs), system availability, and the progress and outcomes ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
Dynatrace into CI/CD pipelines (Jenkins, GitLab, GitHub Actions) and "GitOps" workflows (ArgoCD, Flux).* Help customers implement Service Level Objectives (SLOs) and Error Budgets within their cloud-native stacks to drive SRE maturity.4. Ecosystem Advocacy:* Stay at the forefront ...

Site Reliability Engineer / Senior Engineer

Location
Greater London, England, United Kingdom
support a wide ranging CSR programme + 2 days’ volunteering leave per year Your key responsibilities Defining and implementing Service Level Objectives (SLOs) and embed Site Reliability Engineering (SRE) principles to improve service reliability and operational excellence ...

Head Of Infrastructure and Cloud

Hiring Organisation
Arbuthnot Latham
Location
London, UK
Employment Type
Full-time
service delivery and operational outcomes. Establish and integrate Site Reliability Engineering (SRE) practices, defining and managing service-level objectives (SLOs), error budgets, and proactive reliability engineering across critical services. Ensure end-to-end service ...

Site Reliability Engineer

Hiring Organisation
Bristow Holland Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum - Offering 100% Work from home
with Development and DevOps teams throughout the application release process Balancing the delivery of new functionality with platform reliability and service-level objectives Identifying bottlenecks and proposing improvements across infrastructure and applications Improving the reliability, quality and time-to-market ...

Site Relaibility Engineer

Hiring Organisation
Bristow Holland
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£55,000 - £60,000 per annum
with Development and DevOps teams throughout the application release process Balancing the delivery of new functionality with platform reliability and service-level objectives Identifying bottlenecks and proposing improvements across infrastructure and applications Improving the reliability, quality and time-to-market ...

Head Of Infrastructure and Cloud - Internal Applicants Only

Location
Greater London, England, United Kingdom
delivery and operational outcomes. Establish and integrate Site Reliability Engineering (SRE) practices, defining and managing servicelevel objectives (SLOs), error budgets, and proactive reliability engineering across critical services. Ensure end‐to‐end service ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
infrastructure and/or platform engineering focused teams Expertise in Observability and reliability, operational metrics and data analysis Proven track record architecting monitoring frameworks, SLO platforms, and automated response workflows Datadog (or equivalent observabilty tooling like New Relic, Grafana). Proven experience working on large-scale enterprise software applications Experience ...

AWS Technology Lead – AWS Cloud Modernization & Platform Engineering

Hiring Organisation
Infinity Quest
Location
London Area, United Kingdom
automation using Terraform and/or CloudFormation . Establish automated testing, security scanning, code quality and release controls. Define monitoring, logging, tracing, alerting and SLO practices. Drive DevSecOps and continuous improvement. Technical & Customer Leadership Act as the primary technical contact for customer architects and engineering/programme stakeholders. Lead discovery ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement … reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£67,547 - £83,778 per annum
Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement … reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning ...

Observability SME/Architect/Consultant

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
capabilities across critical business services.The role is responsible for aligning telemetry, monitoring, service health, dashboards, Service Level Objectives (SLOs), and resilience reporting with Operational Resilience outcomes. Working closely with technology, operations, service management ...

Senior Software Engineer (Application Operations)

Location
Greater London, England, United Kingdom
Improvement Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first‐contact resolution, and SLO adherence. Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile‐aligned ways ...

Senior ML Engineer

Location
Greater London, England, United Kingdom
hardware and ensure scalability for exponential user growth. Serve as the technical guardian for the model's quality and Service Level Objectives (SLOs). Provide a hands-on solution architecture for the core ML infrastructure. Select, evaluate, and implement state ...

SVP, Global Head of Platform Operations

Location
Greater London, England, United Kingdom
raising reliability for customers who cannot tolerate downtime. You will own the end-to-end deployment pipeline, production observability and incident management, and the SLO/SLI framework that defines what "good" looks like for every customer-facing service. You will build the leadership layer underneath you, shape … engineering organization. Key Deliverables : Low Mean Time to Resolve (MTTR), minimal production incidents, and near-perfect achievement of Service Level Objectives (SLOs). Fast time-to-market and high throughput velocity through the CI/CD pipeline. Key Responsibilities 1. Delivery ...

Senior AI Quality Engineer

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
production issues using our observability stack (Datadog, Sentry, Grafana) - tying test coverage back to real customer impactEnsure teams have Service Level Objectives set up and are achieving themRun targeted exploratory testing on high-risk releases when it's the right callWhat … contract testing experience, with a point of view on test data and environment strategyHave set up and used Service Level Objectives to define a contract of system reliability with stakeholders, and driven measurable improvements against themNative AI-assisted working style. ...

Senior Product Engineer (Product Manager)

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
allow Lendable to move faster and scale with confidence. What you'll be doingOwn platform stabilityOwn platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce … hold a line with leadership and pod leads alike. Nice to haveExperience in platform, infrastructure, or developer tooling contexts. Familiarity with SLA/SLO frameworks and incident management at scale. Experience enforcing CI/CD standards across multiple teams or pods. Prior work making "invisible" infrastructure legible to non-technical ...

Senior Product Engineer (Product Manager)

Location
Greater London, England, United Kingdom
move faster and scale with confidence. What you'll be doing Own platform stability Own platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce … hold a line with leadership and pod leads alike. Nice to have Experience in platform, infrastructure, or developer tooling contexts. Familiarity with SLA/SLO frameworks and incident management at scale. Experience enforcing CI/CD standards across multiple teams or pods. Prior work making "invisible" infrastructure legible ...

Technical Product Manager (Superapp)

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
allow Lendable to move faster and scale with confidence. What you'll be doingOwn platform stabilityOwn platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce … hold a line with leadership and pod leads alike. Nice to haveExperience in platform, infrastructure, or developer tooling contexts. Familiarity with SLA/SLO frameworks and incident management at scale. Experience enforcing CI/CD standards across multiple teams or pods. Prior work making "invisible" infrastructure legible to non-technical ...

Principal Data Engineer

Location
City Of London, England, United Kingdom
pattern so they are handover-ready by design. Drive data quality as a first-class, firm-wide concern: establish data contracts, observability, SLA/SLO monitoring, and automated alerting and remediation across ingestion and transformation layers, and hold squads to those standards. Act as the senior technical point of contact … standards, conduct code and design reviews, and grow engineers’ capabilities Strong grasp of data quality practices: data contracts, pipeline observability, SLA/SLO definition, and automated alerting and remediation Solid understanding of SQL transformation patterns and modern tooling such as dbt, alongside experience managing ingestion estates with third-party connectors ...