126 to 150 of 158 Service-Level Objective Jobs in the UK

Global SVP, Platform Operations & AI-Driven Reliability

Location
Greater London, England, United Kingdom
consolidate and lead the production platform across SRE, deployment, and platform engineering. You will own the CI/CD pipeline, incident management, and the SLO/SLI framework to raise reliability for critical trading services worldwide. You will build a leadership layer, expand global engineering hubs, and drive AI-driven ...

Director of Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
/or legacy-to-modern transformations. Experience building paved road? reliability capabilities (shared libraries, templates, tooling) that scale across many teams. Familiarity with SLO programs and operational readiness practices at scale. ABOUT US JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses ...

Lead Cloud Site Reliability Engineer

Location
West of England, England, United Kingdom
promoting effective root cause analysis and continuous service improvement. Champion Site Reliability Engineering practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Drive automation initiatives … practices, including metrics, logging and distributed tracing. Incident management, problem management and service reliability improvement. Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Automation and reducing ...

Lead Cloud Site Reliability Engineer

Location
Manchester, England, United Kingdom
promoting effective root cause analysis and continuous service improvement. Champion Site Reliability Engineering practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Drive automation initiatives … practices, including metrics, logging and distributed tracing. Incident management, problem management and service reliability improvement. Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Automation and reducing ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
promoting effective root cause analysis and continuous service improvement. Champion Site Reliability Engineering practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Drive automation initiatives … practices, including metrics, logging and distributed tracing. Incident management, problem management and service reliability improvement. Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Automation and reducing ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Glasgow, Scotland, United Kingdom
durable remediation and preventative actions. Own reliability reporting and operational governance by tracking key performance indicators — including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate — and producing executive-ready summaries. … root cause analysis, problem management, and operational governance in complex production environments. Strong knowledge of reliability engineering concepts including service level objectives and indicators, error budgets, capacity planning, resilience patterns, and observability. Proven ability to influence across teams, drive cross-functional ...

Technical Support Team Manager, EMEA

Location
Greater London, England, United Kingdom
team to provide consistently high-quality service which drives customer loyalty. Meet or exceed established service level objectives to ensure client satisfaction including staffing, volume, response times and quality of service. Act as a subject matter expert ...

Cyber Infrastructure Engineer, Professional Services

Location
Greater London, England, United Kingdom
design and planning. What will I be doing: You will be regularly completing deployment projects on or before expected Service Level Objectives (SLOs) and integrating new systems into existing network architecture. On the support side, you'll efficiently manage customer support ...

Tech Lead Manager - Lakebase

Location
Greater London, England, United Kingdom
technical work. The team you will lead (one of): Lakebase Platform Reliability (LPR) — Constantly evolve Lakebase reliability: cross‐functional infrastructure, observability, SLI/SLO definition and measurement, and higher operational efficiency without humans-in‐the‐loop, for both internal and customer‐facing systems, by building AI‐powered automation and services. ...

Tech Lead Manager - Lakebase

Hiring Organisation
DataBricks
Location
London, UK
Employment Type
Full-time
technical work. The team you will lead (one of):Lakebase Platform Reliability (LPR) — Constantly evolve Lakebase reliability: cross-functional infrastructure, observability, SLI/SLO definition and measurement, and higher operational efficiency without humans-in-the-loop, for both internal and customer-facing systems, by building AI-powered automation and services. ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
Balance feature delivery and reliability using well-defined service level indicators, service level objectives, and error budgets. Improve the reliability of Voice/UC platforms and integrations, with attention to real-time signaling … speech/voice AI, machine-learning services, or LLM-based applications is preferred. Ability to use metrics, logs, traces, and service-level indicators to diagnose complex production issues and drive measurable reliability improvements. A proactive approach to spotting problems, areas for improvement ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty. Proficiency in shell ...

Program Strategy Manager

Location
Reading, England, United Kingdom
services or technologies. Manage customer escalations and advocate internally for resolution. Monitor customer cases for accuracy, progress and compliance with service-level objectives. Help customers mature their security posture and demonstrate measurable program value. Ensure delivery aligns with service ...

Senior ML Engineer

Location
Greater London, England, United Kingdom
hardware and ensure scalability for exponential user growth. Serve as the technical guardian for the model's quality and Service Level Objectives (SLOs). Provide a hands-on solution architecture for the core ML infrastructure. Select, evaluate, and implement state ...

Enterprise Integration Product Manager

Hiring Organisation
Experis
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£480 - £530/day
lifecycle management (upgrade programs, standardization, deprecation, tech debt). Strong operational mindset: incident/problem management, capacity planning, resilience (HA/DR), SLI/SLO-based performance management. Strong stakeholder management and ability to influence across multiple teams and senior stakeholders. Experience managing large-scale, multi-region deployments of messaging ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customersDemonstrates a high level of technical expertise within one or more technical … more technical disciplinesProficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or SplunkProficiency in continuous integration and continuous delivery tools ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers Demonstrates a high level of technical expertise within one or more … more technical disciplines Proficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Proficiency in continuous integration and continuous ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers Demonstrates a high level of technical expertise within one or more … technical processes with emerging depth in one or more technical disciplines Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
scalable, reliable, and observable infrastructure solutions that meet the firm's availability and performance standards Define and enforce service level objectives, error budgets, and reliability targets in partnership with engineering and product stakeholders Drive incident response, root cause analysis, and post … automation and tooling development Experience defining and managing service level indicators, service level objectives, and error budgets in production environments Strong background in observability tooling, including metrics, logging, and distributed tracing platforms Demonstrated ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty. Proficiency in shell ...

SVP, Global Head of Platform Operations

Location
Greater London, England, United Kingdom
raising reliability for customers who cannot tolerate downtime. You will own the end-to-end deployment pipeline, production observability and incident management, and the SLO/SLI framework that defines what "good" looks like for every customer-facing service. You will build the leadership layer underneath you, shape … engineering organization. Key Deliverables : Low Mean Time to Resolve (MTTR), minimal production incidents, and near-perfect achievement of Service Level Objectives (SLOs). Fast time-to-market and high throughput velocity through the CI/CD pipeline. Key Responsibilities 1. Delivery ...

Technical Operations Manager - 6 Month FTC – Kings Cross, London

Location
Greater London, England, United Kingdom
addressing issues to minimize downtime. Manage the implementation and maintenance of Service Level Objectives (SLO's) and Service Level Agreements (SLA’s) with UMG Business Units, other IT groups, and vendors as appropriate. … This includes measuring and reporting on these SLA's and SLO’s on a regular and scheduled basis. Drive automation for all repeatable processes by implementing AI and SRE methodologies to reduce toil in daily operations. Major Incident Management Take ownership, lead and coordinate major incident response efforts, acting ...

Senior AI Quality Engineer

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
production issues using our observability stack (Datadog, Sentry, Grafana) - tying test coverage back to real customer impactEnsure teams have Service Level Objectives set up and are achieving themRun targeted exploratory testing on high-risk releases when it's the right callWhat … contract testing experience, with a point of view on test data and environment strategyHave set up and used Service Level Objectives to define a contract of system reliability with stakeholders, and driven measurable improvements against themNative AI-assisted working style. ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
Omnicell’s Enterprise Reliability Framework, Including Service Level Indicators (SLIs) Service Level Objectives (SLOs) Error Budgets Reliability Design Standards Production Readiness Reviews Capacity Planning Models Failure Mode Analysis Reliability Scorecards Engineering Guardrails Partner … observability, and automation capabilities. Identify engineering opportunities to improve platform resilience. Develop a three‐year Site Reliability Engineering roadmap. First Six Months Establish enterprise SLO, SLI, and Error Budget framework. Launch Production Engineering practice. Standardize Production Readiness Reviews. Define enterprise observability engineering standards. Recruit key Site Reliability Engineering leaders. Publish ...

Director, Site Reliability Engineering

Hiring Organisation
Omnicell
Location
United Kingdom, UK
Employment Type
Full-time
ReliabilityOwn Omnicell's enterprise reliability framework, including: Service Level Indicators (SLIs)Service Level Objectives (SLOs)Error BudgetsReliability Design StandardsProduction Readiness ReviewsCapacity Planning ModelsFailure Mode AnalysisReliability ScorecardsEngineering GuardrailsPartner with Product Engineering and Cloud Platform … architecture, observability, and automation capabilities. Identify engineering opportunities to improve platform resilience. Develop a three-year Site Reliability Engineering roadmap. First Six MonthsEstablish enterprise SLO, SLI, and Error Budget framework. Launch Production Engineering practice. Standardize Production Readiness Reviews. Define enterprise observability engineering standards. Recruit key Site Reliability Engineering leaders. Publish ...