51 to 72 of 72 Service-Level Objective Jobs in the UK excluding London

Lead Site Reliability / DevOps Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customersDemonstrates a high level of technical expertise within one or more technical … least one programming language such as (e.g., Java, Python, Go, etc.)Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab ...

Senior Lead Site Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
with team members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgets Designs, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/… more programming languages (e.g., Java, Python, Go, etc.) Advanced proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc. ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers Demonstrates a high level of technical expertise within one or more … least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab ...

Senior Lead SRE: Reliability, Observability & Resiliency

Location
Auchentibber, Scotland, United Kingdom
with team members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgets Designs, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/… more programming languages (e.g., Java, Python, Go, etc.) Advanced proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc. ...

Principal Data Engineer

Hiring Organisation
Ronald James
Location
Newcastle upon Tyne, UK
Employment Type
Full-time
technical roadmap for the organisation's data platformEnsure systems meet standards for performance, security, maintainability, and reliabilityDrive progress toward Service Level Objectives (SLOs)Collaborate with technical leads and architects to shape data solutions that meet product and business needsSupport cross-functional ...

Director of Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
/or legacy-to-modern transformations. Experience building paved road? reliability capabilities (shared libraries, templates, tooling) that scale across many teams. Familiarity with SLO programs and operational readiness practices at scale. ABOUT US JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses ...

Lead Cloud Site Reliability Engineer

Location
Manchester, England, United Kingdom
promoting effective root cause analysis and continuous service improvement. Champion Site Reliability Engineering practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Drive automation initiatives … practices, including metrics, logging and distributed tracing. Incident management, problem management and service reliability improvement. Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Automation and reducing ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
promoting effective root cause analysis and continuous service improvement. Champion Site Reliability Engineering practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Drive automation initiatives … practices, including metrics, logging and distributed tracing. Incident management, problem management and service reliability improvement. Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. Automation and reducing ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Glasgow, Scotland, United Kingdom
durable remediation and preventative actions. Own reliability reporting and operational governance by tracking key performance indicators — including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate — and producing executive-ready summaries. … root cause analysis, problem management, and operational governance in complex production environments. Strong knowledge of reliability engineering concepts including service level objectives and indicators, error budgets, capacity planning, resilience patterns, and observability. Proven ability to influence across teams, drive cross-functional ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty. Proficiency in shell ...

Program Strategy Manager

Location
Reading, England, United Kingdom
services or technologies. Manage customer escalations and advocate internally for resolution. Monitor customer cases for accuracy, progress and compliance with service-level objectives. Help customers mature their security posture and demonstrate measurable program value. Ensure delivery aligns with service ...

Enterprise Integration Product Manager

Hiring Organisation
Experis
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£480 - £530/day
lifecycle management (upgrade programs, standardization, deprecation, tech debt). Strong operational mindset: incident/problem management, capacity planning, resilience (HA/DR), SLI/SLO-based performance management. Strong stakeholder management and ability to influence across multiple teams and senior stakeholders. Experience managing large-scale, multi-region deployments of messaging ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customersDemonstrates a high level of technical expertise within one or more technical … more technical disciplinesProficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or SplunkProficiency in continuous integration and continuous delivery tools ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers Demonstrates a high level of technical expertise within one or more … more technical disciplines Proficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Proficiency in continuous integration and continuous ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers Demonstrates a high level of technical expertise within one or more … technical processes with emerging depth in one or more technical disciplines Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
scalable, reliable, and observable infrastructure solutions that meet the firm's availability and performance standards Define and enforce service level objectives, error budgets, and reliability targets in partnership with engineering and product stakeholders Drive incident response, root cause analysis, and post … automation and tooling development Experience defining and managing service level indicators, service level objectives, and error budgets in production environments Strong background in observability tooling, including metrics, logging, and distributed tracing platforms Demonstrated ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty. Proficiency in shell ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
Omnicell’s Enterprise Reliability Framework, Including Service Level Indicators (SLIs) Service Level Objectives (SLOs) Error Budgets Reliability Design Standards Production Readiness Reviews Capacity Planning Models Failure Mode Analysis Reliability Scorecards Engineering Guardrails Partner … observability, and automation capabilities. Identify engineering opportunities to improve platform resilience. Develop a three‐year Site Reliability Engineering roadmap. First Six Months Establish enterprise SLO, SLI, and Error Budget framework. Launch Production Engineering practice. Standardize Production Readiness Reviews. Define enterprise observability engineering standards. Recruit key Site Reliability Engineering leaders. Publish ...

Director, Engineering IT Operations

Hiring Organisation
Cirrus Logic
Location
Edinburgh, UK
Employment Type
Full-time
reliabilityDevelop capacity plans and long-range infrastructure roadmaps to support continued growth in compute, storage, and engineering demandEstablish and maintain service-level expectations, performance metrics, and continuous improvement practices across engineering IT operationsWho you areLinux InfrastructureExtensive experience architecting, deploying, and supporting Linux … engineers in demanding technical environmentsOperational ExcellenceDemonstrated success leading:24x7 production operationsIncident management and escalation processesRoot cause analysis and corrective action planningChange management programsSLA and SLO ownershipDisaster recovery planningSecurity and compliance initiativesCross-Functional CollaborationStrong ability to work effectively across: Silicon engineeringCAD and EDA teamsIT and securityProgram managementExecutive leadershipProven ability to translate ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service … Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery practices and tooling ...

Lead SRE - AWS Platform

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service … Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery practices and tooling ...

Software Engineer III - Backend, Fanatics Markets

Location
Leeds, England, United Kingdom
chain and off-chain data sources, building reliable pipelines for blockchain event ingestion and querying Collaborate on reliability practices including on-call participation, SLO definition, alerting, and incident response Conduct thorough code reviews, raising the bar on code quality, readability, and maintainability Leverage AI tools to accelerate development velocity while …/CD pipelines, automated testing, and safe deployment strategies for backend services Experience working with production reliability practices (monitoring, logging, alerting, and basic SLO concepts) Demonstrated experience using AI tools (Claude Code, Cursor, Copilot, etc.) to ship production code Can demonstrate specific examples of workflow improvements from AI-assisted development ...