26 to 50 of 51 Service-Level Objective Jobs in England

Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office)

Hiring Organisation
Hamilton Barnes
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP 450 - 475 Daily
health across Azure services Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience Drive ...

Senior Software Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automation. Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first-contact resolution, and SLO adherence. Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile-aligned ways ...

Cloud Solutions Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
each of which can be managed by individuals or teams in a scalable manner.* Ensure systems and application have efficient and actionable monitoring with SLO aligned with business goals.Qualifications/Requirements:* You have expert knowledge of Microsoft Azure products and capabilities.* You have proven experience with all aspects of Cloud ...

Inference Engineering - Platform Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
work closely with the hardware and orchestration teams to expose heterogeneous backends reliably through the platform. Who You'll Build Service-level objectives, monitoring, alerting, and observability On-call and incident response: runbooks, escalation, blameless postmortems, and follow-through Capacity planning ...

Cloud Solutions Architect

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
each of which can be managed by individuals or teams in a scalable manner. Ensure systems and application have efficient and actionable monitoring with SLO aligned with business goals. Qualifications and Requirements Expert knowledge of Microsoft Azure products and capabilities. Proven experience with all aspects of Cloud service ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
underpinned by a rigorous Site Reliability Engineering (SRE) mindset. The successful candidates will be instrumental in defining and upholding Service Level Objectives (SLOs) and Service Level Indicators (SLIs), implementing effective monitoring and alerting … environments are reproducible, version‐controlled, and auditable. Proactively manage the capacity of the infrastructure to consistently meet or exceed Service Level Objectives for latency, error rates, and availability. Incident Response and Post‐Mortems: Act as first‐line responders for critical system ...

Senior Engineering Manager - Enterprise Trust & Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
enterprise security reviews. Stability & Operations (S&O) keeps the platform dependable as it scales: incident management and the severity model, service-level objectives (SLOs) and error budgets, observability and DORA (DevOps Research and Assessment) delivery metrics, the on-call rotation ...

Solution Architect - Network Automation

Hiring Organisation
Jobleads-UK
Location
Knutsford, England, United Kingdom
monitoring/logging, security and IAM. Proficiency in Agile Methodologies Scrum/Kanban, backlog and workflow mgmt. and SRE specific reporting (MTTR, deployment frequency, SLO and others) The four LEAD behaviours are: L - Listen and be authentic E - Energise and inspire A - Align across the enterprise D - Develop others.. Your ...

Senior Software Engineer – AI

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
external partners, and lead or start communities of practice. Demonstrate a production first attitude, continuously considering observability and maintaining Service Level Objectives, while delivering change at pace. Embrace emerging technologies and trends, and share insights with the organisation, while developing ...

Senior Engineering Manager - Enterprise Trust & Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security reviews. Stability & Operations (S&O) keeps the platform dependable as it scales: incident management and the severity model, servicelevel objectives (SLOs) and error budgets, observability and DORA (DevOps Research and Assessment) delivery metrics, the on‐call rotation ...

Staff Software Engineer, AI Reliability Engineering

Hiring Organisation
Jobleads-UK
Location
England, United Kingdom
Anthropic offer this kind of dynamic, cross-cutting exposure to the systems that matter most. Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems ...

Senior Site Reliability Engineer - Cloud-Native Leader

Hiring Organisation
Jobleads-UK
Location
Colchester, England, United Kingdom
with cloud infrastructure, automation, and operational leadership to deliver resilient systems that enable rapid product delivery. You will own production reliability, define SLI/SLO, lead blameless postmortems, and automate toil. #J-18808-Ljbffr ...

Staff AI Software Engineer – AI Engineering Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
hiring processes and training engineers up to Staff standard. Demonstrate a production first attitude, continuously considering observability and maintaining Service Level Objectives, while delivering change at pace. Embrace emerging technologies and trends, share insights with the organisation, and develop and maintain ...

Platform & Workplace Engineering Director

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
modern enterprise engineering — and the breadth to lead both. Hard Skills Cloud Infrastructure Observability Site Reliability Engineering Automation Performance Metrics Incident Response Practices SLO Management Identity Hygiene Endpoint Security SaaS Management Soft Skills Stakeholder Communication Team Development Performance Standards Setting #J-18808-Ljbffr ...

Senior Data Platform Engineer (Data Lake and Catalog)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
testable, and observable software for our data lake platform.Lead by example in engineering quality through clean, well-tested code and strong coding practices.Promote an SLO-driven culture by contributing to reliability goals, observability, and incident learnings.Partner closely with engineers and other stakeholders to clarify requirements and deliver effectively.Break down complex ...

Senior Data Platform Engineer (Data Lake and Catalog)

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
observable software for our data lake platform. Lead by example in engineering quality through clean, well-tested code and strong coding practices. Promote an SLO-driven culture by contributing to reliability goals, observability, and incident learnings. Partner closely with engineers and other stakeholders to clarify requirements and deliver effectively. Break ...

Senior SRE: Lead Reliability for Scalable Platforms

Hiring Organisation
Jobleads-UK
Location
Watford, England, United Kingdom
provide technical leadership for reliability across the digital estate, ensuring high availability and performance during normal operations and peak lottery events. You will own SLO/SLI definitions, incident leadership, and a roadmap for platform maturity, working with ECS/EKS, Terraform-based automation, and observability tooling. #J-18808-Ljbffr ...

Senior Data Platform Engineer (Data Lake and Catalog)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
observable software for our data lake platform. Lead by example in engineering quality through clean, well‐tested code and strong coding practices. Promote an SLO‐driven culture by contributing to reliability goals, observability, and incident learnings. Partner closely with engineers and other stakeholders to clarify requirements and deliver effectively. Break ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
turn root-cause findings into lasting reliability improvements. Collaborate with application development teams and other stakeholders to meet internal Service Level Objectives and customer-facing Service Level Agreements while improving the tools, frameworks … with related technical knowledge. Experience designing and implementing scalable, resilient, and well-tested distributed systems. Experience with Service Level Objectives, Service Level Agreements, monitoring, alerting, capacity planning, incident management, or disaster-recovery testing. ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
turn root‐cause findings into lasting reliability improvements. Collaborate with application development teams and other stakeholders to meet internal Service Level Objectives and customer‐facing Service Level Agreements while improving the tools, frameworks … with related technical knowledge. Experience designing and implementing scalable, resilient, and well‐tested distributed systems. Experience with Service Level Objectives, Service Level Agreements, monitoring, alerting, capacity planning, incident management, or disaster‐recovery testing. ...

Senior Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Hounslow, England, United Kingdom
Governance, Risk and Standards Ensure compliance with governance, security and engineering standards. Own platform risks and mitigation plans. Establish service level objectives, operational KPIs and reliability targets. Ensure change is delivered safely and effectively across production environments. Promote data-driven decision ...

Cyber Infrastructure Engineer, Professional Services

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
design and planning. **What will I be doing:** You will be regularly completing deployment projects on or before expected Service Level Objectives (SLOs) and integrating new systems into existing network architecture. On the support side, you'll efficiently manage customer support ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers Demonstrates a high level of technical expertise within one or more … technical processes with emerging depth in one or more technical disciplines Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform ...

Site Reliability Engineer (SRE)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
tightly with our Product Engineering teams Following SRE practices and maintaining high standards of compliance Implementing a new standard of observability utilising SLI/SLO/Error Budgets Continually evolving our observability platforms for greater coverage Using a code-first approach to build and changes to reduce TOIL Advocating … with a keen interest to learn and grow as a Site Reliability Engineer Observability product experience (eg Datadog) Managing services using SLI/SLO & Error Budgets Experience with AWS or other cloud providers Experience in HA environments Automation skills through Terraform, Python, Bash or similar Good SRE skills with ...

Software Engineering III - AI/ML Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
running production incident calls and managing incident resolution. Experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others Strong understanding …/SLO/SLA and Error Budgets Hands‐on experience using enterprise‐authorized AI‐assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI‐generated outputs for correctness, performance, and security. Understanding of responsible ...