26 to 50 of 79 Service-Level Objective Jobs in London

Principal DevOps Engineer

Location
Greater London, England, United Kingdom
/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable) Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards) Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning Mentor engineers, review designs/code, and raise overall engineering ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
standards, providing deep insights into system health and performance. GCO Feature Implementation: Aid with the design of GCO dashboards, log‐based metrics, alerts, and SLO/SLI tracking to provide comprehensive visibility. Essential Skills Strong understanding of SRE concepts, including SLOs, SLIs, error budgets, and toil reduction. Observability Instrumentation: Hands ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
DevOps Practices AWS, GCP, Azure Python or Go Automation Helm and Terraform Hard Skills Infrastructure Engineering DevOps Automation Containerization Incident Response Root‐Cause Analysis SLO Definition Monitoring and Logging System Design High‐Availability Systems Soft Skills Excellent Communication Collaboration Problem-Solving Ownership Accountability Industry Keywords Generative AI High‐Traffic Platforms ...

Head Of Infrastructure and Cloud

Hiring Organisation
Arbuthnot Latham
Location
London, UK
Employment Type
Full-time
service delivery and operational outcomes. Establish and integrate Site Reliability Engineering (SRE) practices, defining and managing service-level objectives (SLOs), error budgets, and proactive reliability engineering across critical services. Ensure end-to-end service ...

Site Reliability Engineer

Hiring Organisation
Bristow Holland Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum - Offering 100% Work from home
with Development and DevOps teams throughout the application release process Balancing the delivery of new functionality with platform reliability and service-level objectives Identifying bottlenecks and proposing improvements across infrastructure and applications Improving the reliability, quality and time-to-market ...

sre engineer in fintech

Location
Greater London, England, United Kingdom
resources and Helm for Kubernetes application deployments; Implement and manage monitoring, alerting, and logging solutions; Define, measure, and enforce Service Level Objectives and Service Level Indicators; Participate in the on-call rotation, where ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
easy for downstream teams to use.Raise reliability and data quality• Define data contracts, validation rules, freshness expectations, lineage, and service-level objectives for critical datasets.• Implement automated testing, anomaly detection, alerting, and observability across the data lifecycle.• Own production issues through ...

Observability SRE

Location
Greater London, England, United Kingdom
outside, if need be, towards building and maintaining robust, scalable, highly available production systems in accordance with our service level objectives Preventing production incidents but when they do occur, performing effective incident and problem management and RCA to minimize downtime ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
virtual SRE community of practice) that scales reliability improvements across many application flows. Defines and implements standards for: Service cataloging, SLO/SLI frameworks and error budgets, incident response maturity, blameless post-incident reviews, resiliency patterns, capacity, performance, and scalability engineering. Drives service ...

Lead DevOps Engineer

Hiring Organisation
Collinson Group
Location
London, UK
Employment Type
Full-time
automation so routine operational tasks are eliminated, not managed. Observability - Own the observability strategy across the platform using Datadog. Define what good looks like: SLO/SLA dashboards, alerting thresholds, runbooks, and the feedback loops that let teams act on signals before users feel them. Team Leadership & Mentoring - Lead ...

Site Relaibility Engineer

Hiring Organisation
Bristow Holland
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£55,000 - £60,000 per annum
with Development and DevOps teams throughout the application release process Balancing the delivery of new functionality with platform reliability and service-level objectives Identifying bottlenecks and proposing improvements across infrastructure and applications Improving the reliability, quality and time-to-market ...

Architect & Delivery Lead (68018)

Location
Greater London, England, United Kingdom
native architecture patterns, microservices, event‐driven design Cloud platform expertise across AWS, Azure, and/or GCP at enterprise scale SRE principles: SLI/SLO/error budgets, observability, chaos engineering, and resilience patterns IAM and security architecture: Zero Trust, RBAC/ABAC, PAM, compliance frameworks (SOX, GDPR ...

Head Of Infrastructure and Cloud - Internal Applicants Only

Location
Greater London, England, United Kingdom
delivery and operational outcomes. Establish and integrate Site Reliability Engineering (SRE) practices, defining and managing servicelevel objectives (SLOs), error budgets, and proactive reliability engineering across critical services. Ensure end‐to‐end service ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
infrastructure and/or platform engineering focused teams Expertise in Observability and reliability, operational metrics and data analysis Proven track record architecting monitoring frameworks, SLO platforms, and automated response workflows Datadog (or equivalent observabilty tooling like New Relic, Grafana). Proven experience working on large-scale enterprise software applications Experience ...

AWS Technology Lead – AWS Cloud Modernization & Platform Engineering

Hiring Organisation
Infinity Quest
Location
London Area, United Kingdom
automation using Terraform and/or CloudFormation . Establish automated testing, security scanning, code quality and release controls. Define monitoring, logging, tracing, alerting and SLO practices. Drive DevSecOps and continuous improvement. Technical & Customer Leadership Act as the primary technical contact for customer architects and engineering/programme stakeholders. Lead discovery ...

CD&A Data Engineer

Location
Greater London, England, United Kingdom
technical problems. Experience working with AI-assisted software engineering and rapid prototyping techniques. What will help you on the job Understanding of SRE concepts (SLO/SLI’s, uptime, alerts, MTTR). Experience with GCP DevOps, GitHub, YAML, AI assisted coding, and basic repo/pull request workflows. Exposure ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement … reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£67,547 - £83,778 per annum
Work with teams to define and embed Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets in a pragmatic and user-focused way. Support the development and continual improvement … reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning ...

platform engineer in cloud platforms

Location
Greater London, England, United Kingdom
improve availability and reduce operational risk; Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement; Experience establishing service-level objectives, monitoring, alerting, and operational practices; Experience providing technical leadership across multiple engineering teams or a wider engineering organisation ...

Principal Platform Engineer (12 Month FTC)

Location
Greater London, England, United Kingdom
practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service-level objectives, monitoring, alerting, and operational practices. Technical Leadership Providing technical leadership across multiple engineering teams or a wider engineering ...

Principal Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service-level objectives, monitoring, alerting, and operational practices. Technical Leadership Providing technical leadership across multiple engineering teams or a wider engineering ...

Observability SME/Architect/Consultant

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
capabilities across critical business services.The role is responsible for aligning telemetry, monitoring, service health, dashboards, Service Level Objectives (SLOs), and resilience reporting with Operational Resilience outcomes. Working closely with technology, operations, service management ...

Senior Data Engineer

Hiring Organisation
Just Eat Takeaway.com
Location
London, UK
retail media data platform, including attribution pipelines supporting campaign performance and incrementality measurement. Contribute to Retail Media event tracking infrastructure, schema governance, SLA/SLO monitoring, and data quality validation across the unified events pipeline. Support the development of a semantic layer (Cube.dev/Helix) bridging BigQuery Gold tables ...

Principal Splunk Architect

Hiring Organisation
Bank of America
Location
London, UK
Employment Type
Full-time
troubleshoot and resolve complex distributed-system issues in production environments. Strong experience supporting highly available, mission-critical platforms with demanding service-level objectives. Desirable Skills & ExperienceBanking or Financial Services experience. Experience within Application Production Support, Production Engineering, or Site Reliability Engineering teams. ...

Cloud Solutions Architect

Location
Greater London, England, United Kingdom
each of which can be managed by individuals or teams in a scalable manner.* Ensure systems and application have efficient and actionable monitoring with SLO aligned with business goals.Qualifications/Requirements:* You have expert knowledge of Microsoft Azure products and capabilities.* You have proven experience with all aspects of Cloud ...