51 to 75 of 79 Service-Level Objective Jobs in London

Senior Site Reliability Engineer - Cloud & Observability

Location
Greater London, England, United Kingdom
leading global financial markets company is seeking a Senior Engineer in Site Reliability. This role involves maintaining service level objectives, enhancing system reliability, and automating to ensure scalability. With a strong focus on cloud platforms, particularly Azure, candidates should have extensive ...

Senior Software Engineer (Application Operations)

Location
Greater London, England, United Kingdom
Improvement Help define, measure, and report on operational KPIs such as MTTR, percentage resolved at L2/L2.5, escalation rate, first‐contact resolution, and SLO adherence. Continuously assess processes and workflows, delivering improvements that increase efficiency, consistency, and quality; balance reactive demand with proactive improvement work in Agile‐aligned ways ...

Applied AI ML Data Scientist, Vice President - Payments

Location
Greater London, England, United Kingdom
language model services integrated with strategic platforms and downstream consumers across APIs, batch, streaming, and event‐driven patterns, meeting service level objectives Partner with product, operations, risk and control, and technology teams to influence roadmaps, align on requirements, and deliver data ...

Applied AI ML Data Scientist, Vice President - Payments

Location
Greater London, England, United Kingdom
language model services integrated with strategic platforms and downstream consumers across APIs, batch, streaming, and event-driven patterns, meeting service level objectives Partner with product, operations, risk and control, and technology teams to influence roadmaps, align on requirements, and deliver data ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
constructive collaboration with application engineers. Experience supporting business-critical services through an on-call rota. Added Bonus Experience with Honeycomb, OpenTelemetry , Prometheus, Splunk , including SLO-led practices. Pragmatic use of AI-assisted engineering tools to improve quality and productivity. At Zopa we value flexible ways of working. We value face ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
teams on cost centres, profitability, pricing, and operational financial planning. Key measures Operational Excellence & Service Stability measured by: SLA/SLO attainment, incident reduction, MTTR, problem recurrence rate, quality of monitoring/alerting, operational readiness scores. Practice Growth & Commercial Performance measured by: revenue growth, headcount growth ...

Sr. Lead Data Engineer

Location
Greater London, England, United Kingdom
practices and controls. Job Responsibilities Change and release health - Track deployment frequency, change failure rate, lead time, and rollbacks; correlate changes to incidents/SLO impact; influence safer release practices. Capacity, performance, and scalability - Produce capacity forecasts, headroom and hotspot reporting; partner with engineering to validate scaling policies and performance ...

Data Engineer

Location
Greater London, England, United Kingdom
lifecycle analytics. Apply data quality frameworks, lineage, cataloguing, and GDPR/PII compliance practices. Design secure data pipelines with monitoring, alerting, and SLA/SLO management. Translate business and marketing needs into clear technical specifications, communicating technical topics clearly to non‐technical stakeholders. Partner with analytics, marketing ops, and MarTech ...

Staff Software Engineer, AI Reliability Engineering

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
teams at Anthropic offer this kind of dynamic, cross-cutting exposure to the systems that matter most. ResponsibilitiesDevelop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability ...

Senior SRE Engineer: Reliability, Cloud & Automation

Location
Greater London, England, United Kingdom
Site Reliability who will join a driven team focused on system availability, performance, and scalability. Responsibilities include maintaining service level objectives, writing automation for system resilience, and partnering with development teams. Required qualifications include a Bachelor's degree in computer science ...

Lead GCP Data Platform Architect (Real-Time)

Location
Greater London, England, United Kingdom
divisions to build trusted models using Core Lake data. You will lead a team of data engineers, ensuring data quality, observability, and SLA/SLO compliance #J-18808-Ljbffr ...

Senior Data Platform Engineer (Data Lake and Catalog)

Location
Greater London, England, United Kingdom
observable software for our data lake platform. Lead by example in engineering quality through clean, well-tested code and strong coding practices. Promote an SLO-driven culture by contributing to reliability goals, observability, and incident learnings. Partner closely with engineers and other stakeholders to clarify requirements and deliver effectively. Break ...

Global Platform Operations Leader | SRE & CI/CD Excellence

Location
Greater London, England, United Kingdom
leadership layer and drive platform integration for acquisitions, with a focus on reliability and automated delivery across TT's engineering hubs. You will own SLO/SLI governance, AI-first operations, and cross-functional collaboration with CTPO, Security, Product and line-of-business heads. #J-18808-Ljbffr ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
turn root-cause findings into lasting reliability improvements. Collaborate with application development teams and other stakeholders to meet internal Service Level Objectives and customer-facing Service Level Agreements while improving the tools, frameworks … with related technical knowledge. Experience designing and implementing scalable, resilient, and well-tested distributed systems. Experience with Service Level Objectives, Service Level Agreements, monitoring, alerting, capacity planning, incident management, or disaster-recovery testing. ...

Senior Cyber Infrastructure Engineer, Professional Services

Hiring Organisation
Darktrace
Location
London, UK
Employment Type
Full-time
design and planning. What will I be doing: You will be regularly completing deployment projects on or before expected Service Level Objectives (SLOs) and integrating new systems into existing network architecture. On the support side, you'll efficiently manage customer support ...

Senior Cyber Infrastructure Engineer, Professional Services

Location
Greater London, England, United Kingdom
design and planning. **What will I be doing:** You will be regularly completing deployment projects on or before expected Service Level Objectives (SLOs) and integrating new systems into existing network architecture. On the support side, you'll efficiently manage customer support ...

Global SVP, Platform Operations & AI-Driven Reliability

Location
Greater London, England, United Kingdom
consolidate and lead the production platform across SRE, deployment, and platform engineering. You will own the CI/CD pipeline, incident management, and the SLO/SLI framework to raise reliability for critical trading services worldwide. You will build a leadership layer, expand global engineering hubs, and drive AI-driven ...

Technical Support Team Manager, EMEA

Location
Greater London, England, United Kingdom
team to provide consistently high-quality service which drives customer loyalty. Meet or exceed established service level objectives to ensure client satisfaction including staffing, volume, response times and quality of service. Act as a subject matter expert ...

Cyber Infrastructure Engineer, Professional Services

Location
Greater London, England, United Kingdom
design and planning. What will I be doing: You will be regularly completing deployment projects on or before expected Service Level Objectives (SLOs) and integrating new systems into existing network architecture. On the support side, you'll efficiently manage customer support ...

Tech Lead Manager - Lakebase

Location
Greater London, England, United Kingdom
technical work. The team you will lead (one of): Lakebase Platform Reliability (LPR) — Constantly evolve Lakebase reliability: cross‐functional infrastructure, observability, SLI/SLO definition and measurement, and higher operational efficiency without humans-in‐the‐loop, for both internal and customer‐facing systems, by building AI‐powered automation and services. ...

Tech Lead Manager - Lakebase

Hiring Organisation
DataBricks
Location
London, UK
Employment Type
Full-time
technical work. The team you will lead (one of):Lakebase Platform Reliability (LPR) — Constantly evolve Lakebase reliability: cross-functional infrastructure, observability, SLI/SLO definition and measurement, and higher operational efficiency without humans-in-the-loop, for both internal and customer-facing systems, by building AI-powered automation and services. ...

Senior ML Engineer

Location
Greater London, England, United Kingdom
hardware and ensure scalability for exponential user growth. Serve as the technical guardian for the model's quality and Service Level Objectives (SLOs). Provide a hands-on solution architecture for the core ML infrastructure. Select, evaluate, and implement state ...

SVP, Global Head of Platform Operations

Location
Greater London, England, United Kingdom
raising reliability for customers who cannot tolerate downtime. You will own the end-to-end deployment pipeline, production observability and incident management, and the SLO/SLI framework that defines what "good" looks like for every customer-facing service. You will build the leadership layer underneath you, shape … engineering organization. Key Deliverables : Low Mean Time to Resolve (MTTR), minimal production incidents, and near-perfect achievement of Service Level Objectives (SLOs). Fast time-to-market and high throughput velocity through the CI/CD pipeline. Key Responsibilities 1. Delivery ...

Technical Operations Manager - 6 Month FTC – Kings Cross, London

Location
Greater London, England, United Kingdom
addressing issues to minimize downtime. Manage the implementation and maintenance of Service Level Objectives (SLO's) and Service Level Agreements (SLA’s) with UMG Business Units, other IT groups, and vendors as appropriate. … This includes measuring and reporting on these SLA's and SLO’s on a regular and scheduled basis. Drive automation for all repeatable processes by implementing AI and SRE methodologies to reduce toil in daily operations. Major Incident Management Take ownership, lead and coordinate major incident response efforts, acting ...

Senior AI Quality Engineer

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
production issues using our observability stack (Datadog, Sentry, Grafana) - tying test coverage back to real customer impactEnsure teams have Service Level Objectives set up and are achieving themRun targeted exploratory testing on high-risk releases when it's the right callWhat … contract testing experience, with a point of view on test data and environment strategyHave set up and used Service Level Objectives to define a contract of system reliability with stakeholders, and driven measurable improvements against themNative AI-assisted working style. ...