51 to 75 of 92 Observability Jobs in the East of England

AWS Solutions Architect - Microservices & Event-Driven

Location
Cambridge, England, United Kingdom
translate complex requirements into production-grade solutions. You will define end-to-end architectures, lead domain-driven design, and set patterns for resilience, observability, and security. #J-18808-Ljbffr ...

Remote SRE: Cloud Reliability & AI-Driven Ops

Location
Hemel Hempstead, England, United Kingdom
Haven is seeking a hands-on Site Reliability Engineer to join our Product Technology team. This remote-first role involves shaping CI/CD, observability, and incident response while collaborating with engineers and tech leads to ensure reliable, scalable platforms for guests and colleagues. You’ll tackle infrastructure design, tooling ...

Backend SDE II — Real-Time Data & Event-Driven Systems

Location
Welwyn Garden City, England, United Kingdom
delivering personalised experiences at scale while collaborating with more senior engineers. You will gain hands-on experience with Kafka, Kubernetes, CI/CD, and observability tools, contributing to distributed systems and production readiness in a fast-paced environment. #J-18808-Ljbffr ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Operations Team Lead (Production & Reliability)

Location
Norwich, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Operations Team Lead (Production & Reliability)

Location
Watford, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Senior Platform Engineer

Location
Welwyn Garden City, England, United Kingdom
GitOps workflows to enable safe, fast, and repeatable delivery. Championing DevSecOps principles, embedding security and compliance into the software delivery lifecycle. Establishing and improving observability, monitoring, and incident response practices, including vulnerability management and remediation. Mentoring engineers and contributing to a strong engineering culture through knowledge sharing, documentation, and technical … would be great if you have the following Experience with Helm, Kustomize, and Kubernetes ecosystem tooling. Familiarity with Azure and Azure DevOps. Experience with observability platforms and Kubernetes policy enforcement tools. Proficiency in scripting or programming (e.g. Bash, Python, PowerShell, C#). Experience designing multi‐region or highly available systems. ...

Senior Site Reliability Engineer – Global Platform Ops

Location
Ipswich, England, United Kingdom
improvement across international networks and digital platforms. You will collaborate with engineering, product, suppliers and ops teams to raise platform resilience, implement automation and observability, and deliver outstanding service outcomes on a global stage. #J-18808-Ljbffr ...

Production Reliability Lead

Location
Norwich, England, United Kingdom
team to move from firefighting to proactive reliability engineering. This hands-on role requires leading SRE/DevOps practices, defining SLIs/SLOs, improving observability, and ensuring clear ownership with blameless accountability. #J-18808-Ljbffr ...

Operations Team Lead — Production Reliability & Scale

Location
Cambridge, England, United Kingdom
seeking an Operations Team Lead to own production and scale systems. You will lead operational excellence across live customer-facing platforms, ensuring reliability, observability, and proactive improvements. This hands-on role involves shaping processes, guiding incidents, building the team, and moving from firefighting to sustainable reliability engineering. The ideal candidate ...

Mid Software Engineer - Own End-to-End Features

Location
Cambridge, England, United Kingdom
teams and the Shared Applications & Services group. The role emphasizes learning, code quality, and architecture, with exposure to C#/.NET, React, Azure and observability tooling. Expect a hybrid working pattern with in-office presence about every two weeks and a salary up #J-18808-Ljbffr ...

Mid Software Engineer – End-to-End Platform Ownership (Hybrid)

Location
Cambridge, England, United Kingdom
from day one. The role offers hybrid working (in-office every two weeks) with Cambridge office, exposure to C#/.NET, React, Azure, and observability tooling, and opportunities to grow within a supportive #J-18808-Ljbffr ...

Senior Backend Engineer

Location
Cambridge, England, United Kingdom
product lives and dies by — the ingestion pipelines that turn warehouse data into a model, the APIs, durable storage, background work, and the observability that lets a small team operate them with confidence at 3 am. WareBee runs on two engines: Physical AI — a living, spatial model of the warehouse ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
solve complex reliability challenges, and improve system resilience at scale.Unlike a generalist SRE, this role focuses on a **core domain of expertise**—such as **observability, performance engineering, data infrastructure reliability, security-focused SRE, or network reliability**—while influencing reliability standards across the wider engineering organisation.## ## **Key Responsibilities**### **Domain … Bring**### **Essential*** Proven experience in **Site Reliability Engineering, DevOps, or infrastructure engineering*** Deep expertise in at least one of the following areas: + Observability & monitoring (metrics, logging, distributed tracing) + Performance engineering & capacity planning + Data infrastructure reliability (databases, streaming, pipelines) + Security-focused SRE (hardening, compliance automation, secrets ...

Senior / Lead Site Reliability Engineer

Location
Watford, England, United Kingdom
Peak jackpot proactive monitoring Own end-to-end incident lifecycle: Detection triage resolution post-incident review Ensure blameless post-mortems with clear remediation ownership Observability & service insight Define and evolve observability strategy using: Splunk (log analytics) CloudWatch (AWS telemetry) Grafana (metrics visualisation) Quantum Metric (user behaviour insight) Standardise: Alerting quality … safety (CI/CD, progressive delivery patterns) Key contributor to transition strategy for ECS EKS (Kubernetes adoption) Reduce operational toil through tooling, self-healing, observability and platform improvements Empower Level-1 operational teams with safe, controlled access to the tools they need to operate autonomously Capacity & performance engineering Own capacity ...

ML Data & Platform Engineer — Hybrid ML Ops & Pipelines

Location
Cambridge, England, United Kingdom
infrastructure to production ML—owning problems end-to-end to accelerate model delivery. You’ll collaborate with the ML team to improve data quality, observability, and MLOps practices, while scaling infrastructure for faster iteration and reliability. #J-18808-Ljbffr ...

Staff Engineer - Embedded Accountancy Platform Lead

Location
Norwich, England, United Kingdom
collaborate with external partners to deliver scalable financial tooling. As a platform-focused leader, you will ensure robust integration patterns and high standards for observability, performance, and security across the product ecosystem. #J-18808-Ljbffr ...

Cloud HPC Infrastructure Engineer (Kubernetes)

Location
Cambridge, England, United Kingdom
with burst capability for peak demand. You will simplify access and ensure reliability for the data science team. You will own the compute platform, observability, data infrastructure, and security, embracing automation, robust tooling, and zero‐trust networking to deliver a #J-18808-Ljbffr ...

Kubernetes & HPC Infra Engineer

Location
Cambridge, England, United Kingdom
spans on‐prem and cloud, unified under Kubernetes, with emphasis on security and a reliable, self‐service platform. You will own the compute platform, observability, data infrastructure, and security, enabling scalable, automated workflows while maintaining strong safeguards and ease of use for the team. #J-18808-Ljbffr ...

Senior Backend Engineer - Remote or Hybrid, Warehouse Data

Location
Cambridge, England, United Kingdom
contracts that power the product — ingestion pipelines transforming warehouse data into a model, robust APIs, durable storage, and reliable background work. You will ensure observability that lets a small team operate them confidently at 3 am. Two engines power WareBee: Physical AI and Process AI. You will build systems ...

Data Platform Solution Architect

Location
Basildon, England, United Kingdom
Design Documents (ADDs)*** Deep understanding of **cloud-native design patterns*** Experience in **performance tuning** across:* Snowflake* Airflow* Iceberg* Focus on **platform reliability, scalability, and observability*** Experience designing and operating **data platforms** in production environments #J-18808-Ljbffr ...

Site Reliability Engineer

Location
Hemel Hempstead, England, United Kingdom
Tech Leads to design, implement and support the systems that guests, owners and colleagues rely on every day. From CI/CD pipelines and observability through to database reliability, incident management and disaster recovery, this role touches every layer of our stack.This is also a great time to join. … developers and engineers to troubleshoot build and deployment issues and unblock delivery* Contribute to and maintain internally developed engineering tools* Own monitoring, tracing and observability so we are first to know when something is (or is about to be) an issue, and can diagnose it quickly* Drive database reliability across ...

Senior Software Engineer, ML Infrastructure

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Data Architect

Hiring Organisation
PA Consulting
Location
Pimlico, Hertfordshire, UK
Employment Type
Full-time
Company DescriptionWe believe in the power of ingenuity to build a positive human future. We challenge where it matters and own the outcome. As strategies, technologies, and innovation collide, we create opportunity from complexity. Our ...

Site Reliability Engineering Professional

Location
Ipswich, England, United Kingdom
heart of BT International's next-generation platforms, collaborating with engineering, product, supplier, and operational teams to improve service reliability through automation, observability, and SRE best practices. What you'll be doing Support the 24x7 operation of BT International's core network and platform services, ensuring high availability and performance. … reliability, resilience, and operational readiness. Build and maintain high-quality operational documentation, including runbooks, service maps, playbooks, and handover processes. Develop and enhance monitoring, observability, and operational tooling capabilities. Support the implementation of automation and CI/CD practices to improve operational efficiency. Coach and support your colleagues and customer ...