101 to 125 of 167 Observability Jobs in the East of England

Site Reliability Engineer (SRE)

Location
Cambridge, England, United Kingdom
availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications. Join Altium as a Senior Site Reliability Engineer … ensure the reliability and performance of the Altium Cloud Platforms. Key Responsibilities: Understanding how an Altium Cloud Platform works Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection. Develop and implement reliability frameworks and patterns that standardize and elevate ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
focused on their Azure platform, another focused on their AWS platform. Both roles require hands on experience with IaC, Containerisation and Monitoring/Observability tools. Experience required for the Senior Site Reliability Engineer: Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within … experience with both Azure, and on-premise virtual machines. Experience with Infrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways ...

Senior Platform Engineer

Location
Welwyn Garden City, England, United Kingdom
GitOps workflows to enable safe, fast, and repeatable delivery. Championing DevSecOps principles, embedding security and compliance into the software delivery lifecycle. Establishing and improving observability, monitoring, and incident response practices, including vulnerability management and remediation. Mentoring engineers and contributing to a strong engineering culture through knowledge sharing, documentation, and technical … would be great if you have the following Experience with Helm, Kustomize, and Kubernetes ecosystem tooling. Familiarity with Azure and Azure DevOps. Experience with observability platforms and Kubernetes policy enforcement tools. Proficiency in scripting or programming (e.g. Bash, Python, PowerShell, C#). Experience designing multi‐region or highly available systems. ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
grabjobs
Location
Ipswich, Suffolk, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Senior SRE: Cloud Platform Reliability & Automation

Location
Cambridge, England, United Kingdom
Engineer to ensure the reliability, availability, and performance of our large-scale cloud platforms and SaaS products. Your work will automate operational tasks, improve observability, and contribute to incident management while collaborating with development teams to build more reliable and scalable applications across regions. #J-18808-Ljbffr ...

Production Reliability Lead

Location
Norwich, England, United Kingdom
team to move from firefighting to proactive reliability engineering. This hands-on role requires leading SRE/DevOps practices, defining SLIs/SLOs, improving observability, and ensuring clear ownership with blameless accountability. #J-18808-Ljbffr ...

Operations Team Lead — Production Reliability & Scale

Location
Cambridge, England, United Kingdom
seeking an Operations Team Lead to own production and scale systems. You will lead operational excellence across live customer-facing platforms, ensuring reliability, observability, and proactive improvements. This hands-on role involves shaping processes, guiding incidents, building the team, and moving from firefighting to sustainable reliability engineering. The ideal candidate ...

Senior IP Network Engineer & SRE Lead

Location
Ipswich, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...

Mid Software Engineer – End-to-End Platform Ownership (Hybrid)

Location
Cambridge, England, United Kingdom
from day one. The role offers hybrid working (in-office every two weeks) with Cambridge office, exposure to C#/.NET, React, Azure, and observability tooling, and opportunities to grow within a supportive #J-18808-Ljbffr ...

Senior Backend Engineer

Location
Cambridge, England, United Kingdom
product lives and dies by — the ingestion pipelines that turn warehouse data into a model, the APIs, durable storage, background work, and the observability that lets a small team operate them with confidence at 3 am. WareBee runs on two engines: Physical AI — a living, spatial model of the warehouse ...

Principal Software Architect

Location
Cambridge, England, United Kingdom
/software ecosystem. Assess the architectural impact of new technologies. Be aware of the usability, performance, reliability, maintainability, testability, security and observability constraints on the software architecture. Prototyping and validating architectural concepts through proof-of-concept implementations. Contribute to future and/or related product definitions with a forward-looking ...

Staff Verification Engineer - Media IP

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
Identify technical and schedule risks early, establish mitigation plans and drive them to resolution across teams. Review proposed architecture and design changes for correctness, observability, testability, performance and verification complexity. Lead the resolution of complex unit-level, cross-unit, integration or system-level failures. Align verification activities and dependencies across ...

Infrastructure Monitoring Engineer

Hiring Organisation
BT Group
Location
Ipswich, Suffolk, UK
Employment Type
Full-time
like to see on your CVMandatoryFamiliarity with using Linux operating systemsExperience with one or more of the following areas Administering or deploying monitoring and observability tools such as CheckMK, Prometheus or ZabbixAdministering or deploying SIEM tools such as Elastic Security, Splunk or Trellix ESMWorking with dashboarding and visualisation tools such ...

ML Data & Platform Engineer — Hybrid ML Ops & Pipelines

Location
Cambridge, England, United Kingdom
infrastructure to production ML—owning problems end-to-end to accelerate model delivery. You’ll collaborate with the ML team to improve data quality, observability, and MLOps practices, while scaling infrastructure for faster iteration and reliability. #J-18808-Ljbffr ...

Staff Engineer - Embedded Accountancy Platform Lead

Location
Norwich, England, United Kingdom
collaborate with external partners to deliver scalable financial tooling. As a platform-focused leader, you will ensure robust integration patterns and high standards for observability, performance, and security across the product ecosystem. #J-18808-Ljbffr ...

Cloud HPC Infrastructure Engineer (Kubernetes)

Location
Cambridge, England, United Kingdom
with burst capability for peak demand. You will simplify access and ensure reliability for the data science team. You will own the compute platform, observability, data infrastructure, and security, embracing automation, robust tooling, and zero‐trust networking to deliver a #J-18808-Ljbffr ...

Kubernetes & HPC Infra Engineer

Location
Cambridge, England, United Kingdom
spans on‐prem and cloud, unified under Kubernetes, with emphasis on security and a reliable, self‐service platform. You will own the compute platform, observability, data infrastructure, and security, enabling scalable, automated workflows while maintaining strong safeguards and ease of use for the team. #J-18808-Ljbffr ...

Senior Backend Engineer - Remote or Hybrid, Warehouse Data

Location
Cambridge, England, United Kingdom
contracts that power the product — ingestion pipelines transforming warehouse data into a model, robust APIs, durable storage, and reliable background work. You will ensure observability that lets a small team operate them confidently at 3 am. Two engines power WareBee: Physical AI and Process AI. You will build systems ...

Data Platform Solution Architect

Location
Basildon, England, United Kingdom
Design Documents (ADDs)*** Deep understanding of **cloud-native design patterns*** Experience in **performance tuning** across:* Snowflake* Airflow* Iceberg* Focus on **platform reliability, scalability, and observability*** Experience designing and operating **data platforms** in production environments #J-18808-Ljbffr ...

Remote Senior AI Software Engineer

Hiring Organisation
grabjobs
Location
Ipswich, Suffolk, UK
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Senior SRE (AWS)

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
Strong hands-on experience with both AWS, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Southend-on-Sea, Essex, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Site Reliability Engineer NEW Posted today Hemel Hempstead Haven Haven

Location
Hemel Hempstead, England, United Kingdom
Tech Leads to design, implement and support the systems that guests, owners and colleagues rely on every day. From CI/CD pipelines and observability through to database reliability, incident management and disaster recovery, this role touches every layer of our stack. This is also a great time to join. … developers and engineers to troubleshoot build and deployment issues and unblock delivery Contribute to and maintain internally developed engineering tools Own monitoring, tracing and observability so we are first to know when something is (or is about to be) an issue, and can diagnose it quickly Drive database reliability across ...

Senior Software Engineer, ML Infrastructure

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
Hingham, Norfolk, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...