126 to 150 of 167 Observability Jobs in the East of England

Live Operations Lead - FAST Channel Onboarding

Location
Cambridge, England, United Kingdom
ensure smooth launches and responsive changes. The role blends live operations with platform reliability, requiring to work with infrastructure and cloud teams on observability, automation, and encoding health as the FAST landscape grows. #J-18808-Ljbffr ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, United Kingdom
Employment Type
Permanent
Salary
GBP 50,000 Annual
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, click apply for full ...

Senior Network Engineer- IP

Location
Ipswich, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Live Operations Associate Roku, Inc.

Location
Cambridge, England, United Kingdom
role sits at the intersection of live operations and platform reliability, so you will work alongside infrastructure, engineering teams, and service providers to support observability, automation, and encoding pipeline health as the FAST ecosystem scales. As a trusted point of contact for service partners, you will manage expectations, communicate clearly … resolved efficiently Managing ongoing FAST channel configurations, including schedule updates, metadata changes, feed ingestion, and EPG alignment Monitoring channel and pipeline health using observability tools; validating ingestion, triaging transcode and CDN delivery issues, and escalating incidents with clear context to engineering or peer technical teams Driving automation of repetitive operational ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
grabjobs
Location
Barton-le-Clay, Bedfordshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Data Engineer

Hiring Organisation
Codat
Location
Luton, Bedfordshire, UK
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Remote Senior Data Engineer

Hiring Organisation
grabjobs
Location
St Ives, Cambridgeshire, UK
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Remote Senior Data Engineer

Hiring Organisation
Codat
Location
Great Yarmouth, Norfolk, UK
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Clophill, Bedfordshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Senior Software Engineer, ML Infrastructure Roku, Inc.

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Cambridge, Cambridgeshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Hitchin, Hertfordshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Senior SRE: Cloud Platform Reliability & Observability

Location
Cambridge, England, United Kingdom
reliability, performance, and scalability of its cloud platforms from its Cambridge, UK operations. You will work across DevOps, engineering, and platform teams to implement observability, reliability frameworks, and automated deployment practices. The role emphasizes incident response, capacity planning, and continuous improvement, with a focus on IaC, SOO principles, and cross ...

Remote Senior Software Engineer Infra Agent Systems UK

Hiring Organisation
grabjobs
Location
Huntingdon, Cambridgeshire, UK
infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing … deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience ...

Remote Developer Support Engineer (London)

Hiring Organisation
grabjobs
Location
Jaywick, Essex, UK
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Python Developer – Security – Cambridge/London

Location
Cambridge, England, United Kingdom
Collaborate closely with product managers and researchers to translate ideas into working demos and production features. Help troubleshoot, debug, and optimise systems (performance, reliability, observability). Contribute to CI/CD, containerisation (Docker) workflows, and cloud deployment patterns as required. Communicate trade-offs and design decisions clearly; adapt to changing … demonstrable ability considered). Experience with Kubernetes or container orchestration. Exposure to ML/AI projects, data pipelines, or research-driven engineering. Experience with observability tooling, CI/CD, and testing best practices. Degree in Computer Science, Engineering, or related technical field (or demonstrable equivalent). *Rates depend on experience ...

Remote Senior Microsoft Power Platform and AI Developer

Hiring Organisation
grabjobs
Location
Shefford, Bedfordshire, UK
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offs and safe failure paths. You will ensure that evaluation, observability, identity, permissions and data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineering practices … C# or Python, with experience extending low-code services appropriately. • Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. • Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalate architecture or security risks appropriately. ...

Remote Senior Microsoft Power Platform and AI Developer

Hiring Organisation
grabjobs
Location
Basildon, Essex, UK
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offs and safe failure paths. You will ensure that evaluation, observability, identity, permissions and data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineering practices … C# or Python, with experience extending low-code services appropriately. • Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. • Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalate architecture or security risks appropriately. ...

Remote Senior Microsoft Power Platform and AI Developer

Hiring Organisation
grabjobs
Location
Potters Bar, Hertfordshire, UK
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offs and safe failure paths. You will ensure that evaluation, observability, identity, permissions and data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineering practices … C# or Python, with experience extending low-code services appropriately. • Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. • Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalate architecture or security risks appropriately. ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
matter at once. You enjoy building systems that are clean enough to reason about, but pragmatic enough to ship. You care about tests, observability, and operational safety, but you do not hide behind process. Problems you’ll tackle Building low-latency optimisation APIs that sit on the critical path … Python services across serving, workers, training workflows, and internal tooling Working with Postgres, Redis, queues/streams, migrations, and event-driven workflows Making reliability, observability, and deployment safety part of the product from the beginning Core responsibilities Write production-grade Rust and Python services Design clean domain boundaries around requests ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
grabjobs
Location
Abbots Langley, Hertfordshire, UK
post-incident reviews in a "you build it, you run it" environment. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Partner with Product Management and Design to translate business requirements into scalable technical solutions. Minimum Qualifications Bachelor's degree in Computer Science … HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Kubernetes, Docker, and Helm. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Experience in enterprise SaaS organisations, particularly HR Tech or regulated ...

ML Data & Platform Engineer

Location
Cambridge, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

Platform Technical Lead (OpenShift)

Location
Welwyn Garden City, England, United Kingdom
platform across its full stack: OpenShift (HCP/Virtualization), GitOps delivery (Argo CD), multi‐cluster management and policy (ACM, Gatekeeper/OPA), observability, networking and storage integration, and infrastructure‐as‐code. Set and enforce the paved‐road patterns the team builds to. Evaluate new technologies against product outcomes, not novelty … first platform design: authoring and versioning APIs consumed by other product teams, managing breaking changes and deprecation, Kubernetes API extension patterns (CRDs, operators). Observability engineering: Prometheus, Alertmanager, Grafana or equivalent; defining SLIs/SLOs and designing alerting that reflects service health rather than component noise. Designing and operating fault ...

Senior Product Security Engineer

Location
Cambridge, England, United Kingdom
platform as adoption and analysis volume grow. Integrate and extend SCA capabilities, while building platform foundations that support adjacent security tooling over time. Build observability into services through metrics, dashboards, monitoring, and alerting. Work with security and engineering teams to turn requirements into practical platform capabilities. Contribute to the evolution … tools and the ability to interpret and work with their findings. Experience with large-scale analysis pipelines, data processing, or workflow orchestration. Familiarity with observability tooling, operational metrics, and service health dashboards. Experience with React or similar front-end technologies. Interest in security tooling, vulnerability research, and scalable analysis platforms. ...