201 to 225 of 259 Observability Jobs in the East of England

Remote Data Engineering Manager

Location
Ipswich, Suffolk, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Cambridge, Cambridgeshire, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Norwich, Norfolk, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Chigwell, Essex, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...

Live Operations Lead - FAST Channel Onboarding

Location
Cambridge, England, United Kingdom
ensure smooth launches and responsive changes. The role blends live operations with platform reliability, requiring to work with infrastructure and cloud teams on observability, automation, and encoding health as the FAST landscape grows. #J-18808-Ljbffr ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, United Kingdom
Employment Type
Permanent
Salary
GBP 50,000 Annual
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, click apply for full ...

Senior Network Engineer- IP

Location
Ipswich, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Live Operations Associate Roku, Inc.

Location
Cambridge, England, United Kingdom
role sits at the intersection of live operations and platform reliability, so you will work alongside infrastructure, engineering teams, and service providers to support observability, automation, and encoding pipeline health as the FAST ecosystem scales. As a trusted point of contact for service partners, you will manage expectations, communicate clearly … resolved efficiently Managing ongoing FAST channel configurations, including schedule updates, metadata changes, feed ingestion, and EPG alignment Monitoring channel and pipeline health using observability tools; validating ingestion, triaging transcode and CDN delivery issues, and escalating incidents with clear context to engineering or peer technical teams Driving automation of repetitive operational ...

Remote Senior Director, Engineering- X-Ops Platform

Location
Thetford, Norfolk, United Kingdom
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Data Engineer

Hiring Organisation
Codat
Location
Hertfordshire, United Kingdom
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Hertfordshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Location
Stone, Hertfordshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Senior Software Engineer, ML Infrastructure Roku, Inc.

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Senior SRE: Cloud Platform Reliability & Observability

Location
Cambridge, England, United Kingdom
reliability, performance, and scalability of its cloud platforms from its Cambridge, UK operations. You will work across DevOps, engineering, and platform teams to implement observability, reliability frameworks, and automated deployment practices. The role emphasizes incident response, capacity planning, and continuous improvement, with a focus on IaC, SOO principles, and cross ...

Remote Staff Product Manager, OpenTelemetry | UK | Remote

Hiring Organisation
Grafana Labs
Location
Central bedfordshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … confirm that telemetry is healthy, correctly attributed, and usable in Grafana Cloud The transition from receiving telemetry to activating the appropriate Grafana Cloud observability solution Success does not stop when Grafana Cloud receives the first data point. Telemetry must arrive with the quality, context, and structure required to support meaningful ...

Remote Staff Product Manager, OpenTelemetry UK Remote

Location
Ware, Hertfordshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … confirm that telemetry is healthy, correctly attributed, and usable in Grafana Cloud The transition from receiving telemetry to activating the appropriate Grafana Cloud observability solution Success does not stop when Grafana Cloud receives the first data point. Telemetry must arrive with the quality, context, and structure required to support meaningful ...

Remote Senior Software Engineer Infra Agent Systems UK

Location
Cambridge, Cambridgeshire, United Kingdom
infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing … deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience ...

Remote Developer Support Engineer (London)

Location
Bedford, Bedfordshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Python Developer – Security – Cambridge/London

Location
Cambridge, England, United Kingdom
Collaborate closely with product managers and researchers to translate ideas into working demos and production features. Help troubleshoot, debug, and optimise systems (performance, reliability, observability). Contribute to CI/CD, containerisation (Docker) workflows, and cloud deployment patterns as required. Communicate trade-offs and design decisions clearly; adapt to changing … demonstrable ability considered). Experience with Kubernetes or container orchestration. Exposure to ML/AI projects, data pipelines, or research-driven engineering. Experience with observability tooling, CI/CD, and testing best practices. Degree in Computer Science, Engineering, or related technical field (or demonstrable equivalent). *Rates depend on experience ...

Remote Senior Microsoft Power Platform and AI Developer

Hiring Organisation
Methods Business And Digital Technology
Location
Essex, United Kingdom
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offs and safe failure paths. You will ensure that evaluation, observability, identity, permissions and data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineering practices … C# or Python, with experience extending low-code services appropriately. • Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. • Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalate architecture or security risks appropriately. ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
matter at once. You enjoy building systems that are clean enough to reason about, but pragmatic enough to ship. You care about tests, observability, and operational safety, but you do not hide behind process. Problems you’ll tackle Building low-latency optimisation APIs that sit on the critical path … Python services across serving, workers, training workflows, and internal tooling Working with Postgres, Redis, queues/streams, migrations, and event-driven workflows Making reliability, observability, and deployment safety part of the product from the beginning Core responsibilities Write production-grade Rust and Python services Design clean domain boundaries around requests ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Remote Sr. Software Engineer, Fullstack (UK)

Location
Epping, Essex, United Kingdom
post-incident reviews in a "you build it, you run it" environment. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Partner with Product Management and Design to translate business requirements into scalable technical solutions. Minimum Qualifications Bachelor's degree in Computer Science … HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Kubernetes, Docker, and Helm. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Experience in enterprise SaaS organisations, particularly HR Tech or regulated ...

ML Data & Platform Engineer

Location
Cambridge, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...