51 to 72 of 72 Observability Jobs in Cambridge

ML Data & Platform Engineer — Hybrid ML Ops & Pipelines

Location
Cambridge, England, United Kingdom
infrastructure to production ML—owning problems end-to-end to accelerate model delivery. You’ll collaborate with the ML team to improve data quality, observability, and MLOps practices, while scaling infrastructure for faster iteration and reliability. #J-18808-Ljbffr ...

Cloud HPC Infrastructure Engineer (Kubernetes)

Location
Cambridge, England, United Kingdom
with burst capability for peak demand. You will simplify access and ensure reliability for the data science team. You will own the compute platform, observability, data infrastructure, and security, embracing automation, robust tooling, and zero‐trust networking to deliver a #J-18808-Ljbffr ...

Kubernetes & HPC Infra Engineer

Location
Cambridge, England, United Kingdom
spans on‐prem and cloud, unified under Kubernetes, with emphasis on security and a reliable, self‐service platform. You will own the compute platform, observability, data infrastructure, and security, enabling scalable, automated workflows while maintaining strong safeguards and ease of use for the team. #J-18808-Ljbffr ...

Senior Backend Engineer - Remote or Hybrid, Warehouse Data

Location
Cambridge, England, United Kingdom
contracts that power the product — ingestion pipelines transforming warehouse data into a model, robust APIs, durable storage, and reliable background work. You will ensure observability that lets a small team operate them confidently at 3 am. Two engines power WareBee: Physical AI and Process AI. You will build systems ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Senior Software Engineer, ML Infrastructure

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Live Operations Lead - FAST Channel Onboarding

Location
Cambridge, England, United Kingdom
ensure smooth launches and responsive changes. The role blends live operations with platform reliability, requiring to work with infrastructure and cloud teams on observability, automation, and encoding health as the FAST landscape grows. #J-18808-Ljbffr ...

Live Operations Associate Roku, Inc.

Location
Cambridge, England, United Kingdom
role sits at the intersection of live operations and platform reliability, so you will work alongside infrastructure, engineering teams, and service providers to support observability, automation, and encoding pipeline health as the FAST ecosystem scales. As a trusted point of contact for service partners, you will manage expectations, communicate clearly … resolved efficiently Managing ongoing FAST channel configurations, including schedule updates, metadata changes, feed ingestion, and EPG alignment Monitoring channel and pipeline health using observability tools; validating ingestion, triaging transcode and CDN delivery issues, and escalating incidents with clear context to engineering or peer technical teams Driving automation of repetitive operational ...

Senior Software Engineer, ML Infrastructure Roku, Inc.

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Senior SRE: Cloud Platform Reliability & Observability

Location
Cambridge, England, United Kingdom
reliability, performance, and scalability of its cloud platforms from its Cambridge, UK operations. You will work across DevOps, engineering, and platform teams to implement observability, reliability frameworks, and automated deployment practices. The role emphasizes incident response, capacity planning, and continuous improvement, with a focus on IaC, SOO principles, and cross ...

Python Developer – Security – Cambridge/London

Location
Cambridge, England, United Kingdom
Collaborate closely with product managers and researchers to translate ideas into working demos and production features. Help troubleshoot, debug, and optimise systems (performance, reliability, observability). Contribute to CI/CD, containerisation (Docker) workflows, and cloud deployment patterns as required. Communicate trade-offs and design decisions clearly; adapt to changing … demonstrable ability considered). Experience with Kubernetes or container orchestration. Exposure to ML/AI projects, data pipelines, or research-driven engineering. Experience with observability tooling, CI/CD, and testing best practices. Degree in Computer Science, Engineering, or related technical field (or demonstrable equivalent). *Rates depend on experience ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
matter at once. You enjoy building systems that are clean enough to reason about, but pragmatic enough to ship. You care about tests, observability, and operational safety, but you do not hide behind process. Problems you’ll tackle Building low-latency optimisation APIs that sit on the critical path … Python services across serving, workers, training workflows, and internal tooling Working with Postgres, Redis, queues/streams, migrations, and event-driven workflows Making reliability, observability, and deployment safety part of the product from the beginning Core responsibilities Write production-grade Rust and Python services Design clean domain boundaries around requests ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

ML Data & Platform Engineer

Location
Cambridge, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

Senior Product Security Engineer

Location
Cambridge, England, United Kingdom
platform as adoption and analysis volume grow. Integrate and extend SCA capabilities, while building platform foundations that support adjacent security tooling over time. Build observability into services through metrics, dashboards, monitoring, and alerting. Work with security and engineering teams to turn requirements into practical platform capabilities. Contribute to the evolution … tools and the ability to interpret and work with their findings. Experience with large-scale analysis pipelines, data processing, or workflow orchestration. Familiarity with observability tooling, operational metrics, and service health dashboards. Experience with React or similar front-end technologies. Interest in security tooling, vulnerability research, and scalable analysis platforms. ...

Staff/Lead Python Engineer (FastAPI, Orchestration)

Location
Cambridge, England, United Kingdom
with other apps in aservicearchitecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests Strong experience with observability and tracing. Knowledge of best practices for performance optimisation, memory management. Experience mentoring others, especially in good software development practices, patterns, and fundamentals. Drive ...

Senior SRE - Platform Reliability & Observability Lead

Location
Cambridge, England, United Kingdom
seeking a Senior Site Reliability Engineer to lead reliability, performance and continuous improvement of the Bango Platform. You will own end-to-end reliability, observability and incident response across infrastructure and delivery pipelines, serving as a technical centre of gravity for the SRE function. You will shape the bench across ...

Network Engineer

Hiring Organisation
Third Nexus Group Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Contract
Contract Rate
£375 - £400/annum
network security fundamentals. Automation Practical capability in Python, Ansible and REST APIs. Experience with Terraform, PowerShell and Git/GitHub is desirable. Monitoring & Observability Experience with network monitoring or observability platforms such as netbox, SolarWinds, Auvik, IP Fabric, LogicMonitor, Dynatrace, Azure Monitor or Grafana ...

Software Engineer, Agentic AI Roku, Inc.

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development - from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Senior Director, Data and Information Marketplace

Location
Cambridge, England, United Kingdom
ensuring intuitive experiences for both people and agents across discovery, access, sharing, understanding and use. Advise the development of capabilities for data access, lineage, observability and quality so that data assets are transparent, trusted and usable at scale. Shape enterprise approaches to data, information and knowledge lifecycle management, embedding governance … seamless experiences that are widely adopted by users and machines across multiple enterprise business units. Deep expertise in relevant capability areas, including data quality, observability, lineage, access management, lifecycle management and information governance. Measurable evidence of optimising the data P&L across covering cost/FinOps, value realisation and sustainability. ...

IT Infrastructure Solutions Architect

Location
Cambridge, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...

Agentic AI Engineer: Build Production-Grade TV Agents

Location
Cambridge, England, United Kingdom
Roku TV is seeking a hands-on Agentic AI Engineer to design, build, and maintain intelligent agents and copilots that drive automation and unlock new product capabilities. You will own the full lifecycle from prototyping ...