23 of 23 Permanent Observability Jobs in Cambridgeshire

Staff Full Stack Software Engineer

Location
Cambridge, England, United Kingdom
delivers a high-quality user experience. Design, build, and maintain our developer portal including CI/CD pipelines, documentation, automated testing, security upgrades, and observability integrations. Partner closely with platform, software and hardware teams to integrate services, tooling, and policies into the portal in a user-centric and automated manner. ...

(Senior) Software Engineer

Location
Cambridge, England, United Kingdom
Python and software engineering in a Linux environment Strong familiarity with cloud native software development practices – containerisation, CI/CD, API development, DevOps, observability etc. Detailed knowledge of networking and hardware interfacing Excellent programming and problem-solving skills, including the ability to independently debug issues Familiarity with software development practices ...

Senior Engineer, Data Infrastructure

Location
Cambridge, England, United Kingdom
/CD and environment-management patterns that make production systems easier to build, test, deploy and operate. Own cross-cutting engineering decisions for observability, identity and access management, networking, resilience, capacity, data movement and lifecycle management. Partner with Cloud Engineers and Data Engineers to deliver integrated solutions, providing practical mentoring ...

Senior Machine Learning Engineer

Location
Cambridge, England, United Kingdom
practices. Experience with distributed training, large-scale data loading or high-performance computing environments. Experience supporting research-to-production ML workflows. Experience with monitoring, observability, model versioning or MLflow-like tooling. A comprehensive benefits package that includesan annual bonus plan,private medical insurance, life insurance, and acontributory pension scheme ...

Senior Network Engineer

Hiring Organisation
Infoplus Technologies UK Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £500.00 per day
VLANs WAN/LAN SD-WAN Network Security Fundamentals Network Automation Python Ansible REST APIs Terraform – desirable PowerShell – desirable Git/GitHub – desirable Monitoring & Observability Experience with platforms such as: SolarWinds Auvik IP Fabric LogicMonitor Dynatrace Azure Monitor Grafana Preferred Certifications Cisco CCNP/CCIE Juniper JNCIS/JNCIP ...

Applications Engineer

Hiring Organisation
Hays Specialist Recruitment Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£60.00 - £70.00 per day
auditable. Integrate network automation with source control, IPAM, inventory and other platform systems, developing automated configuration compliance and drift detection. Develop telemetry and observability integrations and dashboards to enable easy debugging, solving and tuning opportunities. Work with Compute, Storage and Kubernetes engineers to automate end-to-end network workflows. Required ...

Principal Software Engineer

Location
Peterborough, England, United Kingdom
build cloud platform services, APIs, data layers, and integrate them with large, established ERP products. Bring research & development output up to enterprise standard: reliability, observability, security, multi‐tenancy, and safe failure behavior. Integrate LLMs and agentic patterns responsibly, establishing guardrails, permissions, auditability, and probabilistic design considerations. Raise the technical ...

AI Architect

Hiring Organisation
TALENT INTERNATIONAL UK LTD
Location
Ely, England, United Kingdom
solutions with frontier modelsincluding prompt engineering, fine-tuning, RAG, and agentic/multi-agent system design. LLMOps/MLOps Grounding: Expertise in evaluation frameworks, observability, cost and latency optimization, guardrails, and safe deployment at scale. Cloud and Model Ecosystems: Direct experience with major model providers (Anthropic, OpenAI, Google) and hyperscaler ...

Principal AI Platform Engineer

Location
Cambridge, England, United Kingdom
scalability, secure execution, capacity management and cost-aware operation. Automate provisioning, configuration, upgrades and lifecycle management using infrastructure-as-code and GitOps patterns. Reliability, observability and support: Define and implement service-level indicators, service-level objectives, alerting, dashboards, runbooks and support workflows. Incident response, post-incident review and vendor outage ...

DevOps Engineer (Platform Engineering & Security0

Hiring Organisation
Sanderson
Location
Peterborough, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£500.00 - £550.00 per day
management solutions. Integrate security scanning, compliance validation, and vulnerability management into engineering workflows. Develop automation solutions that improve efficiency and reduce manual intervention. Support observability, monitoring, logging, and platform reliability initiatives. Collaborate with development, infrastructure, database, and security teams to deliver scalable platform solutions. Contribute to platform standardisation, engineering best ...

Senior Platform Engineer

Location
Cambridge, England, United Kingdom
workflows for Kubernetes-based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application-level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real-world, not greenfield ...

Senior Software Engineer

Location
Cambridge, England, United Kingdom
will be impactful and varied, including: Building new product features (full-stack, with frontend focus) Designing non-functional app features (e.g. offline support, improving observability) Contributing to technical roadmap planning Working with product/UX to iterate on the app design and requirements Participating in bug triage/diagnosis ...

Infrastructure Engineer

Location
Cambridge, England, United Kingdom
including cloud bursting. Making it easy for the team to submit, monitor and debug compute‐heavy jobs without needing to understand the machinery underneath. Observability – Monitoring and alerting that tells you something is wrong before a user does. Data infrastructure – Reliable, well‐organised storage and access for large scientific datasets ...

Senior AI Platform Engineer – Production Infra & Security

Location
Cambridge, England, United Kingdom
drive the production AI runtime platforms that power AI services across Arm engineering teams. You will design and operate patterns for isolation, security, observability and cost-aware operation, across Kubernetes, cloud, identity, secrets, networking and telemetry. You will build and maintain the production infrastructure for Arm's AI platform services ...

Senior Security Platform Architect (SCA & Backend)

Location
Cambridge, England, United Kingdom
design and implement backend services, Python APIs, and workflow components to enable tool onboarding, analysis execution, results processing, and delivery, while improving scalability and observability across the platform. #J-18808-Ljbffr ...

Senior Backend Engineer

Location
Cambridge, England, United Kingdom
product lives and dies by — the ingestion pipelines that turn warehouse data into a model, the APIs, durable storage, background work, and the observability that lets a small team operate them with confidence at 3 am. WareBee runs on two engines: Physical AI — a living, spatial model of the warehouse ...

Kubernetes & HPC Infra Engineer

Location
Cambridge, England, United Kingdom
spans on‐prem and cloud, unified under Kubernetes, with emphasis on security and a reliable, self‐service platform. You will own the compute platform, observability, data infrastructure, and security, enabling scalable, automated workflows while maintaining strong safeguards and ease of use for the team. #J-18808-Ljbffr ...

Senior Backend Engineer - Remote or Hybrid, Warehouse Data

Location
Cambridge, England, United Kingdom
contracts that power the product — ingestion pipelines transforming warehouse data into a model, robust APIs, durable storage, and reliable background work. You will ensure observability that lets a small team operate them confidently at 3 am. Two engines power WareBee: Physical AI and Process AI. You will build systems ...

Python Developer - Security - Cambridge/London

Hiring Organisation
Salt Search
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
Collaborate closely with product managers and researchers to translate ideas into working demos and production features. Help troubleshoot, debug, and optimise systems (performance, reliability, observability). Contribute to CI/CD, containerisation (Docker) workflows, and cloud deployment patterns as required. Communicate trade-offs and design decisions clearly; adapt to changing … Desirable (but not required) Experience with Kubernetes or container orchestration. Exposure to ML/AI projects, data pipelines, or research-driven engineering. Experience with observability tooling, CI/CD, and testing best practices. Degree in Computer Science, Engineering, or related technical field (or demonstrable equivalent). *Rates depend on experience ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Senior Product Security Engineer

Location
Cambridge, England, United Kingdom
platform as adoption and analysis volume grow. Integrate and extend SCA capabilities, while building platform foundations that support adjacent security tooling over time. Build observability into services through metrics, dashboards, monitoring, and alerting. Work with security and engineering teams to turn requirements into practical platform capabilities. Contribute to the evolution … tools and the ability to interpret and work with their findings. Experience with large-scale analysis pipelines, data processing, or workflow orchestration. Familiarity with observability tooling, operational metrics, and service health dashboards. Experience with React or similar front-end technologies. Interest in security tooling, vulnerability research, and scalable analysis platforms. ...

Network Engineer - Network Source of Truth, Automation & Network Reliability

Location
Cambridge, England, United Kingdom
network security fundamentals. Automation Practical capability in Python, Ansible and REST APIs. Experience with Terraform, PowerShell and Git/GitHub is desirable. Monitoring & Observability Experience with network monitoring or observability platforms such as netbox, SolarWinds, Auvik, IP Fabric, LogicMonitor, Dynatrace, Azure Monitor or Grafana. #J-18808-Ljbffr ...

IT Infrastructure Solutions Architect

Location
Cambridge, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...