26 to 50 of 73 Observability Jobs in Cambridge

Service Design Specialist

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
support the service, including incident, major incident, request, problem, change, release, knowledge and continual improvement activities. Embed availability, capacity, resilience, continuity, performance, security, observability, user experience and supportability requirements into service designs. Help define service levels and service health measures, including SLAs, SLOs, SLIs, operational metrics, monitoring, alerting, dashboards ...

Senior DevOps Engineer

Location
Cambridge, England, United Kingdom
workflows for Kubernetes-based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application-level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real-world, not greenfield ...

Senior Platform Engineer

Location
Cambridge, England, United Kingdom
workflows for Kubernetes-based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application-level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real-world, not greenfield ...

Remote Software Engineer (Machine Learning)

Hiring Organisation
grabjobs
Location
Cambridge, Cambridgeshire, UK
product, from reporting and APIs through to the field apps used to verify leaks on site Taking ownership of production systems, including testing, observability, performance, reliability and troubleshooting Exploring and solving problems where the solution isn't always known upfront — combining engineering with experimentation and real-world feedback Application Process ...

Remote Senior Software Engineer - Serverless

Hiring Organisation
Runware
Location
Cambridge, Cambridgeshire, UK
custom LLMs Architect scalable, reliable systems for inference orchestration and workload optimisation Partner with platform and infrastructure teams to drive performance, reliability, and observability improvements Shape technical direction through architectural discussions and long-term planning Mentor engineers through code and design reviews, fostering technical excellence and growth Contribute to documentation ...

Infrastructure Engineer (Cambridge)

Location
Cambridge, England, United Kingdom
including cloud bursting. Making it easy for the team to submit, monitor and debug compute‐heavy jobs without needing to understand the machinery underneath. Observability – Monitoring and alerting that tells you something is wrong before a user does. Data infrastructure – Reliable, well‐organised storage and access for large scientific datasets ...

Principal DevOps Engineer — AWS, Kubernetes & SaaS (Remote)

Location
Cambridge, England, United Kingdom
engineering culture. You’ll collaborate with developers, data engineers and product teams to ensure safe, rapid releases, while shaping the technical roadmap and improving observability and incident response. #J-18808-Ljbffr ...

Senior Platform Engineer: Kubernetes & CI/CD Leader

Location
Cambridge, England, United Kingdom
deploy and scale workloads. You’ll own day-to-day delivery across our on-premise estate, driving platform stability, self-service tooling, and robust observability to help engineering ship value safely and rapidly to production. #J-18808-Ljbffr ...

Business Systems Software Engineer/Architect Product Design & Engineering · ·

Location
Cambridge, England, United Kingdom
Bachelor's degree in Computer Science, Software Engineering, or equivalent practical industry experience. Experience designing 'defensive' system architectures—building for error handling, auditability, and observability, specifically for scenarios where manual operational intervention is required. A collaborative mindset that values the speed of low-code platforms while maintaining the rigor ...

Software Engineer Full Stack

Hiring Organisation
Foundation Partners
Location
Cambridge, Cambridgeshire, East Anglia, United Kingdom
Employment Type
Permanent, Work From Home
will be impactful and varied, including: Building new product features (full-stack, with frontend focus) Designing non-functional app features (e.g. offline support, improving observability) Contributing to technical roadmap planning Working with product/UX to iterate on the app design and requirements Participating in bug triage/diagnosis ...

EDA Optimisation Engineer

Location
Cambridge, England, United Kingdom
Platform Optimization and Systems Engineering: Contributing to improvements in HPC scheduling and infrastructure efficiency. Assist with performance benchmarking activities, to help support platform decisions. Observability and Performance Analysis: Developing dashboards and visualisations. Cost and Resource Optimisation: Analysing Cost and Usage Reports (CUR) and infrastructure metrics to support forecasting, budgeting ...

Artificial Intelligence Engineer

Location
Cambridge, England, United Kingdom
solutions in enterprise or regulated environments — aviation, land transport, public safety, telecommunications, or government Familiarity with cloud platforms, APIs, CI/CD, analytics, and observability Experience shaping early-stage products with users and iterating based on feedback Additional information Why Join NCS? Grow with Us Work on cutting-edge ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Cambridge, Cambridgeshire, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Senior Security Platform Architect (SCA & Backend)

Location
Cambridge, England, United Kingdom
design and implement backend services, Python APIs, and workflow components to enable tool onboarding, analysis execution, results processing, and delivery, while improving scalability and observability across the platform. #J-18808-Ljbffr ...

Senior ML Infra Engineer for Scalable Conversational AI

Location
Cambridge, England, United Kingdom
product, data and platform partners. This hands-on role spans software architecture, ML lifecycle decisions and production operations to improve training, evaluation, deployment, and observability of models at scale. #J-18808-Ljbffr ...

AWS Solutions Architect - Microservices & Event-Driven

Location
Cambridge, England, United Kingdom
translate complex requirements into production-grade solutions. You will define end-to-end architectures, lead domain-driven design, and set patterns for resilience, observability, and security. #J-18808-Ljbffr ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Site Reliability Engineer (SRE)

Location
Cambridge, England, United Kingdom
availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications. Join Altium as a Senior Site Reliability Engineer … ensure the reliability and performance of the Altium Cloud Platforms. Key Responsibilities: Understanding how an Altium Cloud Platform works Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection. Develop and implement reliability frameworks and patterns that standardize and elevate ...

Senior SRE: Cloud Platform Reliability & Automation

Location
Cambridge, England, United Kingdom
Engineer to ensure the reliability, availability, and performance of our large-scale cloud platforms and SaaS products. Your work will automate operational tasks, improve observability, and contribute to incident management while collaborating with development teams to build more reliable and scalable applications across regions. #J-18808-Ljbffr ...

Operations Team Lead — Production Reliability & Scale

Location
Cambridge, England, United Kingdom
seeking an Operations Team Lead to own production and scale systems. You will lead operational excellence across live customer-facing platforms, ensuring reliability, observability, and proactive improvements. This hands-on role involves shaping processes, guiding incidents, building the team, and moving from firefighting to sustainable reliability engineering. The ideal candidate ...

Mid Software Engineer – End-to-End Platform Ownership (Hybrid)

Location
Cambridge, England, United Kingdom
from day one. The role offers hybrid working (in-office every two weeks) with Cambridge office, exposure to C#/.NET, React, Azure, and observability tooling, and opportunities to grow within a supportive #J-18808-Ljbffr ...

Senior Backend Engineer

Location
Cambridge, England, United Kingdom
product lives and dies by — the ingestion pipelines that turn warehouse data into a model, the APIs, durable storage, background work, and the observability that lets a small team operate them with confidence at 3 am. WareBee runs on two engines: Physical AI — a living, spatial model of the warehouse ...

Principal Software Architect

Location
Cambridge, England, United Kingdom
/software ecosystem. Assess the architectural impact of new technologies. Be aware of the usability, performance, reliability, maintainability, testability, security and observability constraints on the software architecture. Prototyping and validating architectural concepts through proof-of-concept implementations. Contribute to future and/or related product definitions with a forward-looking ...

Staff Verification Engineer - Media IP

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
Identify technical and schedule risks early, establish mitigation plans and drive them to resolution across teams. Review proposed architecture and design changes for correctness, observability, testability, performance and verification complexity. Lead the resolution of complex unit-level, cross-unit, integration or system-level failures. Align verification activities and dependencies across ...

ML Data & Platform Engineer — Hybrid ML Ops & Pipelines

Location
Cambridge, England, United Kingdom
infrastructure to production ML—owning problems end-to-end to accelerate model delivery. You’ll collaborate with the ML team to improve data quality, observability, and MLOps practices, while scaling infrastructure for faster iteration and reliability. #J-18808-Ljbffr ...