5 of 5 Permanent Incident Management Jobs in Cambridge

Systems Infrastructure Engineer

Hiring Organisation
SoCode Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£45000 - £50000/annum
join its IT department. Reporting directly to the Head of IT , this role offers the opportunity to lead on infrastructure design, service improvement, major incident management and the delivery of strategic technology projects across a complex and evolving environment. This is an ideal position for a senior technical … secure, resilient and high-performing IT services. The Role As Systems Engineer, you will act as a senior technical authority responsible for the design, management and continual improvement of critical infrastructure platforms, network services and enterprise applications. You will take ownership of key systems, provide third-line support ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
continuous improvement of the Bango Platform end-to-end — from the infrastructure and pipelines that build and deploy it, to the observability and incident response that keep it running to agreed service levels. You combine three things that have historically sat in separate teams: platform and cloud infrastructure engineering … automation and delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. You design, build and operate the automation, observability and platform capabilities that other engineering teams rely on, and you are equally comfortable diagnosing a live incident as you are designing ...

Principal AI Platform Engineer

Location
Cambridge, England, United Kingdom
runtime platforms that make AI services reliable, secure, observable and supportable at Arm scale. You will work across Kubernetes, cloud, identity, secrets, networking, telemetry, incident management and automation to provide the production foundation for Arm's AI platform. Participate in production support and our paid on-call rota … high impact incident response. Production AI runtime platforms: Build/deploy and operate the infrastructure for centrally hosted AI platform services, including MCP server infrastructure, model gateway services and supporting control-plane components. Design runtime patterns for isolation, scalability, secure execution, capacity management and cost-aware operation. Automate ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Senior Infrastructure Engineer – UK & Ireland

Location
Cambridge, England, United Kingdom
business stakeholders, ensuring high availability and security across critical platforms. You will mentor colleagues, oversee data centre operations, and drive resilience through governance and incident management while shaping infrastructure strategy for ZEISS Ltd. #J-18808-Ljbffr ...