3 of 3 Incident Management Jobs in Cambridgeshire

Principal AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Cambridge, England, United Kingdom
runtime platforms that make AI services reliable, secure, observable and supportable at Arm scale. You will work across Kubernetes, cloud, identity, secrets, networking, telemetry, incident management and automation to provide the production foundation for Arm's AI platform. Participate in production support and our paid on-call rota … high impact incident response. Production AI runtime platforms: Build/deploy and operate the infrastructure for centrally hosted AI platform services, including MCP server infrastructure, model gateway services and supporting control-plane components. Design runtime patterns for isolation, scalability, secure execution, capacity management and cost-aware operation. Automate ...

Technical Delivery Manager

Hiring Organisation
Cambridge University Press & Assessment
Location
Cambridge, Cambridgeshire, East Anglia, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
service transition standards. Leading specialist 2nd and 3rd line support services, ensuring service levels are consistently achieved or exceeded. Driving ITIL-aligned service management practices, incident resolution and continuous improvement initiatives. Owning security, access governance, data governance and audit readiness across supported platforms. Building strong relationships across Technology … Strong understanding of Workday security, access governance, data governance, audit readiness and data privacy. Proven experience delivering solutions using agile methodologies. Significant experience of incident management, change management and IT service management practices. Experience leading and developing technical teams. Proactive approach to continuous improvement and ensuring ...

Principal Software Engineer

Hiring Organisation
Jobleads-UK
Location
Peterborough, England, United Kingdom
specialists to convert data, models, and experimentation into reliable, usable product capabilities. Production ownership experience: managing partial failure, resilience, blast‐radius thinking, and incident management. Distributed‐systems fluency: service boundaries, transactional thinking, multi‐tenant design, API design, and versioning. Strong architectural communication skills, able to explain, document, and defend ...