18 of 18 Incident Management Jobs in Cambridgeshire

Systems Infrastructure Engineer

Hiring Organisation
SoCode Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£45000 - £50000/annum
join its IT department. Reporting directly to the Head of IT , this role offers the opportunity to lead on infrastructure design, service improvement, major incident management and the delivery of strategic technology projects across a complex and evolving environment. This is an ideal position for a senior technical … secure, resilient and high-performing IT services. The Role As Systems Engineer, you will act as a senior technical authority responsible for the design, management and continual improvement of critical infrastructure platforms, network services and enterprise applications. You will take ownership of key systems, provide third-line support ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
continuous improvement of the Bango Platform end-to-end — from the infrastructure and pipelines that build and deploy it, to the observability and incident response that keep it running to agreed service levels. You combine three things that have historically sat in separate teams: platform and cloud infrastructure engineering … automation and delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. You design, build and operate the automation, observability and platform capabilities that other engineering teams rely on, and you are equally comfortable diagnosing a live incident as you are designing ...

System Support Engineer

Location
Cambridge, England, United Kingdom
play a critical role in ensuring continuity, stability, and quality of service for all customer installations. This role will provide front‐line technical support, incident management, software maintenance coordination, and proactive system monitoring for Ubisense real‐time location and smart factory systems deployed at major customers. The Support … critical systems (automotive manufacturing experience is a plus). Previous roles in support, NOC, field service, IT operations, or application support. Exposure to change management, incident management, or ITIL practices. Working Conditions Office based providing remote support during operational hours and service windows (coverage varies by customer ...

Principal AI Platform Engineer

Location
Cambridge, England, United Kingdom
runtime platforms that make AI services reliable, secure, observable and supportable at Arm scale. You will work across Kubernetes, cloud, identity, secrets, networking, telemetry, incident management and automation to provide the production foundation for Arm's AI platform. Participate in production support and our paid on-call rota … high impact incident response. Production AI runtime platforms: Build/deploy and operate the infrastructure for centrally hosted AI platform services, including MCP server infrastructure, model gateway services and supporting control-plane components. Design runtime patterns for isolation, scalability, secure execution, capacity management and cost-aware operation. Automate ...

Senior Site Reliability Engineer

Location
Cambridge, England, United Kingdom
Bango Platform. You hold everything expected of a Site Reliability Engineer — owning reliability end-to-end across the infrastructure, delivery pipelines, observability and incident response that keep the platform running to agreed service levels — but at greater scope, complexity and influence, and you take responsibility for lifting the capability … that have historically sat in separate teams: platform and cloud infrastructure engineering, automation and delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. What sets the senior role apart is not simply doing this work to a higher standard, but becoming ...

Data Centre Manager

Location
Peterborough, England, United Kingdom
largest banking organisations. This is a high-profile leadership role with accountability for multiple data centres, critical engineering infrastructure, operational performance, stakeholder management and service excellence across a complex and highly regulated environment. The successful candidate will lead a team of Technical Managers and approximately 55-60 engineers, ensuring … systems that support nationally important banking services. Data centre experience is essential for this position. What You'll Be Doing Leading the operational management of multiple data centres across Peterborough and Corby. Managing Technical Managers and a team of circa 55-60 engineers operating within a 24/ ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Senior Software Development Engineer (SRE)

Location
Cambridge, England, United Kingdom
large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications. Join Altium as a Senior Site Reliability Engineer to ensure … Altium Cloud Platforms. Key Responsibilities: Understanding how an Altium Cloud Platform works Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection. Develop and implement reliability frameworks and patterns that standardize and elevate the resilience of our SaaS products across ...

Performance Manager (6 Months FTC, Shifts)

Hiring Organisation
M Group
Location
St. Ives, Cambridgeshire, East Anglia, United Kingdom
Employment Type
Permanent
best technology, manage assets and refresh systems. With 24/7 national operations, we keep things running smoothly, while operating comprehensive network or service management repair and maintenance to keep everything running operationally. Want to come and be a part of it? We're looking for a proactive … teams, customers, and stakeholders to maintain service excellence. Alongside operational leadership, you'll play a key role in developing your team through coaching, performance management, wellbeing support, and creating an inclusive environment where colleagues feel empowered to succeed. What youll bring Experience leading teams within a Service Desk, National ...

Principal Software Engineer

Location
Peterborough, England, United Kingdom
specialists to convert data, models, and experimentation into reliable, usable product capabilities. Production ownership experience: managing partial failure, resilience, blast‐radius thinking, and incident management. Distributed‐systems fluency: service boundaries, transactional thinking, multi‐tenant design, API design, and versioning. Strong architectural communication skills, able to explain, document, and defend ...

Datacenter Operations Manager

Hiring Organisation
Mitie
Location
Peterborough, England, United Kingdom
senior operational authority across multiple sites, responsible for the resilience, performance and compliance of critical engineering infrastructure and integrated facilities management services. Working closely with the Strategic Client Director and supported by Technical IFM Managers at each location, you'll lead teams, suppliers and stakeholders to deliver outstanding service … role that combines strategic leadership with hands-on operational oversight. You'll drive engineering excellence across electrical and mechanical systems, oversee business continuity and incident management, ensure robust compliance and governance standards, and shape long-term asset and lifecycle strategies. You'll also be accountable for the successful ...

Senior Data Centre Operations Manager - Multi-Site

Location
Peterborough, England, United Kingdom
ensuring safe operation of power, cooling and M&E systems that support banking services. The role requires proven leadership in data centres, strong stakeholder management and a hands-on approach to incident management, performance, and risk across multiple #J-18808-Ljbffr ...

IT Service Desk Team Lead — Hands-On Leadership

Location
Cambridge, England, United Kingdom
Desk Team Lead to enhance IT support experiences for staff. This role combines leadership and hands-on technical skills, focusing on service improvement and incident management. Candidates should have experience in IT support leadership and a strong commitment to user-focused service delivery. The position requires mostly on-site ...

Senior SRE: Observability, Automation & Reliability

Location
Cambridge, England, United Kingdom
Engineer to ensure the reliability, availability, and performance of large-scale SaaS platforms across regions. You will automate operations, improve observability, and contribute to incident management in collaboration with DevOps and engineering teams. Join a team that champions IaC, SOO principles, and scalable deployments, driving reliability improvements … faster incident resolution across the cloud platform. #J-18808-Ljbffr ...

RTLS Support Engineer – Industrial IoT & Smart Factory

Location
Cambridge, England, United Kingdom
Support Analyst to ensure continuity and quality of service for customer installations of Ubisense RTLS systems. You will deliver front-line technical support, incident management, and maintain proactive monitoring across on-site and cloud environments for major customers. The role requires expertise in IT support, log analysis ...

Senior Infrastructure Engineer – UK & Ireland

Location
Cambridge, England, United Kingdom
business stakeholders, ensuring high availability and security across critical platforms. You will mentor colleagues, oversee data centre operations, and drive resilience through governance and incident management while shaping infrastructure strategy for ZEISS Ltd. #J-18808-Ljbffr ...

Smart Factory Support Engineer (Remote & On-Call)

Location
Cambridge, England, United Kingdom
Group is seeking a Support Analyst to ensure continuity, stability and quality of service for customer installations. You will provide front-line technical support, incident management, software maintenance coordination and proactive system monitoring for Ubisense RTLS and smart factory systems at major customers. You will become an expert ...

Executive Assistant to Founder & CEO

Hiring Organisation
MIM®
Location
Cambridge, England, United Kingdom
base + up to £10,000 annual performance bonus (£57,000 OTE). Join the world’s leading institute for Major Incident Management. MIM® is the world’s leading institute for Major Incident Management. Organisations in more than 98 countries use our internationally recognised frameworks, training and certifications ...