16 of 16 Incident Management Jobs in Cambridgeshire

Systems Infrastructure Engineer

Hiring Organisation
SoCode Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£45000 - £50000/annum
join its IT department. Reporting directly to the Head of IT , this role offers the opportunity to lead on infrastructure design, service improvement, major incident management and the delivery of strategic technology projects across a complex and evolving environment. This is an ideal position for a senior technical … secure, resilient and high-performing IT services. The Role As Systems Engineer, you will act as a senior technical authority responsible for the design, management and continual improvement of critical infrastructure platforms, network services and enterprise applications. You will take ownership of key systems, provide third-line support ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
continuous improvement of the Bango Platform end-to-end — from the infrastructure and pipelines that build and deploy it, to the observability and incident response that keep it running to agreed service levels. You combine three things that have historically sat in separate teams: platform and cloud infrastructure engineering … automation and delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. You design, build and operate the automation, observability and platform capabilities that other engineering teams rely on, and you are equally comfortable diagnosing a live incident as you are designing ...

System Support Engineer

Location
Cambridge, England, United Kingdom
play a critical role in ensuring continuity, stability, and quality of service for all customer installations. This role will provide front‐line technical support, incident management, software maintenance coordination, and proactive system monitoring for Ubisense real‐time location and smart factory systems deployed at major customers. The Support … critical systems (automotive manufacturing experience is a plus). Previous roles in support, NOC, field service, IT operations, or application support. Exposure to change management, incident management, or ITIL practices. Working Conditions Office based providing remote support during operational hours and service windows (coverage varies by customer ...

Senior Site Reliability Engineer

Location
Cambridge, England, United Kingdom
Bango Platform. You hold everything expected of a Site Reliability Engineer — owning reliability end-to-end across the infrastructure, delivery pipelines, observability and incident response that keep the platform running to agreed service levels — but at greater scope, complexity and influence, and you take responsibility for lifting the capability … that have historically sat in separate teams: platform and cloud infrastructure engineering, automation and delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. What sets the senior role apart is not simply doing this work to a higher standard, but becoming ...

Data Centre Manager

Location
Peterborough, England, United Kingdom
largest banking organisations. This is a high-profile leadership role with accountability for multiple data centres, critical engineering infrastructure, operational performance, stakeholder management and service excellence across a complex and highly regulated environment. The successful candidate will lead a team of Technical Managers and approximately 55-60 engineers, ensuring … systems that support nationally important banking services. Data centre experience is essential for this position. What You'll Be Doing Leading the operational management of multiple data centres across Peterborough and Corby. Managing Technical Managers and a team of circa 55-60 engineers operating within a 24/ ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Head of IT Operations

Hiring Organisation
Michael Page Technology
Location
Peterborough, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £80,000 per annum
technology operations leader to take ownership of business-as-usual IT services across a complex organisation. This role is responsible for service delivery, service management, application support, supplier performance, operational resilience, and continuous improvement, ensuring reliable, accessible, and cost-effective technology services that support organisational success. Client Details This … Description Lead the end-to-end delivery of technology operations, ensuring high-quality, reliable and accessible IT services. Own and enhance the IT service management framework, including incident, problem, change, request, configuration and knowledge management processes. Define and manage service catalogues, SLAs, OLAs, KPIs and performance reporting. ...

Site Reliability Engineer (SRE)

Location
Cambridge, England, United Kingdom
large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications. Join Altium as a Senior Site Reliability Engineer to ensure … Altium Cloud Platforms. Key Responsibilities: Understanding how an Altium Cloud Platform works Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection. Develop and implement reliability frameworks and patterns that standardize and elevate the resilience of our SaaS products across ...

Performance Manager (6 Months FTC, Shifts)

Hiring Organisation
M Group
Location
St. Ives, Cambridgeshire, East Anglia, United Kingdom
Employment Type
Permanent
best technology, manage assets and refresh systems. With 24/7 national operations, we keep things running smoothly, while operating comprehensive network or service management repair and maintenance to keep everything running operationally. Want to come and be a part of it? We're looking for a proactive … teams, customers, and stakeholders to maintain service excellence. Alongside operational leadership, you'll play a key role in developing your team through coaching, performance management, wellbeing support, and creating an inclusive environment where colleagues feel empowered to succeed. What youll bring Experience leading teams within a Service Desk, National ...

Principal Software Engineer

Location
Peterborough, England, United Kingdom
specialists to convert data, models, and experimentation into reliable, usable product capabilities. Production ownership experience: managing partial failure, resilience, blast‐radius thinking, and incident management. Distributed‐systems fluency: service boundaries, transactional thinking, multi‐tenant design, API design, and versioning. Strong architectural communication skills, able to explain, document, and defend ...

Senior Data Centre Operations Manager - Multi-Site

Location
Peterborough, England, United Kingdom
ensuring safe operation of power, cooling and M&E systems that support banking services. The role requires proven leadership in data centres, strong stakeholder management and a hands-on approach to incident management, performance, and risk across multiple #J-18808-Ljbffr ...

IT Service Desk Team Lead — Hands-On Leadership

Location
Cambridge, England, United Kingdom
Desk Team Lead to enhance IT support experiences for staff. This role combines leadership and hands-on technical skills, focusing on service improvement and incident management. Candidates should have experience in IT support leadership and a strong commitment to user-focused service delivery. The position requires mostly on-site ...

Senior SRE: Cloud Platform Reliability & Automation

Location
Cambridge, England, United Kingdom
reliability, availability, and performance of our large-scale cloud platforms and SaaS products. Your work will automate operational tasks, improve observability, and contribute to incident management while collaborating with development teams to build more reliable and scalable applications across regions. #J-18808-Ljbffr ...

RTLS Support Engineer – Industrial IoT & Smart Factory

Location
Cambridge, England, United Kingdom
Support Analyst to ensure continuity and quality of service for customer installations of Ubisense RTLS systems. You will deliver front-line technical support, incident management, and maintain proactive monitoring across on-site and cloud environments for major customers. The role requires expertise in IT support, log analysis ...

Senior Infrastructure Engineer – UK & Ireland

Location
Cambridge, England, United Kingdom
business stakeholders, ensuring high availability and security across critical platforms. You will mentor colleagues, oversee data centre operations, and drive resilience through governance and incident management while shaping infrastructure strategy for ZEISS Ltd. #J-18808-Ljbffr ...

Smart Factory Support Engineer (Remote & On-Call)

Location
Cambridge, England, United Kingdom
Group is seeking a Support Analyst to ensure continuity, stability and quality of service for customer installations. You will provide front-line technical support, incident management, software maintenance coordination and proactive system monitoring for Ubisense RTLS and smart factory systems at major customers. You will become an expert ...