201 to 225 of 990 Incident Management Jobs in England

Infrastructure Administrator - Enterprise Storage - 170807

Location
Newington, England, United Kingdom
systems, network, database, backup, and cybersecurity teams to ensure storage availability and reliability Perform storage upgrades, patching, and maintenance activities in accordance with change management processes Participate in on-call rotations and scheduled maintenance windows Develop automation solutions to improve storage administration and operational efficiency Qualifications of the Storage … enterprise storage environments Hands-on experience supporting SAN and/or NAS platforms, including NetApp storage solutions Experience with storage provisioning, LUN and volume management, RAID configuration, snapshots, replication, and performance tuning Experience supporting Windows and/or Linux servers with attached storage, including multipathing and file system management ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands‐on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement. We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area … resilient, and scalable platforms across Azure, AWS, or both. Design and improve CI/CD pipelines, infrastructure as code, automated testing, release controls, environment management, and deployment strategies. Work with software engineering and data science teams to productionize AI solutions and ensure services are ready to operate reliably ...

Presales Architect - CSP & MSP

Location
Greater London, England, United Kingdom
hybrid and multi-cloud environments with unprecedented visibility and control. Built on a foundation of industry-leading solutions, including OpsRamp for intelligent monitoring and incident management, Morpheus Enterprise for hybrid cloud orchestration and automation, Morpheus VM Essentials for virtualized infrastructure, and Zerto for data protection and disaster recovery. … suite delivers unified operations management across on-premises, private cloud, and public cloud resources. By combining advanced machine learning capabilities with these proven technologies, HPE Cloud Operations Suite enables IT teams to proactively identify and resolve issues, optimize performance and costs, ensure business continuity, and accelerate digital transformation initiatives ...

Head of Machine Learning Engineering

Hiring Organisation
Trainline
Location
London, UK
Employment Type
Full-time
conversational AI. Ensure your teams ship end-to-end - from prototyping through deployment, monitoring, and iteration. Raise operational maturity by defining SLIs/SLOs, incident management, and on-call practices for ML systems in production. Shape AI practices and governance. Co-own the operating model for modern … instincts - you know what it takes to run reliable ML at scale – including MLOps tooling and practicesExperience raising operational standards: SLIs/SLOs, monitoring, incident management, and production reliabilityBudget ownership, including infrastructure cost management and vendor oversightComfortable operating at both strategic and technical levels, with strong stakeholder ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective. What you'll be doing You'll build Python-based automation around incident management, operational runbooks and routine … with strong hands-on Python automation skills. You'll also need experience with: Monitoring and observability tools such as Prometheus, Grafana or similar Production incident management and/or on-call environments Automating operational runbooks and repetitive infrastructure processes APIs and systems integration Version-controlled automation and operational ...

IT Service Operations Manager

Location
West Yorkshire, England, United Kingdom
repeat) This is a senior operational leadership role within a brand new IT Command Centre, responsible for real-time service stability, Major Incident Management, shift leadership, supplier performance and operational improvement. You’ll need: Proven experience leading 24/7 IT Operations, NOC or Command Centre environments Strong … Major Incident Management experience ITIL 4 knowledge Experience with ServiceNow or similar ITSM platforms Strong supplier and stakeholder management skills Experience with monitoring/observability tools Confidence making high-pressure operational decisions Leeds – on-site 24/7 rotating shifts, including nights and weekends #J-18808-Ljbffr ...

IT Service Operations Manager

Hiring Organisation
Outsource UK
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£60,000 - £65,000 per annum
repeat) This is a senior operational leadership role within a brand new IT Command Centre, responsible for real-time service stability, Major Incident Management, shift leadership, supplier performance and operational improvement. You’ll need: Proven experience leading 24/7 IT Operations, NOC or Command Centre environments Strong … Major Incident Management experience ITIL 4 knowledge Experience with ServiceNow or similar ITSM platforms Strong supplier and stakeholder management skills Experience with monitoring/observability tools Confidence making high-pressure operational decisions Leeds – on-site 24/7 rotating shifts, including nights and weekends ...

ITIL Change & Release Manager

Location
Hatfield, England, United Kingdom
Request for Changes, are promptly and correctly approved by authorized approvers. Ensure all Change Requests & Releases are kept up to date on the Tool Management System with progress that occurs, including any actions to correct problems and/or to take opportunities to improve service quality. Ensure the change … production environments Mediate all conflicts regarding scheduling, lack of approval or lack of documentation prior to implementation of the change Coordinate with Incident Management, Problem Management and Configuration Management to ensure correct and consistent data is provided to the Change and Release Management processes Ensure ...

Head of Machine Learning Engineering

Location
Greater London, England, United Kingdom
conversational AI. Ensure your teams ship end-to-end - from prototyping through deployment, monitoring, and iteration. Raise operational maturity by defining SLIs/SLOs, incident management, and on-call practices for ML systems in production. Shape AI practices and governance. Co-own the operating model for modern … know what it takes to run reliable ML at scale – including MLOps tooling and practices Experience raising operational standards: SLIs/SLOs, monitoring, incident management, and production reliability Budget ownership, including infrastructure cost management and vendor oversight Comfortable operating at both strategic and technical levels, with strong ...

Service Asset & Configuration Manager

Hiring Organisation
Experis
Location
Corsham, Wiltshire, South West, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £550 per day
critical secure environment. This is an excellent opportunity to play a key role in the development, governance, and continual improvement of Service Asset & Configuration Management processes across a complex multi-domain infrastructure. Working within a highly secure environment, you will take ownership of SACM activities, ensuring the accuracy, integrity … effectiveness of the Configuration Management System (CMS) and associated processes. You will work closely with operational teams, service management functions, and stakeholders to drive best practice, governance, and continuous improvement across the estate. Key Experience: * Lead the creation, implementation, and continuous improvement of SACM processes, procedures, governance ...

Manager - DE - Technology Consulting - FS

Location
Greater London, England, United Kingdom
responsible for leading software engineering delivery across EY FSO client engagements. The role combines engineering leadership, solution design, cloud-native development, Agile delivery, stakeholder management and production readiness in regulated financial services environments. The opportunity EY FSO Technology Consulting is expanding its Digital Engineering capability to help financial services … close enough to the code to assure quality and delivery outcomes. Key responsibilities Lead end-to-end engineering workstreams across banking, insurance, wealth, asset management or capital markets clients, ensuring high-quality delivery against scope, plan, risk and value outcomes. Design and guide implementation of scalable digital solutions using ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
Cloud Infrastructure Engineer, your main focus will be to ensure our services remain highly available and performant. Key responsibilities include: Cloud Infrastructure Management Manage and support AWS infrastructure , focusing on scalability, security, and reliability. Handle deployments, managing CI/CD pipelines for both containerised (Docker/ECS) and serverless … data stores (Aurora MySQL, DynamoDB) Drive cost optimisation across the infrastructure — monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues ...

Infrastructure Operations Engineer (Operational Comms)

Location
Devizes, England, United Kingdom
effective delivery of life-critical operations. You’ll contribute to the stability, proactive monitoring and continuous improvement of operational communications services, supporting rapid incident response, service resilience and operational readiness. What you will be doing Support the operation of Airwave radio and associated operational communications infrastructure, ensuring services … procedures. Identify, log and categorise operational communications incidents, ensuring accurate information capture and timely escalation to senior engineers, suppliers or on-call support. Support incident management activities for live operational communications issues, including assisting with diagnostics, restoration actions and communication updates. Assist with major incident support, ensuring ...

Platform Engineering Manager (SRE)

Location
Bracknell, England, United Kingdom
Best Workplaces by Great Place to Work for three years running, and building the next generation of our Intelligent Change Management platform with Klario, we are scaling the engineering foundations that keep our cloud products fast, reliable, and trusted by some of the world's largest SAP-run businesses. … that accelerate software delivery and operational performance. Build and maintain core platform capabilities: CI/CD pipelines, build and release automation, infrastructure automation, environment management, observability, and developer tooling. Drive standardisation and reduce engineering friction through automation and self-service. Site Reliability Engineering Introduce and embed SRE practices across ...

Operations Team Lead (Production & Reliability)

Location
Watford, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Wolverhampton, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Wakefield, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Guildford, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Wollaton, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Southampton, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Slough, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
High Wycombe, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Norwich, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Stockport, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Operations Team Lead (Production & Reliability)

Location
Bath, England, United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High‐signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...