276 to 300 of 1,206 Incident Management Jobs in the UK

Systems Development Manager, AWS Managed Operations

Hiring Organisation
Amazon
Location
London, United Kingdom
Salary
£ 80 K
ready to transform how support engineering operates at scale? Amazon Web Services (AWS) Operations Management is seeking an experienced Support Engineering Manager to lead a global team that's evolving from reactive troubleshooting to proactive automation and engineering excellence using Agentic Al. In this role, you'll eliminate repetitive … engineers grow their careers through strategic problem-solving, not just ticket resolution.You'll have the opportunity to create transformational impact across AWS by leading incident management at massive scale, championing automation-driven solutions, and mentoring engineers to think beyond immediate fixes. This isn't traditional support management ...

Lead DevOps Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, United Kingdom
Salary
£ 80 K
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands-on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement.We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve … observable, resilient, and scalable platforms across Azure, AWS, or both.Design and improve CI/CD pipelines, infrastructure as code, automated testing, release controls, environment management, and deployment strategies.Work with software engineering and data science teams to productionize AI solutions and ensure services are ready to operate reliably at scale.Establish ...

Technical Account Manager – Managed Services

Location
Peterborough, England, United Kingdom
reviews, governance forums, QBRs, and service reviews Track lifecycle, end-of-life, patching, firmware, obsolescence, capacity, and technical debt risks Provide technical governance across incident, problem, and change management Lead major incident escalations, root cause analysis, and post-incident reviews Drive continual service improvement, automation, optimisation … relationships and escalations Record and track service-led opportunities via Salesforce Requirements ITIL v3/v4 Foundation required Proven experience in a technical account management, senior technical support, or technical consulting role within an MSP or IT services environment Strong understanding of managed services and broader technology portfolios, including ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective. What you'll be doing You'll build Python-based automation around incident management, operational runbooks and routine … with strong hands-on Python automation skills. You'll also need experience with: Monitoring and observability tools such as Prometheus, Grafana or similar Production incident management and/or on-call environments Automating operational runbooks and repetitive infrastructure processes APIs and systems integration Version-controlled automation and operational ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands-on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement.We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve … resilient, and scalable platforms across Azure, AWS, or both.* Design and improve CI/CD pipelines, infrastructure as code, automated testing, release controls, environment management, and deployment strategies.* Work with software engineering and data science teams to productionize AI solutions and ensure services are ready to operate reliably ...

Proofpoint SME - Hybrid in Wokingham - 6 Month Initial Contract - Inside IR35

Hiring Organisation
Hamilton Barnes
Location
Wokingham, Berkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP Daily
technologies, threat intelligence, and advanced email threat analysis and remediation. Email authentication protocols - Strong hands-on knowledge of SPF, DKIM, and DMARC configuration and management, with proven experience in email fraud prevention. Microsoft 365 & mail flow architecture - Experience working with Microsoft 365/Exchange Online, Exchange Server, SMTP routing … mail flow architecture within enterprise environments. SIEM & incident management - Familiarity with SIEM platforms such as Splunk, Microsoft Sentinel, or QRadar, with strong analytical, troubleshooting, and incident management capability; PowerShell or Python Scripting is a desirable advantage. Contract Details: Rate: £450 per day Inside IR35 Location: Hybrid ...

Manager, DE , TC, FS

Location
Greater London, England, United Kingdom
responsible for leading software engineering delivery across EY FSO client engagements. The role combines engineering leadership, solution design, cloud-native development, Agile delivery, stakeholder management and production readiness in regulated financial services environments. The opportunity EY FSO Technology Consulting is expanding its Digital Engineering capability to help financial services … close enough to the code to assure quality and delivery outcomes. Key responsibilities Lead end‐to‐end engineering workstreams across banking, insurance, wealth, asset management or capital markets clients, ensuring high‐quality delivery against scope, plan, risk and value outcomes. Design and guide implementation of scalable digital solutions using ...

Head of Machine Learning Engineering

Hiring Organisation
Trainline
Location
London, United Kingdom
Salary
£ 100 K
conversational AI. Ensure your teams ship end-to-end - from prototyping through deployment, monitoring, and iteration. Raise operational maturity by defining SLIs/SLOs, incident management, and on-call practices for ML systems in production.Shape AI practices and governance. Co-own the operating model for modern … instincts - you know what it takes to run reliable ML at scale – including MLOps tooling and practicesExperience raising operational standards: SLIs/SLOs, monitoring, incident management, and production reliabilityBudget ownership, including infrastructure cost management and vendor oversightComfortable operating at both strategic and technical levels, with strong stakeholder ...

Senior Application Engineer (Jira ITSM)

Hiring Organisation
Leonardo DRS
Location
Yeovil, Somerset, United Kingdom
Salary
£ 50 K
underpin critical defence, government and public sector services.What you will do as a Senior Application EngineerSupport, configure, and enhance ITSM platforms (e.g. Jira Service Management) that underpin service delivery and operational workflowsDeliver engineering tasks within defined work packages, ensuring alignment with user needs, security requirements, and service-level objectives.Contribute … application architecture, integration, and lifecycle activities under the guidance of senior engineers.Collaborate with stakeholders across engineering, service management, and business teams to resolve technical issues.Mentor junior engineers and specialists, supporting their development and promoting best practices.Participate in continuous improvement initiatives, including automation, documentation, and knowledge sharing.Ensure compliance with engineering ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
Cloud Infrastructure Engineer, your main focus will be to ensure our services remain highly available and performant. Key responsibilities include: Cloud Infrastructure Management Manage and support AWS infrastructure , focusing on scalability, security, and reliability. Handle deployments, managing CI/CD pipelines for both containerised (Docker/ECS) and serverless … data stores (Aurora MySQL, DynamoDB) Drive cost optimisation across the infrastructure — monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues ...

Data Center Engineer

Hiring Organisation
Radius
Location
City of London, London, United Kingdom
Provide a value-added service to customers. MAIN DUTIES, INCLUDING BUT NOT RESTRICTED TO: Receive and act on all customer Incidents\Events\Request notifications Incident Management – Prioritises and diagnoses, investigates incidents according to agreed procedures. Incident Management – Facilitates recovery, following resolution of Incidents, documents and closes … resolved Incidents, Events and/or Requests via agreed procedures. Implementation of agreed IT Infrastructure changes and following maintenance routines Network Support – Uses Network Management tools where applicable to investigate and diagnose network problems, collect performance statistics, and create reports Customer Service - Liaise with customers regarding fault correction progress ...

Platform Engineering Manager (SRE)

Location
Bracknell, England, United Kingdom
Best Workplaces by Great Place to Work for three years running, and building the next generation of our Intelligent Change Management platform with Klario, we are scaling the engineering foundations that keep our cloud products fast, reliable, and trusted by some of the world's largest SAP-run businesses. … that accelerate software delivery and operational performance. Build and maintain core platform capabilities: CI/CD pipelines, build and release automation, infrastructure automation, environment management, observability, and developer tooling. Drive standardisation and reduce engineering friction through automation and self-service. Site Reliability Engineering Introduce and embed SRE practices across ...

Head of Machine Learning Engineering

Location
Greater London, England, United Kingdom
conversational AI. Ensure your teams ship end-to-end - from prototyping through deployment, monitoring, and iteration. Raise operational maturity by defining SLIs/SLOs, incident management, and on-call practices for ML systems in production. Shape AI practices and governance. Co-own the operating model for modern … know what it takes to run reliable ML at scale – including MLOps tooling and practices Experience raising operational standards: SLIs/SLOs, monitoring, incident management, and production reliability Budget ownership, including infrastructure cost management and vendor oversight Comfortable operating at both strategic and technical levels, with strong ...

Mobile App Delivery & Platform Manager

Location
Milton Keynes, England, United Kingdom
team and with other functions across the organisation. The successful candidate will have a strong understanding of mobile app technologies, delivery methodologies and service management and will have experience working within either/or retail and hospitality businesses. Key Responsibilities Solution Ownership & Delivery Oversight … Collaboration: Partner with Enterprise Architecture to define and validate solutions, ensuring alignment with enterprise standards, integrations and long-term technical direction. Vendor & Delivery Partner Management: Manage third-party development partners and platform providers, ensuring quality of delivery and adherence to agreed standards and outcomes. Integration & Ecosystem Coordination: Oversee integration ...

Senior Cyber Security Engineer

Hiring Organisation
ENSEK
Location
Nottingham, Nottinghamshire, United Kingdom
Salary
£ 70 K
Define and enforce secure configurations, network segmentation, identity and access controls for public cloud (primarily AWS).Application & infrastructure hardening: Implement secure coding practices, vulnerability management, secrets management and runtime protections for services and CI/CD pipelines.Detection & response: Build and maintain monitoring, logging and alerting for security events … lead incident response and post‐incident reviews to drive remediation and lessons learned.Incident Management: Support ENSEK’s 24/7 Incident Management processes to ensure security and stability for clients.Automation & tooling: Automate security checks, policy enforcement and remediation using IaC, CI/CD integrations ...

Operations Team Lead (Production & Reliability)

Hiring Organisation
Complexio
Location
United Kingdom
already run on - turning organisational knowledge into measurable business outcomes. A joint venture between Hafnia and Símbolo, backed by leading maritime partners including Marfin Management, C Transport Maritime, Trans Sea Transport, and BW Epic Kosan. Born in maritime, now scaling rapidly across industries. We are now onboarding enterprise customers … Operational readiness for new releases Safe production access and change coordination Production is a high-discipline environment. You make sure it stays that way. Incident Management You own the full lifecycle: High-signal alerting and fast detection Structured incident response Clear internal and customer communication Blameless postmortems ...

Senior Application Engineer (Jira ITSM)

Hiring Organisation
Hackajob Ltd
Location
Yeovil, Somerset, South West, United Kingdom
Employment Type
Permanent, Work From Home
defence, government and public sector services. What you will do as a Senior Application Engineer Support, configure, and enhance ITSM platforms (e.g. Jira Service Management) that underpin service delivery and operational workflows Deliver engineering tasks within defined work packages, ensuring alignment with user needs, security requirements, and service-level … objectives. Contribute to application architecture, integration, and lifecycle activities under the guidance of senior engineers. Collaborate with stakeholders across engineering, service management, and business teams to resolve technical issues. Mentor junior engineers and specialists, supporting their development and promoting best practices. Participate in continuous improvement initiatives, including automation, documentation ...

Data Centre Operations Manager , LHR Data Center Operations (DCO)

Hiring Organisation
Amazon
Location
London, United Kingdom
Salary
£ 70 K
CentersConstantly improving all our processes and procedures. We believe there is nothing we cannot improveAssisting & managing relationships with external vendors & contractorsLiaising with internal teams & management groupsEnsuring we adhere to and exceed local Health & Safety standards in all our Data CentersCreating and maintaining metrics on all aspects of our Data … Centers and utilizing those metrics to drive positive changesAssisting in implementing service methodologies including incident management, problem management, change management and capacity managementA day in the lifeAbout AWSDiverse ExperiencesAWS values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed ...

DevOps Engineer - Glasgow

Location
Glasgow, Scotland, United Kingdom
roadmaps and deployment strategies. Provide technical leadership and guidance, sharing DevOps best practices across teams. Ensure environments remain secure, compliant and operationally resilient. Support incident management activities and assist with root cause analysis and service restoration. Required Skills & Experience Essential Strong experience working as a DevOps Engineer … enterprise environments. Advanced experience in Infrastructure Service Incident Management . Proven expertise designing and maintaining CI/CD pipelines . Experience with infrastructure automation and deployment tooling. Strong knowledge of cloud technologies and modern DevOps practices. Experience with containerisation and orchestration technologies. Ability to troubleshoot and resolve complex ...

Director, Incident Management & Service Continuity

Location
Greater London, England, United Kingdom
Hines is seeking a Director - REPAC to lead a strategic incident management framework tailored to JPMC’s needs, ensuring high-quality, compliant, and efficient services while mitigating operational risks. You will steer service evolution and align delivery with SLAs and governance. You will supervise REPAC managers and staff … foster accountability and high performance, and collaborate across functions to drive consistency and excellence in incident response and resilience. #J-18808-Ljbffr ...

Senior software engineer (Node.js/TypeScript)

Location
Bath, England, United Kingdom
operate in a build‐and‐run model, with teams responsible for the full software lifecycle, from architecture and delivery through to operational support and incident management. Our core platform is built on AWS serverless technologies including Lambda, SQS, EventBridge, API Gateway, S3, and ECS. We primarily develop in TypeScript … Infrastructure is managed through a combination of Terraform and Serverless Framework. We use GitHub Actions for CI/CD and incident.io to support our incident management processes and operational workflows. We care deeply about engineering quality, pragmatic architecture, and continuous improvement. If you’d like to go deeper ...

Senior software engineer (Node.js/TypeScript)

Location
Greater London, England, United Kingdom
operate in a build-and-run model, with teams responsible for the full software lifecycle, from architecture and delivery through to operational support and incident management. Our core platform is built on AWS serverless technologies including Lambda, SQS, EventBridge, API Gateway, S3, and ECS. We primarily develop in TypeScript … Infrastructure is managed through a combination of Terraform and Serverless Framework. We use GitHub Actions for CI/CD and incident.io to support our incident management processes and operational workflows. We care deeply about engineering quality, pragmatic architecture, and continuous improvement. If you’d like to go deeper ...

Technical Lead

Hiring Organisation
London Stock Exchange Group
Location
London, United Kingdom
Salary
£ 80 K
/Jenkins/GitHub)Configuring and running Code/Binary scans using solutions like SonarQube, Semgrep, Blackbuck, Trivy, GitLeaks Veracode, etc.Configuring and using Secrets management tools like Vault and Cloud native solutionsBroad knowledge of SDLC Tools, specifically Build, Test and Deploy Automation tools, e.g., Maven, Gradle, Selenium, Ansible, etc.Good … supported systems is continually updated. To provide a high level of customer service, whilst working under pressure. To follow and adhere to established Incident Management, Change Management and Problem Management procedures. LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving ...

Associate Medical Director

Location
Oxford, England, United Kingdom
This means setting and maintaining clinical policies and protocols to the standard our regulator expects, chairing clinical governance and risk review, ensuring robust audit, incident reporting and learning cycles are embedded across the team, and acting as the primary clinical point of contact for regulatory engagement and inspection readiness. … clinical workshops, pathway design, and team days. What you’ll do Lead the clinical team and oversee governance processes Provide clinical oversight and management of the multidisciplinary clinical team, including GPs, midwives, and other clinicians Line-manage clinical staff, including providing feedback on clinical care and supporting the recruitment ...

service performance lead

Location
Greater London, England, United Kingdom
drive service performance across internal teams and external vendors, ensuring services meet the needs of baristas and customers. You will focus on service stability, incident management, and continuous improvement, while supporting the transition of new and enhanced services into operations. You will play a key role inanalysingservice performance … driving improvements that enhance the service desk experience and overall technology performance. We’ll look to you to bring your strong experience in service management disciplines (incident, problem, change and vendor management), ideally within a retail or hospitality environment. You’ll be able to apply your ITIL ...