101 to 125 of 498 Incident Management Jobs in London

Data Center Health & Safety Lead - Development

Location
Greater London, England, United Kingdom
testing & commissioning. Your primary mission is to embed a world‐class safety culture across our development sites, ensuring all projects are delivered entirely incident-free. What we are looking for A degree educated professional with a proven track record in the delivery of large scale commercial or industrial projects … demonstrate extensive hands‐on experience managing large‐scale build safety for critical infrastructure, or heavy industrial developments. Policy/Project Planning and Risk Management Develop and implement construction EHS policies, project safety plans, and field procedures to ensure compliance with Verne and customer standards. Oversee project hazard analyses, risk ...

Principal SRE (AWS, Azure, Terraforms, Kubernetes)

Location
Greater London, England, United Kingdom
Story In July 2019, Fourth joined forces with HotSchedules to become the global leader in end-to-end restaurant and hospitality management technology solutions. Together, the merged company now represents the world’s largest and only provider of end-to-end restaurant and hospitality management solutions for customers … global restaurant or hotel chain. The combined company’s complete software-as-service (SaaS) solution suite including scheduling, time & attendance, applicant tracking, training, inventory management/procurement, HR/benefits and payroll services now serves customers in 120,000 locations worldwide and is supported by a dedicated, unified team ...

Head of Clinical Systems

Hiring Organisation
Circle Health Group
Location
London, UK
Employment Type
Full-time
this role include: Lead and manage the Clinical Systems function, including application specialists and support resources. Ensure effective day‐to‐day operation, support, incident management, and change control for all clinical systems. Ensure high‐quality operational support, incident management, and change delivery across all clinical systems. … efficiency and productivitySystem performance, resilience, and availabilityOwn strategic relationships with clinical system suppliers and vendors. Work with Vendor team and be SME lead contract management, supplier performance, and commercial negotiations. Build strong, trusted relationships with senior clinicians, operational leaders, and executive stakeholders. Applicants should meet the following criteria: Significant ...

Head of Clinical Systems

Location
Greater London, England, United Kingdom
this role include: Lead and manage the Clinical Systems function, including application specialists and support resources. Ensure effective day‐to‐day operation, support, incident management, and change control for all clinical systems. Ensure high‐quality operational support, incident management, and change delivery across all clinical systems. … productivity System performance, resilience, and availability Own strategic relationships with clinical system suppliers and vendors. Work with Vendor team and be SME lead contract management, supplier performance, and commercial negotiations. Build strong, trusted relationships with senior clinicians, operational leaders, and executive stakeholders. Applicants should meet the following criteria: Significant ...

TechOps Lead

Location
Greater London, England, United Kingdom
This is a strategically significant role which will require close collaboration with our UK and APAC leaders. You will bring strong leadership skills, service management, and operational experience, with the ability to lead, scale, and mature this function in line with Kraken’s growth trajectory. You will have experience … Significant people leadership experience Scaling teams and services over time Service management, capacity planning, and operational maturity Managing support pipelines, projects, and assignments Building regional processes while balancing global standards Working with internal and external stakeholders and vendors Driving reliability, consistency, and continuous improvement You will bring strong technical ...

SVP, Global Head of Platform Operations

Location
Greater London, England, United Kingdom
production platform. This is a newly consolidated executive role that brings Systems Engineering, SRE/DevOps, Deployment, Platform Engineering, Service Improvement and Release Management into one organization with a single mandate: make TT's delivery faster, safer and more automated while raising reliability for customers who cannot tolerate downtime. … will own the end-to-end deployment pipeline, production observability and incident management, and the SLO/SLI framework that defines what "good" looks like for every customer-facing service. You will build the leadership layer underneath you, shape TT's global operating footprint ...

Field Operations Technical Support Manager

Hiring Organisation
Tiger Resourcing Group
Location
London Area, United Kingdom
Ensure the availability, performance, and reliability of business-critical applications. Act as the senior escalation point for complex technical issues and major incidents. Oversee incident management, problem management, root cause analysis, and service recovery activities. Monitor service performance against agreed KPIs and service levels. Drive continuous improvement … initiatives across support processes, tooling, and operational practices. Work closely with engineering, infrastructure, service management, and business teams to ensure effective service delivery. Support application upgrades, releases, enhancements, and operational readiness activities. Manage workload planning, resource allocation, and team performance. Produce regular reporting on service performance, trends, risks ...

Senior Cloud Engineer, Cloud COE

Location
Greater London, England, United Kingdom
consistent, version‐controlled, and policy‐compliant Manage drift detection, environment consistency, and release governance Azure Core Platform Skills Hands‐on engineering across: Subscriptions, Management Groups, RBAC, Policy Virtual Machines, Storage, Backup, DR Azure Monitor, Log Analytics Private Endpoints and secure service exposure Implement secure‐by‐design configurations aligned … Azure AD) configurations and integrations Implement RBAC, SSO, and least privilege access models Assist with security controls, compliance policies, and remediation activities Monitoring, BAU & Incident Management Support business‐as‐usual (BAU) cloud operations across environments Monitor cloud services using Azure Monitor, Log Analytics, and alerting tools Investigate incidents ...

Assistant Director, Technology Services (EMEA & Asia)

Location
City Of London, England, United Kingdom
improving technology operations, enhancing user experience, and driving the adoption of emerging technologies, including AI-enabled solutions. This role combines leadership, operational oversight, stakeholder management, and transformation initiatives, making it an excellent opportunity for an experienced technology leader looking to make a broad business impact within a complex international … approved AI, automation, and productivity-enhancing technologies. Work closely with regional leadership teams to understand evolving business needs and technology priorities. Participate in vendor management, contract negotiations, and budget planning activities. Support major incident management, problem management, and post-incident review processes. Champion best practice ...

Assistant Director, Technology Services (EMEA & Asia)

Hiring Organisation
Ryder Reid Legal Ltd
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
improving technology operations, enhancing user experience, and driving the adoption of emerging technologies, including AI-enabled solutions. This role combines leadership, operational oversight, stakeholder management, and transformation initiatives, making it an excellent opportunity for an experienced technology leader looking to make a broad business impact within a complex international … approved AI, automation, and productivity-enhancing technologies. Work closely with regional leadership teams to understand evolving business needs and technology priorities. Participate in vendor management, contract negotiations, and budget planning activities. Support major incident management, problem management, and post-incident review processes. Champion best practice ...

Systems Development Manager, AWS Managed Operations

Location
Greater London, England, United Kingdom
Amazon Development Centre (London) Limited Are you ready to transform how support engineering operates at scale? Amazon Web Services (AWS) Operations Management is seeking an experienced Support Engineering Manager to lead a global team that's evolving from reactive troubleshooting to proactive automation and engineering excellence using Agentic Al. … grow their careers through strategic problem-solving, not just ticket resolution. You'll have the opportunity to create transformational impact across AWS by leading incident management at massive scale, championing automation-driven solutions, and mentoring engineers to think beyond immediate fixes. This isn't traditional support management ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands-on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement.We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve … resilient, and scalable platforms across Azure, AWS, or both.* Design and improve CI/CD pipelines, infrastructure as code, automated testing, release controls, environment management, and deployment strategies.* Work with software engineering and data science teams to productionize AI solutions and ensure services are ready to operate reliably ...

Cloud Operations Owner

Location
Greater London, England, United Kingdom
driving continuous improvement across the cloud estate. Key Responsibilities Include Acting as the primary operational contact and escalation point for AWS services. Leading major incident management and driving timely resolution with internal and external partners. Monitoring and managing service performance against SLAs, KPIs, and operational targets. Driving continual … service improvement initiatives and operational excellence programmes. Managing governance across incident, change, release, and transition processes. Conducting root cause analysis and implementing preventative measures. Overseeing AWS infrastructure operations across compute, storage, networking, and security services. Ensuring compliance with security, audit, and regulatory requirements. Working closely with Cloud, DevOps, Engineering ...

Manager, DE , TC, FS

Location
Greater London, England, United Kingdom
responsible for leading software engineering delivery across EY FSO client engagements. The role combines engineering leadership, solution design, cloud-native development, Agile delivery, stakeholder management and production readiness in regulated financial services environments. The opportunity EY FSO Technology Consulting is expanding its Digital Engineering capability to help financial services … close enough to the code to assure quality and delivery outcomes. Key responsibilities Lead end‐to‐end engineering workstreams across banking, insurance, wealth, asset management or capital markets clients, ensuring high‐quality delivery against scope, plan, risk and value outcomes. Design and guide implementation of scalable digital solutions using ...

Head of Machine Learning Engineering

Hiring Organisation
Appcast
Location
London, UK
conversational AI. Ensure your teams ship end-to-end - from prototyping through deployment, monitoring, and iteration. Raise operational maturity by defining SLIs/SLOs, incident management, and on-call practices for ML systems in production.Shape AI practices and governance. Co-own the operating model for modern … instincts - you know what it takes to run reliable ML at scale – including MLOps tooling and practicesExperience raising operational standards: SLIs/SLOs, monitoring, incident management, and production reliabilityBudget ownership, including infrastructure cost management and vendor oversightComfortable operating at both strategic and technical levels, with strong stakeholder ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
Cloud Infrastructure Engineer, your main focus will be to ensure our services remain highly available and performant. Key responsibilities include: Cloud Infrastructure Management Manage and support AWS infrastructure , focusing on scalability, security, and reliability. Handle deployments, managing CI/CD pipelines for both containerised (Docker/ECS) and serverless … data stores (Aurora MySQL, DynamoDB) Drive cost optimisation across the infrastructure — monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues ...

Data Centre Engineer - Service Desk

Location
City Of London, England, United Kingdom
Provide a value-added service to customers. MAIN DUTIES, INCLUDING BUT NOT RESTRICTED TO: Receive and act on all customer Incidents\\Events\\Request notifications Incident Management – Prioritises and diagnoses, investigates incidents according to agreed procedures. Incident Management – Facilitates recovery, following resolution of Incidents, documents and closes … resolved Incidents, Events and/or Requests via agreed procedures. Implementation of agreed IT Infrastructure changes and following maintenance routines Network Support – Uses Network Management tools where applicable to investigate and diagnose network problems, collect performance statistics, and create reports Customer Service - Liaise with customers regarding fault correction progress ...

Head of Machine Learning Engineering

Location
Greater London, England, United Kingdom
conversational AI. Ensure your teams ship end-to-end - from prototyping through deployment, monitoring, and iteration. Raise operational maturity by defining SLIs/SLOs, incident management, and on-call practices for ML systems in production. Shape AI practices and governance. Co-own the operating model for modern … know what it takes to run reliable ML at scale – including MLOps tooling and practices Experience raising operational standards: SLIs/SLOs, monitoring, incident management, and production reliability Budget ownership, including infrastructure cost management and vendor oversight Comfortable operating at both strategic and technical levels, with strong ...

Data Centre Operations Manager, LHR Data Center Operations (DCO)

Hiring Organisation
Appcast
Location
London, UK
CentersConstantly improving all our processes and procedures. We believe there is nothing we cannot improveAssisting & managing relationships with external vendors & contractorsLiaising with internal teams & management groupsEnsuring we adhere to and exceed local Health & Safety standards in all our Data CentersCreating and maintaining metrics on all aspects of our Data … Centers and utilizing those metrics to drive positive changesAssisting in implementing service methodologies including incident management, problem management, change management and capacity managementA day in the lifeAbout AWSDiverse ExperiencesAWS values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed ...

Director, Incident Management & Service Continuity

Location
Greater London, England, United Kingdom
Hines is seeking a Director - REPAC to lead a strategic incident management framework tailored to JPMC’s needs, ensuring high-quality, compliant, and efficient services while mitigating operational risks. You will steer service evolution and align delivery with SLAs and governance. You will supervise REPAC managers and staff … foster accountability and high performance, and collaborate across functions to drive consistency and excellence in incident response and resilience. #J-18808-Ljbffr ...

Senior software engineer (Node.js/TypeScript)

Location
Greater London, England, United Kingdom
operate in a build-and-run model, with teams responsible for the full software lifecycle, from architecture and delivery through to operational support and incident management. Our core platform is built on AWS serverless technologies including Lambda, SQS, EventBridge, API Gateway, S3, and ECS. We primarily develop in TypeScript … Infrastructure is managed through a combination of Terraform and Serverless Framework. We use GitHub Actions for CI/CD and incident.io to support our incident management processes and operational workflows. We care deeply about engineering quality, pragmatic architecture, and continuous improvement. If you’d like to go deeper ...

Technical Lead

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
Jenkins/GitHub)Configuring and running Code/Binary scans using solutions like SonarQube, Semgrep, Blackbuck, Trivy, GitLeaks Veracode, etc. Configuring and using Secrets management tools like Vault and Cloud native solutionsBroad knowledge of SDLC Tools, specifically Build, Test and Deploy Automation tools, e.g., Maven, Gradle, Selenium, Ansible, etc. … supported systems is continually updated. To provide a high level of customer service, whilst working under pressure. To follow and adhere to established Incident Management, Change Management and Problem Management procedures. LSEG is a leading global financial markets infrastructure and data provider. Our purpose is driving ...

service performance lead

Location
Greater London, England, United Kingdom
drive service performance across internal teams and external vendors, ensuring services meet the needs of baristas and customers. You will focus on service stability, incident management, and continuous improvement, while supporting the transition of new and enhanced services into operations. You will play a key role inanalysingservice performance … driving improvements that enhance the service desk experience and overall technology performance. We’ll look to you to bring your strong experience in service management disciplines (incident, problem, change and vendor management), ideally within a retail or hospitality environment. You’ll be able to apply your ITIL ...

service performance lead

Hiring Organisation
Starbucks Corporation
Location
London, UK
Employment Type
Full-time
drive service performance across internal teams and external vendors, ensuring services meet the needs of baristas and customers. You will focus on service stability, incident management, and continuous improvement, while supporting the transition of new and enhanced services into operations. You will play a key role in analysing … driving improvements that enhance the service desk experience and overall technology performance. We'll look to you to bring your strong experience in service management disciplines (incident, problem, change and vendor management), ideally within a retail or hospitality environment. You'll be able to apply your ITIL ...

Control Testing Specialist

Hiring Organisation
SWIFT
Location
London, UK
Employment Type
Full-time
business-relevant insights for stakeholders and senior management. Key ResponsibilitiesControl Testing ExecutionPlan and execute first-line control testing aligned with regulatory requirements (ICT risk management, incident management, resilience, third-party risk) Assess design and operating effectiveness of key controls supporting regulatory compliance Perform walkthroughs, evidence reviews … document control gaps, weaknesses, and non-compliance against legal requirements Translate technical findings into clear, concise business language with actionable recommendations Produce high-quality management reports, working papers, and summaries on control effectiveness and testing outcomes Track remediation actions, validate closure, and support retesting activities Audit & Regulatory Readiness (First ...