7 of 7 Incident Management Jobs in the City of Westminster

Head of Production Management- J.P. Morgan Personal Investing

Location
Westminster, West End, United Kingdom
investing, and the trust that 150 years of J.P. Morgan heritage brings. J.P. Morgan Personal Investing offers award-winning investments, products and digital wealth management services to over 275,000 investors in the UK. We built the business with innovation as a core part of our ethos to give … work in tribes and squads that focus on specific products and projects. About the role: J.P. Morgan Personal Investing is strengthening its production management capability to support a growing portfolio ofwealthproducts built on a common Investment Platform. We are seeking a Head of Production Management to lead ...

Operations and SRE Manager

Location
City of Westminster, England, United Kingdom
external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement., Lead the implementation of the team's strategic direction, translating priorities into clear operational plans, backlogs and deliverables. Line … manage and develop the team leaders, setting clear expectations, providing coaching, challenging unhelpful patterns and helping them build stronger delivery discipline. Ensure incident management, problem management, RCAs, post-mortems and improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation ...

Product Associate - SRE Team - Chase UK

Location
Westminster, West End, United Kingdom
services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas … reliability outcomes. Identify opportunities to improve platform reliability, reduce operational friction, and help teams build and run services confidently. Use data and feedback, including incident trends, service health metrics, adoption data, and stakeholder input, to guide decisions and measure progress. Create clear product materials, including problem statements, user stories ...

Software Engineer Lead - Site Reliability

Location
City of Westminster, England, United Kingdom
testing. Improve deployment safety, release readiness and operational readiness for customer-facing digital services. Apply SRE principles pragmatically to improve availability, recoverability, monitoring and incident learning. Strengthen monitoring, logging, tracing, alerting and service-health dashboards across digitally connected workloads. Reduce single points of failure and improve failover, degradation handling … recovery testing. Embed reliability, security, performance and operational-readiness expectations into engineering delivery. Support major incidents, post-incident reviews and root-cause analysis, ensuring improvement actions are owned and delivered. Use DORA, availability, recovery and operational metrics to identify risks and evidence progress. Enable practical AI adoption across ...

Cyber Incident & Crisis Lead — Remote

Location
City of Westminster, England, United Kingdom
Fosse-1 is seeking an experienced cyber incident and crisis management professional to help strengthen resilience and response capabilities. The role operates at both strategic and tactical levels to mature incident management, crisis response, and organisational resilience within a complex environment. Ideal candidates will lead major … incident activities, deliver resilience initiatives, and work with senior leadership during significant cyber events. #J-18808-Ljbffr ...

Customer Experience Engineering Manager

Location
City of Westminster, England, United Kingdom
complex issues but also invests in engineering practices such as daily scrums and triage to deeply understand platform gaps from customer insights and incident signals. Collaborate with Azure engineering teams using a prioritized set of opportunities to eliminate top issues impacting customer experience and improve Azure quality and security … continuously improve diagnostics and supportability. Lead operational excellence by reinforcing ACE accountability for complex cases, improving Time to Mitigate (TTM), and maturing the ACE Incident-Management function to ensure high-quality, engineering-driven problem resolution. Attract and build a diverse, high-performing team with the capabilities needed ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning and performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure … team, sharing learnings and best practices with the broader production engineering organizationContribute hands-on to technical work including code, system design reviews, and incident response, using AI tooling to expand personal and team reach across disciplinesPartner cross-functionally with software engineering, data science, and product teams to unblock dependencies ...