26 to 50 of 747 Root Cause Analysis Jobs in London

Senior Business Systems Analyst

Location
Greater London, England, United Kingdom
System Integration: Oversee end-to-end supply chain flows and maintain smooth integration between Oracle EBS and downstream platforms (Anaplan, SAP, ASCP).* Escalation & Root Cause Analysis: Act as the primary escalation point for critical ERP issues, conducting deep root cause analysis ...

Production Engineer

Location
City Of London, England, United Kingdom
oversight of member user administration, OMS integration support, and trade lifecycle issues for both the MTF platform and trading desk, with a focus on root-cause analysis and preventative improvements.The successful candidate should possess a positive 'can-do' attitude and an intuitively high level of customer service … skills, along with a proven ability to lead and mentor.Role ResponsibilitiesAct as technical lead and primary escalation point for complex application support issues, driving root-cause analysis and resolution across the trading platformLead and mentor junior engineers, fostering technical growth and ensuring knowledge transfer within the teamDesign ...

Vice President, Identity and Access Management

Location
Greater London, England, United Kingdom
approvals, MFA recovery flows, identity and access status checks) to reduce Service Desk dependency. Implement a continuous improvement loop: analyze top ticket drivers, remove root causes, standardize processes, improve knowledge, and automate recurring issues. Own operational risk posture for IAM services including access outages, mis-provisioning, privileged drift, toxic … telemetry for IAM services and integrations, and partner with SecOps where needed (SIEM, logging, anomaly detection). Drive reduction in repeat incidents through disciplined root cause analysis, prevention, and engineering partnership. Build strong partnerships across Security, Infrastructure, HR, application owners, and enterprise service management teams. WORK EXPERIENCE ...

Lead SRE - Chase UK

Location
Greater London, England, United Kingdom
planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start. Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation. Continuously develop AI skills relevant … across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within ...

Cloud Support Engineer

Location
Greater London, England, United Kingdom
stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight for IC1 and IC2 engineers Contribute to internal documentation, playbooks and root cause analysis reports, with a focus on accuracy and knowledge sharing Support both hosted and self-managed deployments of Vault, collaborating with … experience in technical support, cloud infrastructure or SRE/DevOps roles, ideally within high-scale, client-facing environments Deep experience with incident management, root cause analysis and production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience ...

Senior Specialist Engineer - Networks

Hiring Organisation
UK Health Security Agency
Location
Birmingham, Leeds, Liverpool or London (Canary Wharf), E14 4PU, United Kingdom
Salary
£56185.00 to £70566.00
with security standards and evolving threat landscapes. Act as the senior network authority during Major Incident Response Team (MIRT) activities, leading diagnosis, resolution, communication, Root Cause Analysis (RCA) and Post Incident Reviews (PIR). Develop incident runbooks, continuity plans and mitigation strategies to reduce risk, improve service … major incidents, leading network diagnosis and resolution activities, collaborating within multi-disciplinary Major Incident Response Teams, and managing complex, high-pressure situations. Experience producing Root Cause Analysis (RCA) reports, Post-Incident Reviews (PIRs), and maintaining incident runbooks, playbooks, and escalation procedures while driving continuous service improvement ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
Ansible Experience building and managing CI/CD pipelines using GitLab CI/CD, GitHub Actions, Jenkins or similar Experience with incident response, root cause analysis and driving improvements to the reliability and performance of production systems Commercial experience with Microsoft Azure or another major cloud platform … Linux-based production environments Build and improve containerised environments using Docker and Kubernetes Diagnose complex infrastructure, networking and performance issues, supporting incident response and root cause analysis Develop automation, Infrastructure as Code and CI/CD processes to improve engineering efficiency and reliability Implement and enhance monitoring ...

DevOps Manager

Location
London, United Kingdom
launches and releases, working closely with Development, Product, Testing, Security and IT teams to ensure platforms are ready for live service. Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. Manage monitoring and observability, using tools such … applications, microservices and high-traffic digital platforms. Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. Experience managing incidents, problem management, root cause analysis and live production environments. Understanding of application and cloud security, including OWASP principles, WAF, bot management and security scanning. Experience ...

Customer Experience (CX) Manager

Location
Greater London, England, United Kingdom
champion capability that makes improvement stick. Responsibilities: Lead and facilitate improvement activity and cross-functional ‘Tiger Teams’ end-to-end - from problem framing and root-cause analysis through design, delivery and implementation - turning long-standing pain points into visible, measurable improvements. Own hands-on delivery, not just … that keep delivery moving under uncertainty. Drive change management - stakeholder engagement, adoption and embedding new ways of working so improvements stick. Use data and analysis (e.g. Excel, Power BI) to size problems, prioritise, track progress and evidence impact. Make strong, proactive use of AI to accelerate analysis, delivery ...

Senior Communications Engineer

Location
Greater London, England, United Kingdom
operational incidents. Mentoring and supporting less experienced engineers. Undertaking advanced fault diagnosis across RF, fibre, IP and communications infrastructure. Conducting RF investigations including spectrum analysis, antenna testing, VSWR and Distance-to-Fault measurements. Supporting infrastructure upgrades, equipment commissioning and lifecycle replacement programmes. Planning and overseeing preventative maintenance. Conducting Root Cause Analysis and supporting Problem Management activities. Supporting the introduction and operational acceptance of new communications technologies. Maintaining engineering documentation, configuration records and technical procedures. Managing technical relationships with equipment manufacturers and specialist suppliers. Working closely with airport stakeholders, operational teams and external engineering partners. You will ...

Senior Security Incident Response Analyst

Location
Greater London, England, United Kingdom
through the on-call rota . Conduct forensic investigation and evidence collection across endpoint, identity, cloud, email and network technologies. Produce clear investigation timelines, root cause analysis and post-incident reports for technical and business stakeholders. Work with Threat Intelligence and Detection Engineering teams to apply knowledge … attacker tactics, techniques and procedures, with experience applying this knowledge to investigations or threat hunting. Experience leading technical investigations, building incident timelines and completing root cause analysis and post-incident reporting. Strong analytical and problem-solving skills, with evidence of making sound decisions and coordinating activity during ...

DevOps Team Manager

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
supportability thinking; involve senior engineers and technical leadership where specialist authority is required. Own platform SLAs/SLOs, service health, major-incident command, root-cause analysis, disaster recovery, business continuity testing and continuous service improvement. Operational Controls, Standards & Compliance Take day-to-day ownership within the DevOps … reviewed. Set the strategy and guardrails for AI-assisted and agentic DevOps delivery across CI/CD, Infrastructure-as-Code generation, incident triage, log analysis, runbook automation and documentation. Use and govern AI engineering tools such as Claude, GitHub Copilot or Codex, with clear quality, security, data-protection ...

C# Developer in Test

Location
Greater London, England, United Kingdom
Integration and Continuous Deployment (CI/CD) pipelines.* Mentorship & Backlog Support: Mentor other QEs on automation best practices and proactively support automation backlog efforts.* Root Cause Analysis: Put on your detective hat to investigate bugs, performing deep root-cause analysis and diving straight into ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
improve SLIs, SLOs, and reliability metrics Proactively identify and resolve performance, availability, and reliability issues Lead and contribute to incident response, troubleshooting, and root cause analysis Automate operational processes and eliminate repetitive manual tasks Work closely with software engineers to improve deployment processes, system reliability, and developer … logging, tracing, alerting, and system health Experience troubleshooting complex production environments Understanding of SLIs, SLOs, SLAs, and error budgets Experience with incident management and root cause analysis Good understanding of cloud networking, security, and infrastructure fundamentals Strong scripting/automation skills A strong understanding of reliability, scalability ...

Network & Infrastructure Tooling / Automation Specialist/ Architect - freelance - hybrid, London, UK

Location
Greater London, England, United Kingdom
unified operational and observability platforms. Enable traffic engineering, QoS, and policy enforcement through code-driven workflows. Observability, Operations & Resilience Develop automation for fault detection, root cause analysis, and remediation. Integrate telemetry, logs, and metrics into observability platforms. Support SRE-style practices including error budgets, reliability metrics … OSPF, EIGRP, Hybrid WAN, SD-WAN, MPLS, Traffic engineering, QoS, ExpressRoute, Direct Connect Observability & Operations : Observability platforms, Telemetry, Logs, Metrics, Fault detection automation, Root cause analysis, Remediation automation, SRE practices, Reliability metrics Platform, Cloud & Service Integration: IPAM integration, CMDB integration, Platform engineering, DevOps, Lifecycle management, Cloud networking ...

Data Engineer

Location
Greater London, England, United Kingdom
multiple systems. The successful candidate will combine strong hands-on SQL and data engineering expertise with the ability to investigate complex data issues, identify root causes and work across technical teams to implement sustainable solutions. Key Responsibilities Design, build and optimise ETL/ELT data pipelines across multiple data … optimise complex SQL queries , including joins, window functions and performance tuning Translate business requirements into reliable, auditable datasets Perform data mining, reconciliation and root-cause analysis across complex data sources Map end-to-end data lineage and system flows , from source and ingestion through transformation and serving ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions Ltd
Location
London, United Kingdom
Employment Type
Full-Time
Salary
£780.00 - £830.00 per day
application and cloud teams Lead technical investigations - resolve platform incidents, event processing failures and integration problems, providing SME support during major incidents and driving root cause analysis Build integrations and automation - connect monitoring with ServiceNow, ITSM, reporting and automation tooling via gateways, APIs and event forwarding … administration skills and deep understanding of event processing, correlation and alert lifecycle management Experience integrating monitoring with ServiceNow and automation platforms Excellent troubleshooting and root cause analysis skills, within an ITIL-aligned Incident, Problem and Change environment Confident communication and stakeholder management with both technical ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£780 - £830 per day
application and cloud teams Lead technical investigations - resolve platform incidents, event processing failures and integration problems, providing SME support during major incidents and driving root cause analysis Build integrations and automation - connect monitoring with ServiceNow, ITSM, reporting and automation tooling via gateways, APIs and event forwarding … administration skills and deep understanding of event processing, correlation and alert lifecycle management Experience integrating monitoring with ServiceNow and automation platforms Excellent troubleshooting and root cause analysis skills, within an ITIL-aligned Incident, Problem and Change environment Confident communication and stakeholder management with both technical ...

Technology Support Director, Risk Engagement

Location
Greater London, England, United Kingdom
appropriate handling of sensitive data. Establishes governance standards for AI‐assisted workflows used in incident/problem/change processes (including documentation and trend analysis), ensuring traceability/auditability and alignment to resiliency and security expectations. Advises on complex, cross-domain incident, problem and change management challenges, identifying process … opportunities to improve operational resilience across Infrastructure Platforms. Uses observability, monitoring and diagnostic information to analyse significant, recurring or cross‐domain issues, test root‐cause hypotheses and help accountable teams define sustainable remediation. Leads focused technical investigations into significant, recurring or cross‐domain issues, using support engineering expertise ...

Marketing Experience - Senior Associate

Location
Greater London, England, United Kingdom
develop team leads and front‐line staff, including performance management, workforce planning, and succession readiness Drive continuous improvement initiatives using structured methods such as root‐cause analysis, translating findings into sustained process controls and measurable benefits Establish and maintain operational controls, including documented procedures, control testing … experience managing operational risk, controls, incident response, and business continuity in a regulated environment Demonstrated experience driving continuous improvement using structured methodologies such as root‐cause analysis with measurable outcomes Experience performing product ownership activities, including managing backlogs, defining user stories and acceptance criteria, and partnering with ...

Marketing Experience - Senior Associate

Location
London, United Kingdom
develop team leads and front-line staff, including performance management, workforce planning, and succession readiness Drive continuous improvement initiatives using structured methods such as root-cause analysis, translating findings into sustained process controls and measurable benefits Establish and maintain operational controls, including documented procedures, control testing … experience managing operational risk, controls, incident response, and business continuity in a regulated environment Demonstrated experience driving continuous improvement using structured methodologies such as root-cause analysis with measurable outcomes Experience performing product ownership activities, including managing backlogs, defining user stories and acceptance criteria, and partnering with ...

Marketing Experience - Senior Associate

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
develop team leads and front-line staff, including performance management, workforce planning, and succession readiness Drive continuous improvement initiatives using structured methods such as root-cause analysis, translating findings into sustained process controls and measurable benefits Establish and maintain operational controls, including documented procedures, control testing … experience managing operational risk, controls, incident response, and business continuity in a regulated environment Demonstrated experience driving continuous improvement using structured methodologies such as root-cause analysis with measurable outcomes Experience performing product ownership activities, including managing backlogs, defining user stories and acceptance criteria, and partnering with ...

Head of Complaints

Location
Greater London, England, United Kingdom
automation or technology-led process change Added bonus: experience redesigning an operating model for greater productivity and capability Added bonus: track record of building root-cause analysis partnerships with Product or Customer Experience teams Right to work in the country of choice for working abroad Core Competencies … change, build high-performing teams, and leverage data for operational improvements. Highest-signal resume keywords Complaints-Function Leadership Experience FCA DISP Compliance Operational Data Analysis Stakeholder Influence AI-Enabled Transformation Hard Skills Regulatory Reporting Quality Frameworks Root-Cause Analysis Capacity Planning Change Management Soft Skills Clear ...

Privacy Analyst II

Location
Greater London, England, United Kingdom
diligence, working with InfoSec, Legal, and Procurement to assess vendor posture and manage subprocessor obligations. Lead on data privacy incident response, including logging, investigation, root cause analysis, remediation, and regulatory and client reporting. Develop and deliver training and knowledge content on global data protection. Create and maintain … privacy mailbox, including handling rights requests. What you’ll bring to the party Proven experience in the following: Core requirement Data incident management, incl. root cause analysis, remediation, logging and reporting. Significant operational experience in DPIAs/PIAs, third party vendor management, and risk management. Creating ...

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
criteria, and prioritized backlog items across the AIOps product lifecycle. Ownership of high-impact AIOps capabilities, including noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, and self-healing remediation. Partner across Operations, infrastructure, SRE, platform engineering, service management, application teams, security, and vendors … product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, and automated remediation. Experience defining product roadmaps, managing requirements and backlog priorities, and delivering technical platform capabilities ...