76 to 100 of 1,750 Root Cause Analysis Jobs in the UK

Software Engineer III - Python

Location
Glasgow, Scotland, United Kingdom
management Improve operability of services by adding and using observability tooling including logs, metrics, traces, dashboards, and alerts, and participate in incident response and root-cause analysis Leverage enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g. … secure handling of inputs/outputs, and adherence to resiliency and security expectations Preferred skills Strong debugging and troubleshooting skills in distributed systems, including rootcause analysis and performance bottleneck identification using logs, metrics, and traces Experience improving code quality and reliability through test strategy improvements, refactoring ...

Cloud Support Engineer

Location
Greater London, England, United Kingdom
stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight for IC1 and IC2 engineers Contribute to internal documentation, playbooks and root cause analysis reports, with a focus on accuracy and knowledge sharing Support both hosted and self-managed deployments of Vault, collaborating with … experience in technical support, cloud infrastructure or SRE/DevOps roles, ideally within high-scale, client-facing environments Deep experience with incident management, root cause analysis and production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience ...

Senior Communications Engineer

Location
Greater London, England, United Kingdom
operational incidents. Mentoring and supporting less experienced engineers. Undertaking advanced fault diagnosis across RF, fibre, IP and communications infrastructure. Conducting RF investigations including spectrum analysis, antenna testing, VSWR and Distance-to-Fault measurements. Supporting infrastructure upgrades, equipment commissioning and lifecycle replacement programmes. Planning and overseeing preventative maintenance. Conducting Root Cause Analysis and supporting Problem Management activities. Supporting the introduction and operational acceptance of new communications technologies. Maintaining engineering documentation, configuration records and technical procedures. Managing technical relationships with equipment manufacturers and specialist suppliers. Working closely with airport stakeholders, operational teams and external engineering partners. You will ...

Senior Full Stack Engineer - Managed Service

Location
Leeds, England, United Kingdom
Patching and package upgrades, Secret rotations, certificate renewals and infrastructure upgrades Contribute to the management and resolution of complex L2/L3 incidents Conduct root cause analysis and post-incident investigations Maintain and update runbooks and operational documentation Participate in an on-call rota, including … environment, including incident triage, technical investigation, remediation and post-incident review. Experience implementing and deploying production changes in business-critical environments. Strong troubleshooting and root-cause analysis skills across application, infrastructure and cloud platforms. Experience with monitoring and observability tooling (e.g. Azure Monitor, Application Insights, Grafana, Datadog ...

Senior Incident Response Consultant 2

Location
Oxford, England, United Kingdom
neutralize cyber threats. Specializing in industry-standard forensic tools and Sophos technologies, the team provides comprehensive investigations, response actions, remediation guidance, and root cause analysis to combat a wide range of cybersecurity incidents. As a Senior Incident Response Consultant on the Sophos DFIR team, you will … have been taken by both your team and the customer to effectively neutralize the threat. Additionally, you will be tasked with conducting a thorough root cause analysis to determine the origin of the incident, including identifying whether any data exfiltration occurred, provided the necessary evidence is available. ...

Senior Specialist Engineer - Networks

Hiring Organisation
UK Health Security Agency
Location
Birmingham, Leeds, Liverpool or London (Canary Wharf), E14 4PU, United Kingdom
Salary
£56185.00 to £70566.00
with security standards and evolving threat landscapes. Act as the senior network authority during Major Incident Response Team (MIRT) activities, leading diagnosis, resolution, communication, Root Cause Analysis (RCA) and Post Incident Reviews (PIR). Develop incident runbooks, continuity plans and mitigation strategies to reduce risk, improve service … major incidents, leading network diagnosis and resolution activities, collaborating within multi-disciplinary Major Incident Response Teams, and managing complex, high-pressure situations. Experience producing Root Cause Analysis (RCA) reports, Post-Incident Reviews (PIRs), and maintaining incident runbooks, playbooks, and escalation procedures while driving continuous service improvement ...

DevOps Manager

Hiring Organisation
Stott & May Professional Search Limited
Location
United Kingdom
Employment Type
Permanent
launches and releases, working closely with Development, Product, Testing, Security and IT teams to ensure platforms are ready for live service. Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. Manage monitoring and observability, using tools such … applications, microservices and high-traffic digital platforms. Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. Experience managing incidents, problem management, root cause analysis and live production environments. Understanding of application and cloud security, including OWASP principles, WAF, bot management and security scanning. Experience ...

Systems Administrator (Associate)

Location
Basingstoke, England, United Kingdom
team -You will support critical applications and ensure the stability of services by performing dedicated maintenance activities. You engage in automation activities, perform root cause analysis (RCA), and remediation. Knowledge of production support process including incident/change/problem management, call triaging, and critical issue resolution … DHCP Experience working with Microsoft Office (Word, Excel, Outlook, PowerPoint) and SharePoint Good understanding of Software Engineering concepts and methodologies Perform root cause analysis and track defect resolution to completion Team player who can work independently with general supervision Work closely with application teams in ensuring best ...

Manager, Operational Excellence Programs

Location
Erskine, Scotland, United Kingdom
within a global matrixed organization. Expertise in program and portfolio management disciplines, governance practices, and business transformation frameworks. Experience with Lean, Six Sigma, Kaizen, root cause analysis, and corrective action methodologies. Strong leadership, coaching, and talent development skills. Excellent written and verbal communication skills with the ability … degree in Business, Engineering, Operations, or related field. Additional Skills: Accountability Business Transformation Coaching and Mentoring Continuous Improvement Cross-Functional Leadership Customer Experience Data Analysis Decision Making Executive Communication Financial Management Global Team Leadership Managing Ambiguity Operational Excellence Organizational Change Management People Leadership Portfolio Management Process Governance Program Governance ...

Data Quality Analyst (global role – in a virtual working environment)

Hiring Organisation
Grant Thornton International Ltd
Location
United Kingdom
quality issues Produce regular reporting on data quality of member firms. Support ongoing training for member firms to continually improve data quality Issue Identification, Root Cause Analysis & Resolution Detect data defects and anomalies in member firm GRD data submissions Perform root cause analysis … monitoring and improving data quality Understanding of data quality controls, governance principles and data lifecycle management Strong Excel skills, including Pivot Tables and data analysis Ability to interpret data and communicate findings clearly to both technical and non-technical stakeholders Process improvement mindset with the ability to identify inefficiencies ...

Customer Experience (CX) Manager

Location
Greater London, England, United Kingdom
champion capability that makes improvement stick. Responsibilities: Lead and facilitate improvement activity and cross-functional ‘Tiger Teams’ end-to-end - from problem framing and root-cause analysis through design, delivery and implementation - turning long-standing pain points into visible, measurable improvements. Own hands-on delivery, not just … that keep delivery moving under uncertainty. Drive change management - stakeholder engagement, adoption and embedding new ways of working so improvements stick. Use data and analysis (e.g. Excel, Power BI) to size problems, prioritise, track progress and evidence impact. Make strong, proactive use of AI to accelerate analysis, delivery ...

C# Developer in Test

Location
Greater London, England, United Kingdom
Integration and Continuous Deployment (CI/CD) pipelines.* Mentorship & Backlog Support: Mentor other QEs on automation best practices and proactively support automation backlog efforts.* Root Cause Analysis: Put on your detective hat to investigate bugs, performing deep root-cause analysis and diving straight into ...

Endpoint Support Team Lead

Location
Cardiff, Wales, United Kingdom
deployment through Intune, ensuring robust testing, documentation and successful releases. Technical Escalation: Provide expert 2nd line support for endpoint and application issues, supporting root cause analysis, permanent fixes and proactive remediation. Security & Compliance: Support endpoint security, vulnerability remediation and compliance initiatives, ensuring agreed security and change management … support teams. Strong Microsoft Intune, Entra ID, Windows 10/11 and Windows Autopilot experience. Experience with application packaging and deployment . Strong troubleshooting, root cause analysis and escalation management skills. Experience working within Incident, Change and Problem Management processes and ITSM platforms such as Freshservice ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
Ansible Experience building and managing CI/CD pipelines using GitLab CI/CD, GitHub Actions, Jenkins or similar Experience with incident response, root cause analysis and driving improvements to the reliability and performance of production systems Commercial experience with Microsoft Azure or another major cloud platform … Linux-based production environments Build and improve containerised environments using Docker and Kubernetes Diagnose complex infrastructure, networking and performance issues, supporting incident response and root cause analysis Develop automation, Infrastructure as Code and CI/CD processes to improve engineering efficiency and reliability Implement and enhance monitoring ...

Senior Application Developer

Location
Basildon, England, United Kingdom
What you’ll do: Provide application development and production support for Numia‐owned systems and integrations Act as first‐line advisor for application incidents, analysis, and resolution during Italian business hours. Participate in out‐of‐hours (OOH) support for critical production issues as required Diagnose and resolve application defects … performance issues, and integration failures Perform root cause analysis (RCA) and implement corrective and preventive actions Collaborate closely with Numia stakeholders, internal development teams, and infrastructure partners Support deployments, patches, and minor enhancements to ensure platform stability Ensure incidents and changes are managed in line with agreed ...

Senior Systems Engineer

Location
Harwell, England, United Kingdom
demonstrate system performance against requirements.Identify and manage technical risks early in the development cycle.Ensure robust engineering practices are applied throughout product development.Problem Solving & Root Cause AnalysisLead investigation of complex system issues using structured troubleshooting methodologies.Perform system-level fault finding across software, firmware, electronics, optics, mechanics, and instrument control … systems.Drive root cause analysis and implementation of corrective actions.Support customer escalations and sustaining engineering activities where required.Team DevelopmentProvide coaching and mentoring to engineers across disciplines.Contribute to the development of systems engineering best practices within the organisation.Support the technical development of project teams.Potentially undertake line management responsibilities depending ...

Senior Software Engineers (ADA 95)

Location
Portsmouth, England, United Kingdom
locations for integration, testing and deployment activities. Responsibilities Develop, modify and maintain software applications supporting test & simulation systems. Lead or undertake software defect investigation, root-cause analysis and fault resolution. Analyse change requirements and produce engineering assessments, effort estimates and technical recommendations. Design and implement approved software … using C++ and/or C# . Ability to understand, modify and review existing and long-life software baselines. Structured software defect investigation and root-cause analysis. Code-review and peer-review skills. Unit, integration and regression testing. Verification of software fixes and production of objective test evidence. ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
improve SLIs, SLOs, and reliability metrics Proactively identify and resolve performance, availability, and reliability issues Lead and contribute to incident response, troubleshooting, and root cause analysis Automate operational processes and eliminate repetitive manual tasks Work closely with software engineers to improve deployment processes, system reliability, and developer … logging, tracing, alerting, and system health Experience troubleshooting complex production environments Understanding of SLIs, SLOs, SLAs, and error budgets Experience with incident management and root cause analysis Good understanding of cloud networking, security, and infrastructure fundamentals Strong scripting/automation skills A strong understanding of reliability, scalability ...

DevOps Team Manager

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
supportability thinking; involve senior engineers and technical leadership where specialist authority is required. Own platform SLAs/SLOs, service health, major-incident command, root-cause analysis, disaster recovery, business continuity testing and continuous service improvement. Operational Controls, Standards & Compliance Take day-to-day ownership within the DevOps … reviewed. Set the strategy and guardrails for AI-assisted and agentic DevOps delivery across CI/CD, Infrastructure-as-Code generation, incident triage, log analysis, runbook automation and documentation. Use and govern AI engineering tools such as Claude, GitHub Copilot or Codex, with clear quality, security, data-protection ...

Network & Infrastructure Tooling / Automation Specialist/ Architect - freelance - hybrid, London, UK

Location
Greater London, England, United Kingdom
unified operational and observability platforms. Enable traffic engineering, QoS, and policy enforcement through code-driven workflows. Observability, Operations & Resilience Develop automation for fault detection, root cause analysis, and remediation. Integrate telemetry, logs, and metrics into observability platforms. Support SRE-style practices including error budgets, reliability metrics … OSPF, EIGRP, Hybrid WAN, SD-WAN, MPLS, Traffic engineering, QoS, ExpressRoute, Direct Connect Observability & Operations : Observability platforms, Telemetry, Logs, Metrics, Fault detection automation, Root cause analysis, Remediation automation, SRE practices, Reliability metrics Platform, Cloud & Service Integration: IPAM integration, CMDB integration, Platform engineering, DevOps, Lifecycle management, Cloud networking ...

Lead Network Engineer - Secure Infrastructure

Location
Salisbury, England, United Kingdom
adoption of automation using tools such as Ansible and Terraform. Supporting operational requirements within a fast-paced, secure environment. Investigating complex issues, producing root cause analysis, and implementing solutions. Proactively monitoring and maintaining production systems while managing escalations. Mentoring colleagues and contributing to service improvement and innovation. … such as BGP and network segmentation using VRFs. CCNA certification with CCNP achieved or in progress. Ability to produce technical documentation including designs and root cause analysis reports. It would be great if you had: Exposure to cloud networking principles in Azure or AWS. Experience working within ...

IT Project Manager

Location
Biggleswade, England, United Kingdom
support model, KPIs).* Provide Level 3 functional/technical expertise for complex incidents/major problems (Run remains the ticket owner).* Support root cause analysis and corrective action plans with Run teams and vendors to prevent recurrence.**Business analysis support for small/medium … clean operational handover and limited disruption to plants.* High stakeholder satisfaction and measurable adoption of delivered solutions.* Reduced recurrence of major issues through effective root cause analysis and corrective actions.**Your Profile*** 6/8 years of IT project management experience (with an industrial background)* Proven ...

Insight Partner

Location
York and North Yorkshire, England, United Kingdom
usable and aligned to stakeholder requirements. Using Power BI to create compelling, high-quality visualisations and self-service reporting products. Diagnosing business issues through root cause analysis, performance investigation and structured analytical thinking. Developing scenario models, forecasts and expected outcomes to support strategic and operational decisions. Translating … reporting solutions. Ability to design, build, test and iterate reports to ensure accuracy, quality and practical business use. Strong business problem-solving skills, including root cause analysis and structured analytical thinking. Commercial awareness and the ability to link insight to business outcomes, benefits and performance improvement. Experience ...

Data Engineer

Location
Greater London, England, United Kingdom
multiple systems. The successful candidate will combine strong hands-on SQL and data engineering expertise with the ability to investigate complex data issues, identify root causes and work across technical teams to implement sustainable solutions. Key Responsibilities Design, build and optimise ETL/ELT data pipelines across multiple data … optimise complex SQL queries , including joins, window functions and performance tuning Translate business requirements into reliable, auditable datasets Perform data mining, reconciliation and root-cause analysis across complex data sources Map end-to-end data lineage and system flows , from source and ingestion through transformation and serving ...

Software Engineer Lead - Site Reliability

Location
Telford, England, United Kingdom
failover, degradation handling and recovery testing. Embed reliability, security, performance and operational-readiness expectations into engineering delivery. Support major incidents, post-incident reviews and root-cause analysis, ensuring improvement actions are owned and delivered. Use DORA, availability, recovery and operational metrics to identify risks and evidence progress. … expertise in Terraform, GitHub, GitHub Actions, CI/CD pipeline design, deployment automation and release support. Practical knowledge of incident management, problem management, root-cause analysis and operational readiness. Good awareness of Site Reliability Engineering principles, with the ability to apply them pragmatically to DevOps and platform ...