251 to 275 of 1,737 Root Cause Analysis Jobs in the UK

3rd Line Network Engineer

Hiring Organisation
IMT Resourcing Solutions
Location
London, Bishopsgate, United Kingdom
Employment Type
Permanent
Salary
£45000 - £50000/annum
established Monitor network events with the potential to impact multiple customers and correlate these against incoming incidents Lead problem management for recurring faults, identifying root causes and driving permanent fixes Produce clear Root Cause Analysis (RCA) and Reason for Outage documentation for customer-facing incidents Develop … Line Network Engineer or advanced 2nd Line Network Engineer who enjoys getting into complex technical problems, taking ownership of incidents and finding the root cause rather than applying temporary fixes. Apply now to take on a senior technical role supporting business-critical network services across London. ...

Application Support Analyst

Location
Bradford, England, United Kingdom
Discover Technology & Change at Vanquis: Job Board | Dayforce Jobs How you’ll make an impact: Own Level 2 incidents from diagnosis through resolution and root cause analysis. Business‐critical services recover quickly and repeat issues are reduced. Troubleshoot issues across Salesforce, MuleSoft APIs, GoAnywhere MFT workflows and Genesys … with clear technical evidence. Use logs, transaction traces, API payloads and Sumo Logic dashboards to investigate failures and performance issues. The team can identify root causes across application, integration and data layers. Manage incidents, problems, changes, defects and technical backlog items through ITSM tooling and Azure DevOps. Work ...

Technical Customer Engineer

Location
United Kingdom
stale. Troubleshoot and resolve complex technical issues across Meraki, Fortinet, HP Aruba, Cisco WLC, and Zscaler environments and more to come – digging into root cause, not just symptoms. Use every resolved ticket as a learning artifact: update the knowledge base, flag recurring patterns, and turn individual fixes into … looking for Must have 3 years minimum in a support role CP/IP fundamentals: OSI, routing, DNS, NAT VPN Troubleshooting (IPSEC, GRE) Root cause analysis (RCA) methodology Ticketing systems and structured documentation (ITIL Standard) Technical onboarding and customer training Fluent or native English (client‐facing level ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
incident response for outages and model-quality regressions Participate in an on-call rotation, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-ups Identify recurring operational issues and automate remediation to improve platform stability and developer experience Build and maintain multi … methods Hands-on experience with AWS and Terraform for infrastructure delivery and lifecycle management Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns Practical knowledge of observability and instrumentation across metrics, logs, and traces Comfort with on-call operations ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
incident response for outages and model-quality regressions Participate in an on-call rotation, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-ups Identify recurring operational issues and automate remediation to improve platform stability and developer experience Build and maintain multi … methods Hands-on experience with AWS and Terraform for infrastructure delivery and lifecycle management Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns Practical knowledge of observability and instrumentation across metrics, logs, and traces Comfort with on-call operations ...

Data Analyst

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
Leicester, Leicestershire, United Kingdom
Employment Type
Full-Time
Salary
£350.00 - £400.00 per day
provides reporting, investigations, and insight to stakeholders across multiple functions. This contract opportunity offers exposure to a fast-paced, data-rich environment where your analysis will have a direct impact. The Role and Deliverables Analyse operational, stock, warehouse, logistics, and finance data to support business performance. Investigate stock variances … inventory discrepancies, promotional performance, and banking data issues. Perform stock reconciliation and root cause analysis to identify and explain operational inconsistencies. Develop and maintain Power BI dashboards and reports to communicate findings to stakeholders. Support KPI reporting across supply chain, deliveries, and operational performance. Respond ...

Implementation & Technical Account Manager

Location
Greater London, England, United Kingdom
peak assessment windows, coordinating internal teams and providing real-time monitoring during live events. Lead post-event and post-implementation reviews, including incident/root-cause analysis and recommendations for improvement. Incident, Problem & Service Management Take ownership of critical and high-impact incidents, coordinating Support and Development … access management, data integrations/imports and browser/client troubleshooting Skills: Excellent communication across technical and non-technical audiences; resilience and structured root-cause problem solving under incident pressure Eligibility: Eligible to work in the UK and able to satisfy BPSS-aligned pre-employment screening (identity verification ...

Senior CNS Engineer

Location
Greater London, England, United Kingdom
this role. You will leverage AI-assisted tools, intelligent data sanitisation utilities, and modern analytics to accelerate incident triage, automate log de-sensitisation, enhance root-cause analysis, and drive operational efficiency. As a senior member of the team, you will also provide vital support and mentoring … assisted workflows to enable secure information transfer to global Cisco Support teams. Act as a technical coordinator during critical escalations, producing comprehensive incident reports, root-cause analyses, and stakeholder updates. Configure, diagnose, and troubleshoot technical issues logged by CNS customers across multi-vendor and Cisco enterprise environments. Issue ...

AI Data Scientist - Autonomous Network

Location
United Kingdom
models for network performance, service quality, fault behaviour, customer impact, capacity, and resilience.* Build statistical and machine learning models for anomaly detection, fault prediction, root-cause analysis, degradation detection, and proactive assurance.* Develop data aggregation, cleansing, enrichment, and feature engineering pipelines for network telemetry and OSS data. …/MPLS, SD-WAN, fixed, and cloud network KPIs.* Support AIOps use cases such as alarm reduction, incident prioritisation, predictive maintenance, and automated root-cause analysis.* Work with OSS and inventory teams to align data models with TMF SID concepts and TMF Open API structures.* Use BigQuery ...

CIS Engineer

Location
Greater London, England, United Kingdom
compliance requirements. Monitor infrastructure health, identify risks, and proactively address potential issues before they impact service availability. Participate in incident management activities, including root cause analysis and implementation of preventative measures. Work within strict change control processes to plan, document, test, and implement infrastructure changes. Support infrastructure … secure, regulated or mission‐critical environments. Proven experience delivering infrastructure patching programmes and vulnerability remediation activities. Strong background in infrastructure fault finding, troubleshooting, and root cause analysis. Experience working within formal change management and incident management processes. Experience supporting enterprise‐scale infrastructure environments. Personal Attributes Analytical and methodical ...

SIAM Problem Analyst: Root Cause & Service Stability

Location
United Kingdom
Capgemini is seeking a SIAM Problem Analyst to own end-to-end problem management across a multi-supplier IT environment. You will drive root cause analysis with suppliers and internal teams, manage problem records, and deliver permanent fixes to reduce incidents and improve service stability. The role ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
tooling, and review pull requests Encode recurring infrastructure tasks as reusable internal skills for human and agent teammates Lead incident response, post-mortems, and root-cause analyses Own reliability, performance, and efficiency of core services end-to-end Define and uphold SLOs and error budgets and carry … keywords Infrastructure Engineering DevOps Practices AWS, GCP, Azure Python or Go Automation Helm and Terraform Hard Skills Infrastructure Engineering DevOps Automation Containerization Incident Response RootCause Analysis SLO Definition Monitoring and Logging System Design High‐Availability Systems Soft Skills Excellent Communication Collaboration Problem-Solving Ownership Accountability Industry ...

Electronics Engineer - Reliability & Supportability

Location
City of Edinburgh, Scotland, United Kingdom
understand failure mechanisms and improve product resilience Conduct reliability, maintainability and supportability assessments throughout the engineering lifecycle Develop FMEA/FMECA and failure analysis activities at component, board and equipment level Support diagnostics, fault-finding and testability strategies for complex electronic systems Use RAMT techniques to help shape engineering … decisions and optimise through-life support solutions Contribute to root cause investigations and continuous improvement activities across development and in-service products What you’ll bring Electronics engineering, hardware design, electronic test, systems engineering or product support engineering Understanding of electronic components, circuit design principles and schematic interpretation ...

Enterprise IT Architect

Location
Basildon, England, United Kingdom
senior technical contact for client, handling complex queries, escalations, and design discussions Lead major incident management, including deep‐dive diagnostics, cross‐team coordination, and rootcause analysis. Provide hands‐on support for critical production issues across online and batch processing Design and govern distributed‐to‐mainframe integrations, including … including authorisation, clearing, settlement, billing, and disputes Strong hands‐on expertise with mainframe technologies, including: CICS (online transaction processing). DB2 (schema design, performance analysis, and data troubleshooting) Proven experience designing and supporting API‐based and event‐driven integrations with mainframe systems Strong background in production support, complex incident ...

Senior Technical Specialist

Location
Livingston, Scotland, United Kingdom
expertise, mentor junior colleagues, share knowledge, influence stakeholders and provide clear updates to technical and non-technical audiences. Lead complex investigations, major incidents and root cause analysis, while supporting service requests, project delivery, platform upgrades, change activity and improvements aligned with business, operational and security needs. Shape … with a strong understanding of endpoint security, compliance and enterprise management practices. Experience leading technical change, service improvement initiatives, complex investigations, major incidents and root cause analysis. Strong fault-finding, diagnostic and problem-solving skills, supported by scripting or automation experience using PowerShell and/or APIs. Clear ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
/7 enterprise where uptime, performance and stability are critical. Practical experience using LLM platforms and coding assistants safely to improve productivity, quality and root-cause analysis. Additional Information Develop and maintain resilient tools, operational APIs and automation for effective system management. Use orchestration and scripting to remove … trace issues from the edge through to origin systems and coordinate effective remediation. Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence. Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows. Drive initiatives that improve reliability, observability, performance ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
/7 enterprise where uptime, performance and stability are critical. Practical experience using LLM platforms and coding assistants safely to improve productivity, quality and root-cause analysis. Additional Information Develop and maintain resilient tools, operational APIs and automation for effective system management. Use orchestration and scripting to remove … trace issues from the edge through to origin systems and coordinate effective remediation. Participate in live incident response, post-mortems and root-cause analysis to prevent recurrence. Maintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows. Drive initiatives that improve reliability, observability, performance ...

“Techno-Functional” Analyst– Capital Markets Data Transformation

Location
Greater London, England, United Kingdom
stakeholders. Contribute to and maintain the controls gap assessment, supporting prioritization and ensuring closure plans are realistic and tracked. Identify data sourcing gaps, coordinate root-cause discussions, and support agreement of closure plans with business and technology teams. Support delivery of manual controls automation waves, ensuring requirements, testing … front-to-back awareness is a plus). Proficiency in SQL (advanced querying, performance tuning, reconciliation logic). Strong proficiency in Python for data analysis and automation (pandas, validation frameworks, scripting). Experience supporting or validating ETL/ELT pipelines and data quality frameworks (rules, thresholds, exception handling). ...

Senior Manager, Data Center Facilities Engineering

Location
Greater London, England, United Kingdom
procedure documentation, rollback plans, stakeholder communications, and execution controls protect data center availability. The role leads incident response and post-incident reviews, directs root cause analysis and corrective and preventive action plans, validates commissioning and design deliverables, oversees site assessments and build activities, and establishes consistent operational … compliance with safe working practices, providing direction to uphold company safety standards. Oversees post-incident review processes by guiding cross‐functional teams through root cause analyses, institutionalizing lessons learned, and driving the development and governance of Corrective and Preventive Action Plans (CAPAs) to mitigate systemic risks and enhance ...

Quality Improvement Engineer - Contract

Location
West of England, England, United Kingdom
supply chain and supplier operations. This is a highly visible role where you’ll work closely with internal stakeholders and external suppliers to identify root causes, implement corrective actions, improve quality performance, and ensure compliance with ISO 9001 standards. What You'll Be Doing Drive Quality Excellence Maintain … with suppliers to investigate issues, drive improvements, and prevent recurrence. Monitor supplier performance and influence long-term quality improvements. Solve Complex Quality Challenges Lead root cause investigations using 8D, 5 Why, Fishbone and other structured problem-solving techniques. Implement effective containment, corrective and preventive actions. Ensure sustainable solutions ...

Senior Analyst - Microsoft 365

Location
Glasgow, Scotland, United Kingdom
senior technical escalation point for Microsoft Teams, Exchange Online, SharePoint Online, OneDrive, Microsoft Copilot, and related services, resolving complex incidents and problems. Lead root cause investigations and problem management activities, driving preventative actions and continuous service improvement. Define and maintain Microsoft 365 governance, technical standards, documentation, and operational … complex issues across Microsoft Teams, Exchange Online, SharePoint Online, OneDrive, Microsoft Copilot, and related technologies. A proactive problem‐solver with experience applying structured troubleshooting, root cause analysis, and continuous improvement practices to enhance service performance, stability, and user experience. An effective collaborator and communicator who can build ...

Quality Improvement Engineer - Contract

Location
Severn Beach, England, United Kingdom
supply chain and supplier operations. This is a highly visible role where you'll work closely with internal stakeholders and external suppliers to identify root causes, implement corrective actions, improve quality performance, and ensure compliance with ISO 9001 standards. What You'll Be Doing Drive Quality Excellence Maintain … with suppliers to investigate issues, drive improvements, and prevent recurrence. Monitor supplier performance and influence long-term quality improvements. Solve Complex Quality Challenges Lead root cause investigations using 8D, 5 Why, Fishbone and other structured problem-solving techniques. Implement effective containment, corrective and preventive actions. Ensure sustainable solutions ...

Technical Support Engineer

Location
Greater London, England, United Kingdom
site implementations. Ensure SLAs, uptime targets, and performance KPIs are consistently achieved. Provide mentorship and guidance to the junior team member. Support root cause analysis and produce post-incident reports where required. Logistics & Asset Management Oversight Oversee device procurement, stock control, and lifecycle planning. Supervise preparation, configuration … switches, and firewalls. Solid Windows OS knowledge (event logs, command line, system configuration). Experience with remote device management tools. Ability to perform structured root cause analysis. Understanding of endpoint security and device compliance best practices. Strong documentation and asset management discipline. Desirable Technical Knowledge ANPR systems ...

Head of Service Delivery New Solihull, England, United Kingdom

Location
Metropolitan Borough of Solihull, England, United Kingdom
raise service standards across the function. Major incident and escalation leadership: Own the major incident and executive escalation model, ensuring clear command, communication, root cause follow‐up and business impact management for high‐priority service issues. Global support and scalability: Shape and continuously improve a global support model … communication skills, with confidence influencing senior leaders, business partners, suppliers and cross‐functional teams. Data‐led service improvement mindset, using service metrics, user insight, root cause analysis and trend reporting to improve experience, reliability, cost and operational performance. Excellent judgement in ambiguous situations, balancing user experience, operational ...

Head of Service Delivery

Location
Metropolitan Borough of Solihull, England, United Kingdom
raise service standards across the function. Major incident and escalation leadership: Own the major incident and executive escalation model, ensuring clear command, communication, root cause follow-up and business impact management for high-priority service issues. Global support and scalability: Shape and continuously improve a global support model … communication skills, with confidence influencing senior leaders, business partners, suppliers and cross-functional teams. Data-led service improvement mindset, using service metrics, user insight, root cause analysis and trend reporting to improve experience, reliability, cost and operational performance. Excellent judgement in ambiguous situations, balancing user experience, operational ...