626 to 650 of 1,945 Root Cause Analysis Jobs in the UK

L2/L3 Support Engineer| Document Publishing| Stevenage, UK

Location
Stevenage, England, United Kingdom
diagnose and resolve incidents; provide workarounds where permanent fixes require change, and maintain clear communication with users and stakeholders throughout the ticket lifecycle. Perform root cause analysis for recurring/major incidents, document findings, maintain known error records and implement preventive actions through problem management. Support service … health checks, certificate/service account tracking, vulnerability remediation, patching coordination and license/renewal reporting where applicable. Contribute to service improvement through trend analysis, automation opportunities, backlog reduction, knowledge reuse and operational reporting. Required Skills/Experience: Application Management Services (AMS) experience in a managed services or enterprise ...

Senior Embedded Software Engineer

Location
Greater London, England, United Kingdom
help shape and evolve the overall software architecture, ensuring it remains scalable, maintainable and aligned with future product goals Drive debugging and root cause analysis - You’ll investigate complex issues across the hardware and software stack, identifying root causes and implementing robust, long-term fixes Support ...

Lead SRE - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start. Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation. Continuously develop AI skills relevant ...

Lead SRE - Chase UK

Location
Westminster, West End, United Kingdom
planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start. Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation. Continuously develop AI skills relevant ...

Senior Security Engineer, Security Incident Response Team (SIRT) - EMEA

Hiring Organisation
GitLab
Location
United Kingdom, UK
Employment Type
Full-time
enhance automation and AI-assisted workflows to improve triage, investigation speed, and response consistencyPartner with Threat Intelligence to contextualize threats and improve detection coverageConduct root cause analysis (RCA) and lead post-incident reviews to drive continuous improvement and risk reductionDevelop and maintain runbooks, playbooks, and operational documentationCollaborate … GitLab cloud and corporate environments. We are both reactive and proactive, leading security investigations, incident response support and response resolution, through to cyber threat analysis and detection and response engineering. Even though we're a global team, we work together in a cross-regional manner and have automation ...

Service Delivery Manager

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
purpose by overseeing service operations, coordinating support activities, and driving service excellence. Continuous Service Improvement Analyse incidents, trends, and service performance to identify root causes and opportunities for improvement. Work collaboratively with delivery and product teams to implement enhancements that improve user experience, service resilience, and operational efficiency. Customer … Problem Management - identify trends and patterns relating to repeating service problems, and develop and enact plans to resolve these. Sponsor and when appropriate undertake Root Cause Analysis (RCA) of problems, informing stakeholders of findings and driving adoption of new measures required to solve them. Service Design - support ...

devops engineer infrastructure platform

Location
Greater London, England, United Kingdom
design reviews, answer questions, and share knowledge through guides, playbooks, and internal tech talks; Serve as a first responder during on-call rotations, lead root cause analysis, and drive incident mitigation and follow-through; Drive reliability, operational readiness, and developer experience objectives for the Infrastructure platform; Support … CNCF ecosystem. Условия: Benefits include healthcare, well-being support, parental leave, pensions, and generous annual leave, including time off for a charitable cause; Benefits are country-specific. #J-18808-Ljbffr ...

Operations Delivery Lead

Hiring Organisation
Hackajob Ltd
Location
Farnborough, Hampshire, South East, United Kingdom
Employment Type
Permanent
Personnel Security policies, including Personnel Reliability Framework (PRF) and relevant ITAR/export control regulations. Lead incident management, providing accurate status reporting and ensuring root cause analysis is completed. Team Leadership & Development Lead, develop, and motivate a multi-disciplined team of engineers across Wintel, Linux, Virtualisation … Service Improvement (CSI) Champion a culture of continual improvement across all infrastructure and platform services. Identify, document, and implement improvement initiatives using data-driven analysis and operational insight. Maintain and prioritise a CSI register, tracking progress and demonstrating measurable value to the client. Leverage automation, self-healing, and proactive ...

High Level Technician

Location
Wimbledon, England, United Kingdom
maintenance challenges. Monitor warranty issues and collaborate with supplier technical teams. Attend to faulty trains in service, conducting component replacements and modifications. Support root cause analysis using methodologies like Failure Reporting, Analysis, and Corrective Action System (FRACAS), 8D (a problem-solving methodology), and the Production Part ...

Technical Account Manager (Managed Services)

Location
Peterborough, England, United Kingdom
technical governance across Incident, Problem, and Change management. Act as the technical escalation point for major incidents and complex problems. Lead and contribute to root cause analysis (RCA) and post-incident reviews. Monitor estate health, performance, and capacity, driving proactive improvements. Ensure changes are technically assessed, risk … problem technical resolution. Strong ITSM toolset capability (ServiceNow). Ability to translate technical issues into business‐relevant insight. Highly competent in technical reporting, analysis, and presentation. Experience delivering governance meetings, technical reviews, and QBRs. Desirable ITIL Intermediate or equivalent. Relevant technical certifications (e.g. Microsoft, Cisco, AWS, security). Salesforce ...

SAP FICO Consultant

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
SupportProvide Level 2/Level 3 service desk support for the SAP PR1 environment, ensuring timely resolution of incidents, defects, and service requests. Perform root cause analysis and implement sustainable solutions to minimize recurring issues. Monitor system performance and proactively identify opportunities for process and system improvements. … service management frameworks. Experience working within large multinational organizations and shared services environments. Key CompetenciesSAP FICO Functional LeadershipFinancial Process ExpertiseProduction Support & Incident ManagementIntegration ManagementRequirements Analysis & Solution DesignProject Delivery & Stakeholder ManagementSAP S/4HANA Transformation SupportProcess Improvement & OptimizationCross-Functional CollaborationStrong Communication and Documentation SkillsAbout NTT DATANTT DATA ...

Micromobility Planning Analyst, AMZL - ORION - Onroad Resourcing, Intelligence & OptimizatioN (Fulfillment & Operations)

Location
Greater London, England, United Kingdom
that impact daily business, has a strong delivery record and has experience in driving execution in a cross‐functional environment, backed by well defined analysis and research. They thrive in a fast‐paced environment, relish working with data, and enjoys the challenge of solving complex and ambiguous problems. … architectures, data modeling, infrastructure components, ETL/ELT and reporting/analytic tools and environments, data structures and hands‐on SQL coding Experience in root cause analysis and troubleshooting or problem solving Experience using strong customer service, communication, and interpersonal skills Knowledge of SQL/Python/ ...

Microsoft Dynamics D365 Developer

Location
Leeds, England, United Kingdom
Agile ceremonies including sprint planning, backlog refinement, demos, and retrospectives. Support CI/CD deployments using Azure DevOps and Git repositories. Perform troubleshooting, root cause analysis, and production support activities. Produce technical design documentation and contribute to solution design discussions. Required Skills Technical Microsoft Dynamics ...

4th Line Cloud Support Engineer

Location
Basingstoke, England, United Kingdom
technical leadership. The Role You will be involved in: Responding to complex escalations from 3rd Line engineers Assisting in Problem Management investigation and root cause analysis Extensive work with VMware products and associated monitoring Creating opportunities to automate manual tasks to drive efficiency and reduce downtime Identifying ...

Site Reliability Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
issues and measure system performance. Working alongside development and product teams to build scalable and resilient services. Responding to production incidents and carrying out root-cause analysis and improvements. Contributing to the wider DevOps and SRE community. The focus is on engineering solutions rather than simply managing ...

Site Reliability Engineer

Hiring Organisation
Anson Mccade
Location
Gloucester, Gloucestershire, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
issues and measure system performance. Working alongside development and product teams to build scalable and resilient services. Responding to production incidents and carrying out root-cause analysis and improvements. Contributing to the wider DevOps and SRE community. The focus is on engineering solutions rather than simply managing ...

Senior DevOps Engineer

Location
Leeds, England, United Kingdom
Deliver automation initiatives end-to-end, working with DevOps engineers and cross-functional stakeholders to standardize practices and improve efficiency. Troubleshoot environment issues, perform root cause analysis, and drive long-term stability and repeatability improvements. Produce comprehensive documentation including architecture diagrams, runbooks, automation scripts, and operational processes. ...

Snowflake DevOps Engineer

Location
Greater London, England, United Kingdom
Reliabilityo Implement monitoring, logging, and alerting across the data platformo Ensure high availability and performance of pipelines and integrationso Troubleshoot production issues and drive root cause analysis Security & Complianceo Embed security best practices across infrastructure and pipelineso Manage secrets, encryption, and access controls in line with financial ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
Grafana and Splunk. Proactively identify network bottlenecks, performance issues, and reliability risks, implementing long‐term fixes rather than reactive solutions. Support incident response, root cause analysis, and post‐incident reviews with a focus on continuous improvement. Collaborate with cross‐functional engineering, security, and operations teams to ensure ...

Cloud & Infrastructure IT Systems Engineer

Hiring Organisation
Sanderson
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£550.00 - £650.00 per day
Responsibilities Act as an escalation point for complex infrastructure and systems issues. Support, maintain and enhance on-premise and cloud-based services. Perform root cause analysis and major incident resolution. Develop and maintain automation, operational tooling and reusable infrastructure components. Support infrastructure security, vulnerability remediation and compliance ...

Infrastructure Engineer - Azure/M365

Hiring Organisation
The Curve Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
enterprise PKI Supporting Azure infrastructure, networking, monitoring and automation Using PowerShell to automate administration, remediation and operational tasks Investigating complex incidents, carrying out root-cause analysis and driving permanent fixes Supporting monitoring and observability across infrastructure and applications, identifying issues before they become major incidents Working with ...

Desktop Support Engineer

Location
Greater London, England, United Kingdom
regulated environment. Work with third-party technology vendors and maintain technical documentation and user guides. Identify recurring issues and drive continuous improvement through root-cause analysis, knowledge sharing and user training. Experience Required Experience providing IT support within a financial services environment. Strong knowledge of Windows ...

Integration Architect - SAP SAAS Products

Location
Greater London, England, United Kingdom
integrations. Operations & Reliability • Set SLAs/SLOs; implement monitoring and alerting via Cloud ALM and SIEM. • Drive incident/problem management and root-cause analysis; maintain runbooks. Release & Change Management • Plan for SaaS release impacts; maintain compatibility and regression tests. • Standardize CI/CD for integrations with ...

DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
availability, with a focus on zero downtime. Proactively monitor and maintain Kubernetes clusters, identifying and resolving issues before they impact service. Investigate incidents, perform root cause analysis and implement preventative measures. Plan, test and deliver infrastructure and application upgrades. Develop, test and maintain backup and disaster recovery ...

Technical Account Manager

Location
West of England, England, United Kingdom
solving and proactive service improvement. What You Will Do: Provide advanced technical support for critical incidents, infrastructure issues, and escalated service requests. Perform troubleshooting, root cause analysis, problem management, and technical documentation. Maintain, monitor, patch, and improve infrastructure systems to support availability, resilience, security, and performance. Prepare ...