751 to 775 of 1,943 Root Cause Analysis Jobs in the UK

Team Lead - Full Stack - Manchester

Location
Manchester, England, United Kingdom
work closely with the Software Development Manager to plan sprints, prioritise the backlog and manage delivery risk, and you’ll own incident response and root cause analysis for issues within your team’s applications. This role will be based out of our central Manchester office 3 days ...

DevOps Engineer

Location
Wallingford, England, United Kingdom
repeatable environments, reducing manual processes and increasing automation. Support monitoring, logging, observability and alerting systems, via Grafana, to support production health, incident detection, root-cause analysis and rapid recovery. Drive a culture of DevOps best-practice across teams: infrastructure reliability, deployment-velocity, automated testing, rollback strategies, security ...

Senior Systems Engineer

Location
Greater London, England, United Kingdom
Online Act as a Technical Lead for migrations, implementations, relocations, etc Take ownership of complex technical challenges and provide long term fixes based on root cause analysis Document processes and procedures which explain system functionality or regular tasks to contribute to our internal knowledge base including architecture ...

Data Operations Specialist

Location
Greater London, England, United Kingdom
best practices and standards for ongoing data management.* Incident Management & Resolution: Serve as the escalation point for complex or high-impact data issues. Lead root cause analysis, post-mortems, and the design of long-term preventative solutions to reduce operational risk.* Knowledge Leadership: Act as a subject ...

Data Governance Manager

Location
Manchester, England, United Kingdom
data quality frameworks, including data quality rules, controls, issue management and remediation processes. Identify data quality issues and work with relevant stakeholders to drive root-cause analysis and resolution. Establish appropriate data controls, governance processes and standards across key data domains. Support the development and ongoing management ...

Senior Site Reliability Engineer

Location
Swindon, England, United Kingdom
practices. Experience with observability and monitoring tools such as Azure Monitor, Application Insights, Log Analytics, Grafana, Prometheus, or similar technologies. Understanding of incident management, root cause analysis, service reliability metrics, and continuous service improvement. Experience with configuration management solutions such as Puppet. #LI-MP #LI-Hybrid About ...

Senior Data Governance Analyst

Location
Cheadle, England, United Kingdom
there is no direct authority. Experience developing and embedding data governance policies, standards, procedures and proportionate controls. Strong analytical and problem-solving capability, including root cause analysis and the ability to translate complex findings into practical recommendations. Intermediate Excel and SQL skills, with the ability to interrogate ...

Lead Software Engineer - Cloud Platform Engineering

Location
Greater London, England, United Kingdom
work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team. ...

Data Analyst

Hiring Organisation
Experian
Location
Nottingham, UK
Employment Type
Full-time
interrogate, manipulate, and validate complex datasets. Apply robust validation and quality assurance processes to all deliverables. Investigate data-related issues and client queries, conducting root cause analysis and communicating findings to both technical and non-technical audiences. Support the delivery of proof-of-concepts, new products ...

SRE Engineer with .Net C# - Glasgow, UK

Location
Glasgow, Scotland, United Kingdom
requiredYour Skills:Site Reliability Operational SupportMonitor maintain and support businesscritical applications and cloud infrastructureEnsure high system availability performance scalability and reliabilityParticipate in incident management root cause analysis and problem resolutionImplement proactive monitoring alerting and observability solutionsReduce operational overhead through automation and selfhealing mechanismsSupport production releases and deployment ...

Site Reliability Engineer (SRE) - Glasgow, UK

Location
Glasgow, Scotland, United Kingdom
Reliability Operational Support* Monitor maintain and support businesscritical applications and cloud infrastructure* Ensure high system availability performance scalability and reliability* Participate in incident management root cause analysis and problem resolution* Implement proactive monitoring alerting and observability solutions* Reduce operational overhead through automation and selfhealing mechanisms* Support production ...

ITSM Process Support Analyst

Location
Bangor, Caernarfonshire, United Kingdom
ensuring records are logged and kept up to date, assisting with prioritisation and progress tracking, and coordinating inputs from technical teams. Arrange and document root cause analysis sessions under guidance and maintain the known error record and remediation action log, escalating overdue actions and risks ...

ITSM Process Support Analyst

Location
Carrickfergus, County Antrim, United Kingdom
ensuring records are logged and kept up to date, assisting with prioritisation and progress tracking, and coordinating inputs from technical teams. Arrange and document root cause analysis sessions under guidance and maintain the known error record and remediation action log, escalating overdue actions and risks ...

ITSM Process Support Analyst

Location
Downpatrick, County Down, United Kingdom
ensuring records are logged and kept up to date, assisting with prioritisation and progress tracking, and coordinating inputs from technical teams. Arrange and document root cause analysis sessions under guidance and maintain the known error record and remediation action log, escalating overdue actions and risks ...

Senior Data Engineer (Cloud, Data Integration & Migration)

Location
Greater London, England, United Kingdom
premise, and third-party systems. Design and implement robust data migration solutions, including extraction, transformation, validation, reconciliation, and cutover activities. Perform data profiling, root cause analysis, and data quality assessments. Develop and optimise data warehouse, data lake, and lakehouse solutions. Produce source-to-target mappings, transformation rules ...

DevOps Engineer (Linux, PHP)

Location
Greater London, England, United Kingdom
production (Docker). 2+ years’ experience with monitoring and observability tooling (New Relic, Grafana/Loki or equivalent), including responding to production incidents — triage, root-cause analysis and post-mortems. 1+ years’ experience operating a container orchestration platform (Docker Swarm; Kubernetes an advantage). Comfortable using ...

Senior Lead Software Engineer - Python, Data, Cloud, AIML

Location
Greater London, England, United Kingdom
across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
About You You bring deep infrastructure and build systems expertise alongside an SRE mindset. You are systematic in how you approach fault isolation and root cause analysis, and you apply metrics-driven thinking to build reliability and developer productivity. You thrive on variety (ML, compilers, kernel drivers ...

Workplace Platforms Principal Engineer

Location
Greater London, England, United Kingdom
scripting. Have experience producing and reviewing high‐level designs, detailed technical designs, operational documentation and service transition deliverables. Possess extensive experience supporting major incidents, root cause analysis, problem management and complex technical escalations. Be experienced in infrastructure lifecycle management, including patching, vulnerability remediation, backup, restoration and disaster ...

Sr Lead Software Engineer - Python/ AWS/Databricks

Location
Greater London, England, United Kingdom
across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within ...

Platform Site Reliability Engineer

Location
Gloucester, England, United Kingdom
stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations Mentor junior engineers and act as an Operational requirements consultant to other departments Communicate technical ...

Cloud Platform Lead

Location
Glasgow, Scotland, United Kingdom
services are: Fully supportable (aligned to ITSM/Halo) Monitored, documented, and resilient Define: Runbooks, alerting, and operational standards Support major incident resolution and root cause analysis Lead adoption of: Automation, AI-enabled operations, and platform tooling Identify opportunities to: Reduce operational toil Increase platform reuse ...

Software Engineer III - Full Stack, Global Banking Tech

Location
Glasgow, Scotland, United Kingdom
jobs and batch processing such as Spring Batch) where applicable to platform capabilities Support operational stability and reliability by contributing to incident response and root cause analysis, analyzing diverse datasets, logs, and telemetry to identify patterns, build visualizations and reporting, and reduce repeat incidents via preventative controls ...

Technical Operations Manager - 6 Month FTC – Kings Cross, London

Location
Greater London, England, United Kingdom
influencing positive change. Thorough knowledge of IT operations, information technology best practices, and industry trends. Professional ITIL based problem-solving approach that includes root cause analysis, corrective action, and buy-in Proven experience in supporting mission critical infrastructure and applications on a global scale Proven ability ...

Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations Mentor junior engineers and act as an Operational requirements consultant to other departments Communicate technical ...