251 to 275 of 876 Root Cause Analysis Jobs in England

Principal Java Engineer

Hiring Organisation
Jobleads-UK
Location
Wallingford, England, United Kingdom
secure, scalable and maintainable software design across all teams. Support and mentor team members to drive technical excellence and continuous improvement. Production systems Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed ...

Lead Software Engineer - Java/Python/AI-ML

Hiring Organisation
Jobleads-UK
Location
Bournemouth, England, United Kingdom
work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team. ...

AWS Cloud Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Leeds, England, United Kingdom
Terraform) Configure and maintain AWS services including compute, storage, networking, and managed services Monitor infrastructure health, availability, and performance; respond to incidents and perform root cause analysis Partner with application engineering and security teams to support system reliability and scalability Implement and maintain backup, disaster recovery ...

Engineering Lead

Hiring Organisation
Moorepay
Location
Manchester, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Hands-On Contribution: Remain an active contributor to the codebase, setting the standard for code quality, testing, and documentation. Support troubleshooting, incident response, and root-cause analysis for squad-owned services. Promote best practices in cloud-native development, CI/CD, observability, and reliability engineering. Pair-program ...

Head of Security Engineering (DevSecOps & CISO)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
strategy across cloud infrastructure, applications, trading systems, and corporate environment Lead or support the investigation, containment and remediation of security incidents, including post-incident root cause analysis and corrective actions. Work with Engineering teams to build security into the SDLC — SAST, DAST, SCA, container/IaC scanning ...

Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
align with organisational standards and regulatory requirements.* **Operational Excellence:** Monitor, optimise, and support cloud infrastructure, ensuring high availability and performance. Contribute to incident response, root cause analysis, and continuous improvement.* **Collaboration & Documentation:** Work closely with architects, developers, and security teams to deliver integrated solutions. Produce clear technical ...

Senior Software Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
infrastructure decisions and service architecture within our Azure environment, support observability, monitoring and alerting for production services, and participate in incident response and root cause analysis when issues arise. You will take end-to-end ownership of features from technical design through to delivery and iteration, participate ...

Senior Software Engineer

Hiring Organisation
17918
Location
Westminster, West End, United Kingdom
infrastructure decisions and service architecture within our Azure environment, support observability, monitoring and alerting for production services, and participate in incident response and root cause analysis when issues arise. You will take end-to-end ownership of features from technical design through to delivery and iteration, participate ...

Senior Cloud Engineer, AI Platform SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automated eval and regression checks before release. Define and track SLOs/SLAs for AI platform services, and bring SRE rigour to incident response, root cause analysis, and postmortems. Participate in on-call rotations and maintain clear, usable runbooks. Build observability for AI-specific concerns — latency, token ...

Staff / Lead Software Engineer - Backend Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within ...

Test Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
front-end cross-browser functional testing, exploratory testing, system integration testing, and user acceptance testing as the engagement requires Accurate defect recording, monitoring, and root cause analysis, feeding quality data back into the team to drive continuous improvement Reviewing and improving test environments, CI/CD pipeline ...

Digital & Print Production Engineer

Hiring Organisation
Financial Times
Location
Greater London, United Kingdom
Employment Type
Full Time
escalating business-critical incidents where appropriate. Providing timely communication and operational updates during production-critical activities and service disruptions. Supporting post-incident reviews, root cause analysis and preventative improvements. Working with internal teams and third-party suppliers to restore services as quickly as possible. Continuous Improvement & Engineering ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop, self-healing systems that autonomously execute ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate rootcause analysis, anomaly detection, and semantic log clustering. Self‐Healing Infrastructure Engineering: Experience designing closed‐loop, self‐healing systems that autonomously execute ...

Head of Analytics

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
home, in line with the three shifts (hospital to community, analogue to digital, and sickness to prevention). The post holder will oversee analysis and reporting to publication standard where needed, with day to day work involving close collaboration with lead clinicians, academic partners, the CDO and CCDO … proactive and comprehensive approach to risk management, and be responsible for service continuity in own area, participating in the Informatics service continuity planning. Provide Root Cause Analysis (RCA) for allocated incidents and problems, instigating emergency action when required and liaising with other Trust managers as appropriate. Business ...

Head of Analytics

Hiring Organisation
The Christie NHS FT
Location
Manchester, M20 4BX, United Kingdom
Salary
£66582.00 to £77368.00
closer to home, in line with the three shifts (hospital to community, analogue to digital, and sickness to prevention).The post holder will oversee analysis and reporting to publication standard where needed, with day to day work involving close collaboration with lead clinicians, academic partners, the CDO and CCDO … proactive and comprehensive approach to risk management, and be responsible for service continuity in own area, participating in the Informatics service continuity planning. Provide Root Cause Analysis (RCA) for allocated incidents and problems, instigating emergency action when required and liaising with other Trust managers as appropriate. Business ...

Engagement Manager

Hiring Organisation
TEKsystems
Location
London, UK
Employment Type
Full-time
consultant onto Customer and TGS IT systems (Email, MS Teams, SharePoint etc.,) \n\n \n Requisite Abilities and/or Skills \n \n \n Analysis and problem-solving skills \n Time management and organizational skills \n Personnel and team management skills \n Demonstrable engagement data/risk analysis … engagement success \n\n \n Action Orientated \n \n Identifies concerns, such as sourcing gaps, and quickly communicates these \n Facilitates issue resolution using root cause analysis and identifies proper parties to communicate to \n Proactively anticipates customer needs, creates solutions and contingency plans to limit issues ...

Technology Architect - Infrastructure Architect (Capacity & Performance, Multi-Platform) - Lond[...]

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
memory, storage IO, and network utilization; recommend and implement capacity optimization; collaborate with application, database, and infrastructure teams. Performance Engineering & Optimization – Lead performance analysis and tuning across AIX/Linux/Windows and VMware environments; identify bottlenecks in CPU, memory, disk IO, and network throughput; drive proactive performance improvement … reporting, and performance tuning; improve observability across the infrastructure stack. Incident, Problem & Change Management – Lead resolution of complex L2/L3 infrastructure issues; conduct root cause analysis; work within ITIL processes (Incident, Problem, Change). Disaster Recovery & High Availability – Design DR/HA solutions for critical infrastructure ...

SOC Shift Lead - London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
application. Note: The above information relates to a specific client requirement Role Description As a SOC Shift Lead , you will provide advanced investigation and analysis, acting as the escalation point for complex and high-severity incidents. You will conduct root cause analysis, guide and mentor … lead. Essential skills: Education: Bachelor’s degree in Cybersecurity, Computer Science, or related field. Experience: 7–10 years in SOC, Incident Response, or Threat Analysis roles (for Lead role we’ll consider 5 - 7years) Strong analytical mindset, in-depth knowledge of SIEM/EDR tools, malware behaviour, and incident ...

Supply Chain Data Analyst

Hiring Organisation
Sanderson Recruitment
Location
Hampshire, South East, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £550 per day
with Supply Chain, Logistics and Commercial teams to understand business challenges and identify opportunities for improvement. Analyse large and complex datasets to identify trends, root causes, risks and opportunities across supply chain operations. Deliver deep-dive analysis into key performance areas, including service levels, inventory, supplier performance, logistics … skills for querying and analysing large datasets. Strong Power BI experience, including dashboard development and DAX. Advanced Microsoft Excel skills. Proven ability to perform root cause analysis and identify business improvement opportunities. Experience translating data into actionable insights and recommendations. Strong stakeholder management and communication skills. Ability ...

Cloud Infrastructure Engineer

Hiring Organisation
Nigel Wright Group
Location
Tyne and Wear, United Kingdom
Employment Type
Full-Time
Salary
£47,000 per annum
improvement across cloud and infrastructure operations. Key Responsibilities: Manage and optimise public cloud infrastructure, ensuring scalability, availability, and security. Implement monitoring solutions, conduct performance analysis, and automate improvements where possible. Ensure high availability and data integrity through proactive alerting, backups, and robust disaster recovery planning. Own major incident response … troubleshooting and root-cause analysis, implementing long-term fixes. Maintain security best practice across cloud and on-premise environments, including vulnerability management and compliance monitoring. Contribute to and sometimes lead the design and implementation of DR and business continuity plans. Build and maintain automation scripts and configuration ...

Senior Data Engineering Analyst

Hiring Organisation
St. James's Place
Location
Gloucestershire, United Kingdom
Employment Type
Full Time
passionate about harnessing the power of data to deliver meaningful business outcomes. The ideal candidate will bring a strong background in data engineering or analysis, a proven ability to deliver complex projects successfully, and a drive to continuously improve processes and technologies. You will be a collaborative team player … support. Advanced SQL skills with the ability to write, optimise, and troubleshoot complex database queries. Strong analytical and problem-solving skills, with experience conducting root cause analysis and implementing effective solutions. Experience working with and managing data across multiple enterprise data repositories and platforms. A passion ...

Lead SRE - Chase UK

Hiring Organisation
17918
Location
London, United Kingdom
planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start. Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation. Continuously develop AI skills relevant ...

Lead SRE - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
planning to ensure metrics, alerting, logging, automation, resiliency, capacity, and performance needs are built in from the start. Use approved AI tools to accelerate root-cause analysis, log and trace investigation, runbook drafting, post-incident analysis, test scaffolding, and documentation. Continuously develop AI skills relevant ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
single points of failure. Troubleshoot complex issues across infrastructure and platform services, participate in the on‐call rotation and incident‐management process, and turn rootcause findings into lasting reliability improvements. Collaborate with application development teams and other stakeholders to meet internal Service Level Objectives and customer‐facing … recovery testing. Experience building automation that reduces repetitive work, improves release safety, or increases infrastructure efficiency. Strong communication and documentation skills, including experience writing rootcause analyses and working across engineering teams. Strong sense of ownership, sound technical judgment, and attention to operational and security details. Why Cisco ...