76 to 100 of 747 Root Cause Analysis Jobs in London

Production Support Engineer

Hiring Organisation
Tata Consultancy Services
Location
City of London, London, United Kingdom
maintain manufacturing documentation, including work instructions, process plans, control plans and inspection requirements. Investigate production disruptions, non-conformances and process deviations, implementing containment, root-cause analysis and effective corrective actions. Drive continuous improvement initiatives using Lean methods and production data to reduce recurring issues, defects, waste … FAIR, process capability studies, SPC, MSA and Gauge R&R activities. Demonstrated ability to investigate production issues using structured problem-solving methods such as Root Cause Analysis, 8D and 5-Why, and to implement effective corrective actions. Practical experience supporting precision manufacturing, machining, assembly or inspection operations ...

Technical Support Team Manager, EMEA

Location
Greater London, England, United Kingdom
managerial responsibility, you must actively take on all forms of post-sales technical support to customers and partners, such as troubleshooting, problem solving, root cause analysis, setup, and installation assistance for a wide range of electronics, loudspeakers, and software. You will contribute to our public knowledge management … coaching. Managestaffreal time and inone-on-onesettings to increase staff performance levels. Createreference materials, help articlesandengineering reportstomaintaina robustknowledgebase. Provide advanced-level technical support, troubleshooting, root cause analysis, installation assistanceand programming assistanceto our customers and partners, as well as our sales managers onor off-site. Execute service-related ...

Ops Lead Engineer – Big Data Platform

Location
Greater London, England, United Kingdom
Azure data platform (Synapse, Databricks, ADF, Power BI). Act as the primary point of contact for incidents and outages — driving resolution, root cause analysis, and clear stakeholder communication. Define, implement, and enforce SLAs for critical pipelines, datasets, and reporting assets. Run FinOps forums with business stakeholders … Proven track record in leading operations for large-scale data platforms, ensuring stability, performance, and stakeholder trust. Incident & SLA Management: Skilled in incident triage, root cause analysis, escalation handling, and defining/enforcing SLAs with cross-functional teams. Azure Data Stack: Hands-on experience with Azure Synapse ...

Senior Network & Infrastructure Lead Engineer

Location
Greater London, England, United Kingdom
office, warehouse, data centre and remote‐site connectivity. Improve network monitoring, alerting, visibility and capacity management. Lead network lifecycle, refresh and modernisation initiatives. Conduct root cause analysis following major incidents and recurring faults. Review technical designs, network changes and implementation plans. Nutanix Infrastructure Take technical ownership … OSPF and BGP . Experience with SD‐WAN, firewalls and VPN technologies . Experience supporting large, multi‐site or distributed environments. Strong troubleshooting and root cause analysis skills. Experience leading technical resolution during major incidents. Experience designing resilient, highly available network and infrastructure solutions. Experience with monitoring ...

Regulatory Compliance Manager

Hiring Organisation
Allica Bank
Location
London, United Kingdom
Salary
£ 70 K
standards, evidence management and senior governance reporting.Out of scope: Conduct Risk ownership, Consumer Duty/customer outcome monitoring, fair value, vulnerable customer oversight, complaints root cause, DISP ownership, customer redress, Financial Crime, prudential risk ownership, operational resilience ownership, formal legal advice, Internal Audit and 1LOD control operation.Principal AccountabilitiesThe … business changes, partnerships, customer segments, distribution models, countries of operation and structural changes.Own or support regulatory change and horizon scanning, including impact assessment, gap analysis, implementation oversight and reporting to governance forums.Review and challenge new product papers, including product compliance checklists, financial promotions, product information, and any rule applicability.Oversee ...

Site Reliability Engineer / Production Support

Location
Greater London, England, United Kingdom
actively debug incidents, escalate effectively and drive permanent fixes. You will also be a builder: using AI tools for automated alert correlation, root cause analysis and runbook generation. For the right person, this is a rare opportunity to own production reliability at a pre-IPO challenger bank … reliability engineering: resilience patterns, performance tuning, capacity planning. Facilitate post-incident reviews and track actions to completion. Use AI tools for automated alert correlation, root cause analysis, and runbook generation. Hunt for routine/common tasks and formulate plans on how to automate and then execute them. ...

Reliability Engineer SME (M&E)

Location
Greater London, England, United Kingdom
technical support to site teams during live incidents, remote or on site, helping diagnose faults and guiding safe recovery of M&E systems. Own Root Cause Analysis (RCA) creation and closure, and drive Corrective and Preventive Actions (CAPA) through to completion. Lead structured problem solving on complex … Comfortable providing technical sign‐off on change requests, understanding the operational risk of works on live M&E systems. Experience leading or contributing to Root Cause Analysis (RCA) and Corrective and Preventive Action (CAPA) following M&E system incidents, with a structured approach to problem solving. Confident ...

Software Engineer, ChatGPT Infrastructure

Location
Greater London, England, United Kingdom
safe rollout and rollback. Strengthen monitoring, alerting, and diagnostics to detect problems early and reduce customer impact. Participate in on‐call, incident response, and root‐cause analysis, and turn operational learnings into lasting engineering improvements. Work across product and infrastructure teams to ensure new systems are reliable … maintainable production code. Experience designing, building, or improving services, platforms, or shared infrastructure at scale. Familiarity with production reliability practices, including monitoring, incident response, root‐cause analysis, and operational readiness. Experience diagnosing performance, scalability, or reliability issues in production environments. Understanding of distributed systems, data storage, concurrency ...

Validation Engineer II

Location
Greater London, England, United Kingdom
Reports: Author clear engineering test reports — data tables, plots, test conditions, pass/fail outcomes, and observations that engineering teams can act on directly. Root Cause Analysis: Participate in fault investigation, applying systematic thinking to identify failure modes in hardware under test. External Test Houses: Prepare samples … mill, pillar drill) and hand tools for light fabrication of fixtures and jigs. Scripting: Python or MATLAB for test data processing or basic automation. Root Cause Analysis: Familiarity with structured RCA techniques (5-Why, Fishbone, FMEA). 5S/Lean: Experience running or auditing a 5S programme ...

Head of Systems Engineering, Gill Group, Lymington, Hampshire (Hybrid)

Location
London, United Kingdom
systems engineering capability across the business Defining and governing product and system architectures Owning requirements management, verification and validation activities Improving reliability engineering, root cause analysis and product performance Introducing scalable systems engineering processes and governance frameworks Supporting the effective use of outsourced engineering and development partners … three disciplines Proven experience delivering complex products to market Expertise in systems architecture, requirements management and verification & validation A track record of improving reliability, root cause analysis and engineering governance Experience building teams, developing capability and influencing stakeholders The ability to balance strategic thinking with practical delivery ...

Head of UAV Integration

Hiring Organisation
Electus Recruitment
Location
London, United Kingdom
Employment Type
Permanent
teams to resolve build and integration problems Managing engineering changes through the build process Investigating integration failures, non-conformances and recurring build issues Driving root-cause analysis and corrective actions Coordinating system-level testing and fault-finding activity Ensuring completed aircraft are ready for acceptance and subsequent … control Engineering drawings and technical documentation Bills of Materials/BOM control Engineering Change Requests/Orders Manufacturing and assembly instructions Non-conformance management Root Cause Analysis/corrective action Production engineering System-level fault finding Verification, test and acceptance activity Supplier or subcontract manufacturing management Crucially ...

Senior Business Analyst

Location
Greater London, England, United Kingdom
Acting as an information source and communicator between multiple parties/teams, including but not limited to: Gathering business requirements from customers. Providing business analysis based on the requirements for solution design. Assisting Client PM in providing business impact analysis and milestone definition for projects. Preparing documents … interpreting business requirements. Assisting internal tech teams for requirements details explanation during development and testing. Supporting customer implementations, including issue capturing/preliminary analysis/reporting/cause analysis/business impact analysis/etc. Coordination with different teams for trouble shooting. Organizing and summarizing issue ...

Forward Deployed Engineer

Location
Greater London, England, United Kingdom
actionable technical requirements and solution designs, building strong relationships as their primary technical advisor and leading technical discussions.* ****Conduct Deep-Dive Technical Debugging and Root Cause Analysis:**** Perform detailed analysis, debugging, and root cause identification for complex system interactions, data flows, and AI model ...

TikTok Commerce - Country Business Partner (Governance and Experience)

Location
Greater London, England, United Kingdom
loop between country and regional Policy, Product, Operations, Governance, and other cross-functional teams Proactively identify and resolve governance and operational risks, conducting structured root-cause analyses and driving cross-functional mitigation plans through to measurable outcomes Own the market health and operational performance, monitoring key indicators - including … strategy, program management, governance, policy, or platform operations, preferably in technology, e-commerce, or digital platforms Strong analytical and problem-solving skills, including trend analysis, root-cause analysis, and translating insights into action Proven ability to build operating mechanisms, improve processes, and monitor business or operational ...

Technical Support Engineer Tier 3 - London, UK

Location
Greater London, England, United Kingdom
screening process. Your Daily Adventures Will Include Serve as the primary point of escalation for technical support issues requiring advanced troubleshooting, accurate reproduction, log analysis, and technical context Serve as a subject matter expert in at least one product-functional area, and be able to conduct advanced troubleshooting across … Engineering team Regularly resolves moderately complex customer problems requiring investigation across multiple systems, logs, configurations, or workflows, applying structured analytical thinking to identify probable root causes Participate in the incident response rotation throughout the entire incident lifecycle, including incident detection, incident management, customer communications, and root cause ...

WAF Engineer

Location
Greater London, England, United Kingdom
services within the WTW environment in a Tier 3 capacity. The role is London based with a hybrid work style. The Role Perform analysis and tuning of WAF policies to minimise false positives and false negatives while maintaining an appropriate security posture. Design, implement, maintain and optimise WAF policies … rule exclusions, rate-limiting policies and bot protection controls. Support transition of WAF policies from Detection mode to Prevention/Block mode through structured analysis, tuning, testing and stakeholder engagement. Analyse attack patterns, logs and telemetry to identify emerging threats and implement effective mitigations. Work closely with application owners ...

Application Support Engineer

Location
Greater London, England, United Kingdom
manage application enhancements through to live operation. Key Responsibilities Participate in proactive knowledge transfer activities with incumbent suppliers Conduct code review and quality analysis, including the review of complete services and implementation of code scanning tooling Review and improve technical documentation such as architecture overviews, deployment process definition … support are dealt with according to set standards and procedures, and suggest ways to improve those processes over time Participate in incident investigation/root‐cause analysis and deliver technical solutions targeting the root cause within agreed SLAs Implement application enhancements to improve business performance ...

Azure Data Support Engineer

Hiring Organisation
Avanade
Location
London, UK
Employment Type
Full-time
PowerShell, or Bash to reduce manual intervention and improve service reliability. SQL and Data Warehouse OperationsWrite, optimize, and troubleshoot SQL queries for: Data validationRoot cause analysisReport troubleshootingSupport and maintain data warehouse environments, such as: Azure Synapse AnalyticsSQL Server/Azure SQL DBSnowflake or BigQuery (optional, if used in hybrid … investigate slow-running queries and data load failures. Issue Investigation & RCAInvestigate job failures and performance issues across data pipelines, ML endpoints, and dashboards. Perform root cause analysis (RCA) and provide short-term and long-term solutions. Develop and implement self-healing automation for recurring failures. Service Operations ...

Lead Applications Support

Location
Greater London, England, United Kingdom
activities. Ensure application availability, performance, and stability. Act as an escalation point for complex technical issues. Incident, Problem & Change Management Perform detailed troubleshooting and root cause analysis. Drive problem management activities and identify opportunities to reduce recurring incidents. Support change, release, and deployment activities. Contribute to major incident … application engineering experience within financial services or another complex, highly regulated environment. Experience supporting business-critical applications in production. Excellent troubleshooting, diagnostic, and root cause analysis skills. Strong understanding of Incident, Problem, Change, and Release Management practices. Experience working with third-party vendors and technology suppliers. Advanced ...

Senior Risk Manager - Technology

Location
Greater London, England, United Kingdom
operational risks. This role integrates risk management into strategic and change initiatives, ensuring alignment with business objectives. It also provides expert advice, delivers insightful analysis and reporting, ensures compliance with the Enterprise Risk Management Framework and Business Continuity Management requirements, and promotes a risk-aware culture through training … across business-as-usual, emerging threats, strategic initiatives, and change programs.* Support and challenge risk assessment activities (e.g., RCSA, Top-Down Risk Assessment, Scenario Analysis, Root Cause Analysis).* Deliver timely and accurate risk insights to inform decision-making.Risk Appetite* Define and review risk appetite thresholds ...

Senior Risk Manager - Technology

Location
City Of London, England, United Kingdom
operational risks. This role integrates risk management into strategic and change initiatives, ensuring alignment with business objectives. It also provides expert advice, delivers insightful analysis and reporting, ensures compliance with the Enterprise Risk Management Framework and Business Continuity Management requirements, and promotes a risk-aware culture through training … across business-as-usual, emerging threats, strategic initiatives, and change programs. Support and challenge risk assessment activities (e.g., RCSA, Top-Down Risk Assessment, Scenario Analysis, Root Cause Analysis). Deliver timely and accurate risk insights to inform decision-making. Risk Appetite Define and review risk appetite ...

Tivoli Netcool OMNIbus SME – 11850CF

Location
Greater London, England, United Kingdom
incidents, problems and service requests. Lead investigations into monitoring failures, event‐processing issues and integration problems. Provide technical support during major incidents and conduct root cause analysis. Integrate monitoring platforms with ITSM, ticketing, automation and reporting tools, including ServiceNow . Develop automation scripts and tooling to improve operational … event rules, triggers, Probe rules files and enrichment solutions. Strong Linux/UNIX administration and Bash, KornShell or similar scripting skills. Excellent troubleshooting and root cause analysis capabilities. Experience working within ITIL‐aligned Incident, Problem, Change and Service Request processes . Strong communication and stakeholder management skills. ...

Tivoli Netcool OMNIbus SME

Hiring Organisation
Proactive Appointments
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£750.00 - £815.00 per day
incidents, problems and service requests. Lead investigations into monitoring failures, event-processing issues and integration problems. Provide technical support during major incidents and conduct root cause analysis. Integrate monitoring platforms with ITSM, ticketing, automation and reporting tools, including ServiceNow . Develop automation scripts and tooling to improve operational … event rules, triggers, Probe rules files and enrichment solutions. Strong Linux/UNIX administration and Bash, KornShell or similar scripting skills. Excellent troubleshooting and root cause analysis capabilities. Experience working within ITIL-aligned Incident, Problem, Change and Service Request processes . Strong communication and stakeholder management skills. ...

3rd Line Network Engineer

Hiring Organisation
IMT Resourcing Solutions
Location
London, Bishopsgate, United Kingdom
Employment Type
Permanent
Salary
£45000 - £50000/annum
established Monitor network events with the potential to impact multiple customers and correlate these against incoming incidents Lead problem management for recurring faults, identifying root causes and driving permanent fixes Produce clear Root Cause Analysis (RCA) and Reason for Outage documentation for customer-facing incidents Develop … Line Network Engineer or advanced 2nd Line Network Engineer who enjoys getting into complex technical problems, taking ownership of incidents and finding the root cause rather than applying temporary fixes. Apply now to take on a senior technical role supporting business-critical network services across London. ...

Team Leader - Data Engineering & Integration - Commodities Data

Hiring Organisation
Bloomberg
Location
London, United Kingdom
Salary
£ 70 K
datasets across Power and Gas, Oil, Carbon, Agriculture and Metals. The team provides relevant, timely and accurate data to empower customers to drive their analysis of commodity markets, both pricing and fundamentals.The Role:This role will own the modernization, stability, and scalability of core commodities data manufacturing capabilities across … agreed entities, identifiers, relationships, taxonomies, metadata, and lifecycle rules.Ensure datasets are delivered with clear ownership, controls, lineage, documentation, support models, and auditability.Improve monitoring, alerting, root-cause analysis, recovery processes, and preventative controls to reduce operational risk.Reduce duplicate workflows, redundant processes, manual workarounds, and fragmented ownership across commodities ...