276 to 300 of 484 Incident Management Jobs in London

Privacy Analyst II

Location
Greater London, England, United Kingdom
audits. Third-party vendor due diligence, working with InfoSec, Legal, and Procurement to assess vendor posture and manage subprocessor obligations. Lead on data privacy incident response, including logging, investigation, root cause analysis, remediation, and regulatory and client reporting. Develop and deliver training and knowledge content on global data protection. … subsidiaries. Managing the privacy mailbox, including handling rights requests. What you’ll bring to the party Proven experience in the following: Core requirement Data incident management, incl. root cause analysis, remediation, logging and reporting. Significant operational experience in DPIAs/PIAs, third party vendor management, and risk ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
capability.This will include:• AI Platform Strategy & Architecture: Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices ...

Lead IT Service Engineer

Hiring Organisation
Salt Search
Location
London, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
currently recruiting for an IT Service Engineer to join a modern, globally established wealth management organisation. Role context Works within the IT End User Experience team and reports to the IT End User Experience Team Manager - EMEA. The role acts as the principal on-site point of coordination … guiding other team members Competencies, skills and behaviours Fluent written and spoken English; additional languages are advantageous A calm, structured approach to prioritisation, incident management and problem solving Strong customer-service orientation with the confidence to communicate effectively at all levels Ability to mentor colleagues and deliver effective ...

Expert Service Delivery Manager

Location
Greater London, England, United Kingdom
tower estate. This is a senior, client-facing role for someone who is equally comfortable in an ITIL governance forum, a Sev-1 mainframe incident bridge, and a boardroom QBR. Key responsibilities Acts as a client advocate and a point of escalations for client service delivery needs, including … well as exceeding expectations. Establishes and leads operational meetings focused on ITSM governance and SLA adherence. Required Qualifications 8+ years of IT Service Management experience in a client-facing role, with meaningful exposure to mainframe (z/OS) managed service environments. Operational ability in diverse, large-scale, multi-platform ...

Senior Cloud Engineer, Cloud COE

Hiring Organisation
Janus Henderson
Location
London, UK
Employment Type
Full-time
Ensure infrastructure deployments are consistent, version-controlled, and policy-compliantManage drift detection, environment consistency, and release governanceAzure Core Platform SkillsHands-on engineering across: Subscriptions, Management Groups, RBAC, PolicyVirtual Machines, Storage, Backup, DRAzure Monitor, Log AnalyticsPrivate Endpoints and secure service exposureImplement secure-by-design configurations aligned to enterprise cloud controlsAzure … Microsoft Entra ID (Azure AD) configurations and integrationsImplement RBAC, SSO, and least privilege access modelsAssist with security controls, compliance policies, and remediation activitiesMonitoring, BAU & Incident ManagementSupport business-as-usual (BAU) cloud operations across environmentsMonitor cloud services using Azure Monitor, Log Analytics, and alerting toolsInvestigate incidents, perform root cause analysis ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
will collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one.While this is not a people‐management or shift‐based role, you will work closely with global teams (UK and US) and may occasionally be called upon for major incidents …/SLOs.Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.Partner with Security teams to ensure our platforms meet compliance, security, and risk‐management expectations.Lead seamless handovers ...

Platform Support Lead

Location
Greater London, England, United Kingdom
week) Employment: Fulltime Role Summary: The lead will act as the primary onshore contact for IDMC and TIBCO EBX support, coordinating platform operations, incident management, deployments, service transition, offshore delivery, and client communication. Key Responsibilities: Lead daily monitoring of IDMC and EBX platform health, Secure Agents, metadata scans … data-quality workflows, integrations, and scheduled jobs. Own high-priority incident coordination, perform technical triage, drive restoration, manage escalations, and support root-cause analysis. Coordinate Dev, UAT, and Production activities, including Secure Agent configuration, SSO/access issues, deployments, rollback planning, and post-deployment validation. Provide functional and technical ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve performance, availability, and reliability issues Lead and contribute to incident response, troubleshooting, and root cause analysis Automate operational processes and eliminate repetitive manual tasks Work closely with software engineers to improve deployment processes, system … understanding of metrics, logging, tracing, alerting, and system health Experience troubleshooting complex production environments Understanding of SLIs, SLOs, SLAs, and error budgets Experience with incident management and root cause analysis Good understanding of cloud networking, security, and infrastructure fundamentals Strong scripting/automation skills A strong understanding ...

Production Engineer – Trading & Electronic Trading Systems

Location
Greater London, England, United Kingdom
live trading and electronic trading platforms. These environments are production‐critical, operate in real time and require strong ownership of system stability, performance and incident management. The role involves close interaction with traders, developers, IT support and infrastructure teams. Role Overview We are looking for a Front Office Production … Responsibilities Production Ownership & Reliability Operate and support Front Office trading systems in production Ensure high availability, stability and performance during market hours Own incident management from detection to resolution and root cause analysis Participate in on‐call or production support rotations when required Contribute to post‐incident ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
distributed and event-driven systems using messaging and asynchronous processing patterns. Drive engineering excellence through architecture, coding standards, testing, and automation. Support production systems, incident management, troubleshooting, and operational resilience. Balance long-term architecture goals with short-term delivery priorities. Collaborate across engineering teams, product owners, architects … production support. Non-Technical Skills Strong leadership, ownership, and problem-solving mindset. Ability to thrive in ambiguous and rapidly changing environments. Excellent stakeholder management, communication, and collaboration skills. Proven ability to influence engineering practices beyond immediate teams. Experience coaching and developing engineers while driving continuous improvement. If you like ...

Site Reliability Engineer

Hiring Organisation
Trainline
Location
London, UK
Employment Type
Full-time
cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices. The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery … Trainline, you'll be working on...Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platformParticipating in production incident response, supporting investigation, mitigation, communication, and coordinated service restorationContributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilienceTaking part ...

Onshoring - Capacity Manager

Hiring Organisation
Hays Technology
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600/day Up to £600pd inside IR35 via umbrella
Identify potential capacity constraints and performance risks before they impact users or services. Coordinate capacity planning activities with technical and operational teams. Support major incident management by assessing capacity and resilience impacts. Contribute to service recovery planning where capacity and performance are affected. Produce demand forecasts and growth … related risks and issues to operational and leadership stakeholders. Maintain capacity dashboards, reports, and continuous improvement plans. Key Skills & Experience Proven experience in Capacity Management within a complex IT environment. Strong understanding of infrastructure, application performance, and service management principles. Experience producing capacity forecasts, utilisation reports, and risk ...

Software Developer - Risk Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
stability and availability of the Risk software systems, as well as their day-to-day operations. Squarepoint's Risk platform is responsible for position management, profit/loss computation, inventory/locate management and internal order routing. These critical systems need to be performant, resilient, and capable … business hours, people on-duty will prioritise responding to incidents over their project work. On average, people are on-duty one day per week. Incident management: Root cause analyses are performed to understand the source of incidents and to raise appropriate remedial actions. Day-to-day operations: Until ...

Technical Support Engineer

Location
Greater London, England, United Kingdom
issues through to resolution Present complex technical information to non-technical audiences Provide accurate and complete problem resolution documentation for future reference and management reporting Take part in the creation and maintenance of knowledge base data Increase subject matter knowledge on Medidata products Work with other Medidata teams … experience querying and modifying data using MYSQL, MS SQL Server, MongoDB, Snowflake Experience with C#, Javascript, Visual Studio, Ruby, Python and ASP.NET Experience using Incident Management and/or Project Management software Experience with multiple integrated software products/services Experience working with customers to resolve complex ...

Trainee DevOps Engineer | No experience needed (Ref: 7501)

Hiring Organisation
Qualify Nation Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£28,000 - £38,000 per annum
Continuous Deployment (CD) Infrastructure as Code (IaC) Containerisation with Docker Container Orchestration Cloud Computing Fundamentals Cloud Platforms (AWS, Microsoft Azure and Google Cloud) Configuration Management Monitoring and Logging Security Best Practices (DevSecOps) Networking Fundamentals Automation and Scripting Incident Management and Reliability Engineering Practical Experience You will work ...

People Systems Analyst (6 month FTC)

Location
Greater London, England, United Kingdom
Benefits or Payroll is desirable. Comfortable analysing and resolving system and process issues in a structured manner. Strong organisational skills, with experience in operating incident management tools. Ability to prioritise and manage multiple requests within agreed service levels. Familiarity with Microsoft Office suite. Experience using ServiceNow and JIRA … work well within a global and sometimes virtual team. Understanding of HR data governance, confidentiality requirements and system administration best practices. Good stakeholder management skills, with the ability to communicate clearly with technical and non-technical audiences. Ability to handle sensitive people data with discretion and in line with ...

Application Support Analyst | Operations Technology

Location
Greater London, England, United Kingdom
business users offering second- and third-line support. Knowledge of scripting language (such as PowerShell, Python). Manage new system analysis and implementation. Provide incident management per ITIL standards. Liaison between technology department and business groups to communicate system changes. Manage process and trading system documentation in existing … report any breaches of policy to Compliance and/or your supervisor as required To elevate risk events immediately To provide input to risk management processes, as required. Competencies , Skills and Experience Skills and Experience: Essential: Solid background in Windows, Linux/Unix OS, including one of the following ...

Application Support Analyst | Energy Trading Operations

Location
Greater London, England, United Kingdom
Support business users offering second‐ and third‐line support. Knowledge of scripting language (PowerShell, Python, etc.). Manage new system analysis and implementation. Provide incident management per ITIL standards. Liaison between technology department and business groups to communicate system changes. Manage process and trading system documentation in existing … responsibility. Report any breaches of policy to Compliance and/or your supervisor as required. Escalate risk events immediately. Provide input to risk management processes, as required. Competencies, Skills and Experience Competencies: A collaborative team player, approachable, self‐efficient and influences a positive work environment. Demonstrates curiosity. Resilient ...

Network Engineer

Location
Greater London, England, United Kingdom
Python, Ansible, Jinja2 and Git Supporting and extending Infrastructure-as-Code and CI/CD practices for network provisioning Owning break/fix and incident response for complex enterprise and campus network issues Driving incident management practices, including on-call rotation, escalation procedures, postmortems and root-cause ...

Software Engineer III- JPM Personal Investing

Location
Greater London, England, United Kingdom
investing, and the trust that 150 years of J.P. Morgan heritage brings. J.P. Morgan Personal Investing offers award-winning investments, products and digital wealth management services to over 275,000 investors in the UK. We built the business with innovation as a core part of our ethos to give … Kotlin Write unit, component, integration, end-to-end & performance tests Support the products you've built through their entire life cycle, including production and incident management Required qualifications, capabilities and skills Formal training or certification on Java programming concepts and proficient advanced experience Recent hands-on professional experience ...

Trade Desk Associate

Location
Greater London, England, United Kingdom
client access to platforms, provisioning new services and providing ongoing oversight and monitoring. Certification testing, logical port creation and modification of default settings. Coordinating incident management across teams and conducting post-incident review, then tracking the follow-up actions to closure. Documenting issue resolution and communication ...

Cloud Platforms Engineer

Location
City Of London, England, United Kingdom
than a queue of requests. Keep production healthy. Help diagnose and resolve production issues when they arise, bring a clear view of what good incident management looks like, and fix the class of problem rather than just the instance. Make sure we can recover. Look after backup … Azure or Google Cloud Platform experience is fine. Practical experience with infrastructure as code in a team setting, Terraform preferred – modules, code review, state management and applying through a pipeline, rather than clicking in a console. Desirable: experience of cloud identity and access at organisation scale, and single sign ...

Forward Deployed Engineering Lead - Global Broking

Hiring Organisation
Liquidnet
Location
London, UK
Employment Type
Full-time
trusted relationships with desks and stakeholders, ensuring priorities, progress, and risks are clearly communicated. Deliver within TP ICAP's engineering, governance, security, and change management frameworks. Experience/CompetencesEssentialExperience leading people, engineering teams, or owning the delivery of an engineering squad. Strong stakeholder management, communication, and judgement. Ability … deliver in fast-moving environments with evolving requirements. Experience operating production services, including deployment, monitoring, incident management, and support. Strong hands-on engineering skills in Python and TypeScript, with the ability to review code, set technical direction, and guide implementation. Experience using modern software development tools and practices ...

DevOps Engineer x 5

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £552.00 per day
capabilities Troubleshoot and resolve infrastructure, platform and service issues Monitor platform health, availability and performance Contribute to infrastructure improvements and operational resilience Support incident management and service recovery activities Produce and maintain technical documentation Collaborate with engineering teams to improve deployment processes and service reliability Participate in support ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. * Support live service operations, incident management, troubleshooting and platform resilience. * Collaborate with engineers, architects, product teams and stakeholders whilst coaching and mentoring colleagues. Essential Skills for the Senior ...