276 to 300 of 497 Incident Management Jobs in London

IT Support Analyst

Hiring Organisation
GroupNexus
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£30,000 - £35,000 per annum
Full time: Mon – Fri 0900-1700 Reporting to: Head of Customer Support & Monitoring About us: GroupNexus is an established, leading operator in the parking management sector. We are innovative, industry leading and a forward-thinking company that has an exceptional culture with strong values and a hunger for growth. … bottom of the issue. This position includes performing 1st and 2nd level troubleshooting, manual alert analysis and resolving issues. Key Objectives & Responsibilities: Daily management of support tickets Daily Checks Compiling reports/ad-hoc reports Resolving tickets to SLA On-site equipment sign-off Ability to coordinate across teams ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
shape and drive how the firm builds and operates reliable, observable, secure, and cost-efficient systems on AWS. Working closely with development, platform, and incident management teams, you will define reliability in measurable terms and build the tooling and processes to achieve it, improving platform speed, stability … chaos engineering experiments to strengthen resilience and recovery. Automate operational processes to reduce manual intervention and toil across the stack. Support major incident response, root-cause analysis, and continual improvement actions. Collaborate cross-functionally to raise standards for stability, security, performance, and compliance. Required skills & experience 3+ years' experience ...

Principal Systems Architect - M365 and EUC

Hiring Organisation
Finastra
Location
London, UK
Employment Type
Full-time
Role OverviewThe Microsoft 365 Architect at Finastra is a senior, hands-on engineering and operations role responsible for the design and day-to-day management of the global Microsoft 365 environment. This role combines strategic platform architecture with practical administration and L3/L4 support, ensuring services are secure … Establish tenant-wide policies for data protection, retention, and regulatory compliance. Platform Administration & Monitoring: Administer and troubleshoot M365 services to ensure performance and availability. Incident Management & Escalation: Act as SME and escalation point for complex issues, performing root cause analysis. Migrations & Projects: Lead tenant-to-tenant migrations ...

Azure Databricks Engineer

Location
Greater London, England, United Kingdom
Databricks Platform Engineering*** Hands-on implementation and troubleshooting of Azure Databricks* Databricks Serverless configuration and workload optimisation* Databricks SQL and Delta Lake development* Cluster management, compute policies, autoscaling, and workload monitoring**Data Engineering*** Strong Python, PySpark, and SQL development skills* Design and build ingestion and transformation pipelines* Full … reconciliation, and error handling**Integration Development*** REST API integration experience* Integration with databases, SaaS applications, cloud storage, and enterprise systems* Secure authentication and credential management**Azure Platform Services*** Azure Identity and Access Management* Azure Networking and Private Endpoints* Azure Security and Key Vault* Azure Monitoring and Operational Support ...

Privacy Analyst II

Location
Greater London, England, United Kingdom
audits. Third-party vendor due diligence, working with InfoSec, Legal, and Procurement to assess vendor posture and manage subprocessor obligations. Lead on data privacy incident response, including logging, investigation, root cause analysis, remediation, and regulatory and client reporting. Develop and deliver training and knowledge content on global data protection. … subsidiaries. Managing the privacy mailbox, including handling rights requests. What you’ll bring to the party Proven experience in the following: Core requirement Data incident management, incl. root cause analysis, remediation, logging and reporting. Significant operational experience in DPIAs/PIAs, third party vendor management, and risk ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
capability.This will include:• AI Platform Strategy & Architecture: Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices ...

Lead IT Service Engineer

Hiring Organisation
Salt Search
Location
London, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
currently recruiting for an IT Service Engineer to join a modern, globally established wealth management organisation. Role context Works within the IT End User Experience team and reports to the IT End User Experience Team Manager - EMEA. The role acts as the principal on-site point of coordination … guiding other team members Competencies, skills and behaviours Fluent written and spoken English; additional languages are advantageous A calm, structured approach to prioritisation, incident management and problem solving Strong customer-service orientation with the confidence to communicate effectively at all levels Ability to mentor colleagues and deliver effective ...

Expert Service Delivery Manager

Location
Greater London, England, United Kingdom
tower estate. This is a senior, client-facing role for someone who is equally comfortable in an ITIL governance forum, a Sev-1 mainframe incident bridge, and a boardroom QBR. Key responsibilities Acts as a client advocate and a point of escalations for client service delivery needs, including … well as exceeding expectations. Establishes and leads operational meetings focused on ITSM governance and SLA adherence. Required Qualifications 8+ years of IT Service Management experience in a client-facing role, with meaningful exposure to mainframe (z/OS) managed service environments. Operational ability in diverse, large-scale, multi-platform ...

Senior Cloud Engineer, Cloud COE

Hiring Organisation
Janus Henderson
Location
London, UK
Employment Type
Full-time
Ensure infrastructure deployments are consistent, version-controlled, and policy-compliantManage drift detection, environment consistency, and release governanceAzure Core Platform SkillsHands-on engineering across: Subscriptions, Management Groups, RBAC, PolicyVirtual Machines, Storage, Backup, DRAzure Monitor, Log AnalyticsPrivate Endpoints and secure service exposureImplement secure-by-design configurations aligned to enterprise cloud controlsAzure … Microsoft Entra ID (Azure AD) configurations and integrationsImplement RBAC, SSO, and least privilege access modelsAssist with security controls, compliance policies, and remediation activitiesMonitoring, BAU & Incident ManagementSupport business-as-usual (BAU) cloud operations across environmentsMonitor cloud services using Azure Monitor, Log Analytics, and alerting toolsInvestigate incidents, perform root cause analysis ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
will collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one.While this is not a people‐management or shift‐based role, you will work closely with global teams (UK and US) and may occasionally be called upon for major incidents …/SLOs.Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.Partner with Security teams to ensure our platforms meet compliance, security, and risk‐management expectations.Lead seamless handovers ...

Platform Support Lead

Location
Greater London, England, United Kingdom
week) Employment: Fulltime Role Summary: The lead will act as the primary onshore contact for IDMC and TIBCO EBX support, coordinating platform operations, incident management, deployments, service transition, offshore delivery, and client communication. Key Responsibilities: Lead daily monitoring of IDMC and EBX platform health, Secure Agents, metadata scans … data-quality workflows, integrations, and scheduled jobs. Own high-priority incident coordination, perform technical triage, drive restoration, manage escalations, and support root-cause analysis. Coordinate Dev, UAT, and Production activities, including Secure Agent configuration, SSO/access issues, deployments, rollback planning, and post-deployment validation. Provide functional and technical ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve performance, availability, and reliability issues Lead and contribute to incident response, troubleshooting, and root cause analysis Automate operational processes and eliminate repetitive manual tasks Work closely with software engineers to improve deployment processes, system … understanding of metrics, logging, tracing, alerting, and system health Experience troubleshooting complex production environments Understanding of SLIs, SLOs, SLAs, and error budgets Experience with incident management and root cause analysis Good understanding of cloud networking, security, and infrastructure fundamentals Strong scripting/automation skills A strong understanding ...

Production Engineer – Trading & Electronic Trading Systems

Location
Greater London, England, United Kingdom
live trading and electronic trading platforms. These environments are production‐critical, operate in real time and require strong ownership of system stability, performance and incident management. The role involves close interaction with traders, developers, IT support and infrastructure teams. Role Overview We are looking for a Front Office Production … Responsibilities Production Ownership & Reliability Operate and support Front Office trading systems in production Ensure high availability, stability and performance during market hours Own incident management from detection to resolution and root cause analysis Participate in on‐call or production support rotations when required Contribute to post‐incident ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
distributed and event-driven systems using messaging and asynchronous processing patterns. Drive engineering excellence through architecture, coding standards, testing, and automation. Support production systems, incident management, troubleshooting, and operational resilience. Balance long-term architecture goals with short-term delivery priorities. Collaborate across engineering teams, product owners, architects … production support. Non-Technical Skills Strong leadership, ownership, and problem-solving mindset. Ability to thrive in ambiguous and rapidly changing environments. Excellent stakeholder management, communication, and collaboration skills. Proven ability to influence engineering practices beyond immediate teams. Experience coaching and developing engineers while driving continuous improvement. If you like ...

Site Reliability Engineer

Hiring Organisation
Trainline
Location
London, UK
Employment Type
Full-time
cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices. The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery … Trainline, you'll be working on...Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platformParticipating in production incident response, supporting investigation, mitigation, communication, and coordinated service restorationContributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilienceTaking part ...

Onshoring - Capacity Manager

Hiring Organisation
Hays Technology
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600/day Up to £600pd inside IR35 via umbrella
Identify potential capacity constraints and performance risks before they impact users or services. Coordinate capacity planning activities with technical and operational teams. Support major incident management by assessing capacity and resilience impacts. Contribute to service recovery planning where capacity and performance are affected. Produce demand forecasts and growth … related risks and issues to operational and leadership stakeholders. Maintain capacity dashboards, reports, and continuous improvement plans. Key Skills & Experience Proven experience in Capacity Management within a complex IT environment. Strong understanding of infrastructure, application performance, and service management principles. Experience producing capacity forecasts, utilisation reports, and risk ...

Software Developer - Risk Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
stability and availability of the Risk software systems, as well as their day-to-day operations. Squarepoint's Risk platform is responsible for position management, profit/loss computation, inventory/locate management and internal order routing. These critical systems need to be performant, resilient, and capable … business hours, people on-duty will prioritise responding to incidents over their project work. On average, people are on-duty one day per week. Incident management: Root cause analyses are performed to understand the source of incidents and to raise appropriate remedial actions. Day-to-day operations: Until ...

Technical Support Engineer

Location
Greater London, England, United Kingdom
issues through to resolution Present complex technical information to non-technical audiences Provide accurate and complete problem resolution documentation for future reference and management reporting Take part in the creation and maintenance of knowledge base data Increase subject matter knowledge on Medidata products Work with other Medidata teams … experience querying and modifying data using MYSQL, MS SQL Server, MongoDB, Snowflake Experience with C#, Javascript, Visual Studio, Ruby, Python and ASP.NET Experience using Incident Management and/or Project Management software Experience with multiple integrated software products/services Experience working with customers to resolve complex ...

Trainee DevOps Engineer | No experience needed (Ref: 7501)

Hiring Organisation
Qualify Nation Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£28,000 - £38,000 per annum
Continuous Deployment (CD) Infrastructure as Code (IaC) Containerisation with Docker Container Orchestration Cloud Computing Fundamentals Cloud Platforms (AWS, Microsoft Azure and Google Cloud) Configuration Management Monitoring and Logging Security Best Practices (DevSecOps) Networking Fundamentals Automation and Scripting Incident Management and Reliability Engineering Practical Experience You will work ...

Network Engineer

Hiring Organisation
Appcast
Location
London, UK
tasks using Python, Ansible, Jinja2 and GitSupporting and extending Infrastructure-as-Code and CI/CD practices for network provisioningOwning break/fix and incident response for complex enterprise and campus network issuesDriving incident management practices, including on-call rotation, escalation procedures, postmortems and root-cause analysisImplementing ...

People Systems Analyst (6 month FTC)

Location
Greater London, England, United Kingdom
Benefits or Payroll is desirable. Comfortable analysing and resolving system and process issues in a structured manner. Strong organisational skills, with experience in operating incident management tools. Ability to prioritise and manage multiple requests within agreed service levels. Familiarity with Microsoft Office suite. Experience using ServiceNow and JIRA … work well within a global and sometimes virtual team. Understanding of HR data governance, confidentiality requirements and system administration best practices. Good stakeholder management skills, with the ability to communicate clearly with technical and non-technical audiences. Ability to handle sensitive people data with discretion and in line with ...

Application Support Analyst | Operations Technology

Location
Greater London, England, United Kingdom
business users offering second- and third-line support. Knowledge of scripting language (such as PowerShell, Python). Manage new system analysis and implementation. Provide incident management per ITIL standards. Liaison between technology department and business groups to communicate system changes. Manage process and trading system documentation in existing … report any breaches of policy to Compliance and/or your supervisor as required To elevate risk events immediately To provide input to risk management processes, as required. Competencies , Skills and Experience Skills and Experience: Essential: Solid background in Windows, Linux/Unix OS, including one of the following ...

Application Support Analyst | Energy Trading Operations

Location
Greater London, England, United Kingdom
Support business users offering second‐ and third‐line support. Knowledge of scripting language (PowerShell, Python, etc.). Manage new system analysis and implementation. Provide incident management per ITIL standards. Liaison between technology department and business groups to communicate system changes. Manage process and trading system documentation in existing … responsibility. Report any breaches of policy to Compliance and/or your supervisor as required. Escalate risk events immediately. Provide input to risk management processes, as required. Competencies, Skills and Experience Competencies: A collaborative team player, approachable, self‐efficient and influences a positive work environment. Demonstrates curiosity. Resilient ...

Telecoms Specialist Engineer

Hiring Organisation
Flotek
Location
London, United Kingdom
Salary
30000
cannot offer sponsorship or relocation assistance, so you must already have the right to work in the UK. What you will be doing: Incident management: Be the first point of contact for telecoms incidents, logging, categorising and prioritising faults in line with ITIL best practice. Service level management … customers and stakeholders, managing expectations during incidents. On-call support: Support telecoms services during scheduled weekend and out-of-hours cover, following major incident escalation paths. What we are looking for: Essential Requirements: You will be a logical problem-solver with a customer-first mindset who thrives ...

Network Engineer

Location
Greater London, England, United Kingdom
Python, Ansible, Jinja2 and Git Supporting and extending Infrastructure-as-Code and CI/CD practices for network provisioning Owning break/fix and incident response for complex enterprise and campus network issues Driving incident management practices, including on-call rotation, escalation procedures, postmortems and root-cause ...