326 to 350 of 496 Incident Management Jobs in London

Senior Backend Engineer - Payments

Location
Greater London, England, United Kingdom
users Take ownership of product development, from feature discovery, to the breakdown of work, and its implementation End-to-end application support, including production incident management Embrace agile methodologies and user-centred thinking Engage in a culture of continuous improvement by attending events such as blameless post-mortems ...

Site Reliability Engineer

Hiring Organisation
GoCardless
Location
London, UK
Employment Type
Full-time
Building new platform components, from initial design to deployment. Operational Support: Ensuring the health and stability of our systems through on-call rotations and incident management. Business as Usual (BAU): Continuous maintenance, improvements, and support for audits and compliance. Tech stack and tools: Python, Ruby, Golang;Terraform;Atlantis ...

Staff Engineer (Architect) - Researcher Operations

Location
Greater London, England, United Kingdom
pairing using tools like Git and GitHub. Experience of operationally managing software components once live, including observability, logging, metrics, error reporting, debugging and live incident management. Experience of working with sensitive personal data. Competencies Experience working in/with cross‐functional teams consisting of engineers, product ...

PROJECT LEAD L1(CONTRACT)

Location
Greater London, England, United Kingdom
Lead and oversee complex IT projects focused on Data Center and Infrastructure Migrations aligned with business goals. Manage project governance including budget control, risk management, scope maintenance, and delivery timelines. Coordinate project deliverables and proactively identify and resolve risks and issues. Engage with diverse stakeholders, facilitating decision-making … securing approvals at critical milestones. Optimize resource allocation and enhance team productivity. Track progress using advanced project management tools and report outcomes to senior management. Ensure compliance with organizational policies and industry best practices. Requirements Experience: 5–10 years in IT project management, specifically in Infrastructure and Data ...

Desktop Support Engineer - Quant Trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Desktop Engineer will: Manage raised requests and incidents in a timely manner following well-rehearsed internal request runbooks/knowledge based and incident management and request fulfillment processes;Partner with Level 2 & 3 internal and external teams/vendors as appropriate to help troubleshoot more complex incidents;Provide … alert management to pro-active resolve failure alerts, degraded services and unavailable services through internal monitoring tools;You will need: Detailed knowledge and experience in a Microsoft environment supporting Windows 10, Windows Server 2012+, Office Suite including Outlook, Office 365 and Azure, including OneDrive, DLP and SharePointDetailed knowledge ...

Principal Data Engineer

Location
Greater London, England, United Kingdom
authority for our clients, ensuring their data platforms are resilient, scalable, and aligned with their long-term business goals. Key Responsibilities Customer Excellence & Stakeholder Management: Act as the primary strategic partner to client business owners, aligning technical data solutions with corporate goals while championing reliability and \"clean data … Execution: Oversee \"team-of-teams\" delivery performance, ensuring the quality and resilience of high-impact data pipelines while managing capacity planning and team velocity. Incident Management & Resilience: Lead the triage and resolution of critical platform incidents, taking full ownership of root-cause prevention and maintaining transparent communication with ...

Azure Data Support Engineer

Hiring Organisation
Avanade
Location
London, UK
Employment Type
Full-time
solutions. Develop and implement self-healing automation for recurring failures. Service Operations & Support (Managed Services)Provide L2/L3 support aligned with ITIL practices (incident, problem, change management).Participate in on-call rotations and handle critical incident response. Maintain detailed SOPs, runbooks, knowledge base articles, and client … Data WarehousingSQL skills for debugging, data validation, and optimizationExperience with Azure SQL DB, or SQL ServerFamiliarity with data modeling concepts and warehouse performance tuningSupport & Incident ManagementStrong troubleshooting and analytical skills for root cause analysisExposure to ITSM tools (e.g., ServiceNow, Jira)Preferred Qualifications: Microsoft Certifications (e.g. ...

Senior AI Quality Engineer

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
board to solve quality problems togetherExperience shaping CI/CD pipelines for test reliability (GitHub Actions or similar), including flakiness reduction, parallelisation, artefact management, and runtime controlProactive, low-ego, and clear communicator, able to chase things down when blockedAble to operate independently and drive improvements end-to-end (frameworks … test data, reporting)Can point to 3-4 projects where you moved the needle on quality - for example improving SLO coverage or reducing incident rate - and can speak to the measurable impact of eachNice to haveHave worked with an incident management process to root cause issues with ...

Application Security Engineer

Hiring Organisation
The Pokémon Company
Location
London, UK
Employment Type
Full-time
know The Pokémon Company InternationalThe Pokémon Company International manages the Pokémon property outside of Asia and is responsible for brand management, licensing and marketing, the Pokémon Trading Card Game, the animated TV series, home entertainment, and the official Pokémon website. Pokémon was launched in Japan in 1996 and today … facing vulnerabilities. Collaborate with Information Security partners to advance roadmaps and shared initiatives. Integrate security testing and controls into development lifecycle phases. Provide management with insights on business impact related to data compromise or system unavailability. Work within an agile framework to deliver iterative improvements toward team and organizational ...

Integration Engineering Manager

Location
Greater London, England, United Kingdom
copilots, automation platforms, and intelligent workflows to securely discover, access, and act on enterprise systems and data. MuleSoft AI Gateway and secure AI traffic management Model Context Protocol (MCP) enablement for enterprise APIs and tools MuleSoft Agent Fabric adoption, including tool discovery, agent governance, and agent-to-system orchestration … Anypoint Monitoring, Exchange, and CI/CD practices for AI-enabled integration delivery. Drive best practices for error handling, retries, idempotency, throttling, resiliency, traffic management, and graceful failure modes. Partner with SRE and DevOps teams on observability, capacity planning, incident management, and operational runbooks for agent-driven ...

Senior/Lead SAP FICO Consultant

Location
Greater London, England, United Kingdom
other teams. Provide expert guidance on S/4HANA conversion, Universal Journal, New Asset Accounting, and Margin Analysis setups. 2. Client Advisory & Stakeholder Management Act as the trusted advisor to CFOs, Finance Directors, and business leads. Translate complex business requirements into strategic SAP solutions. Present solution options, impact assessments … ensure adherence to consulting methodologies. Review functional specs, design documents, solution gaps, and enhancement proposals. 5. Support & Continuous Improvement Provide leadership for L3 incident management and complex issue resolution. Recommend continuous process improvements and roadmap enhancements using SAP innovations. Lead post-implementation reviews to ensure business adoption ...

Support Engineer - £50K - Remote

Location
City Of London, England, United Kingdom
operational health of the Digital team’s products and end user support and project work will cover security and IAM, incident management and resolution and even automation. Requirements: Good, broad IT support and service management experience Experience of operational support, DevOps and/or platform AWS/ ...

Service Systems Engineer

Hiring Organisation
ARM
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
responsibilities Provide advanced troubleshooting and support for live CAD systems in mission-critical public safety environments. Investigate and resolve complex CAD application issues affecting incident management, mobilisation, mapping, geospatial services and external interfaces. Lead technical investigations during major incidents and act as an escalation point for operational issues. … third-party systems. Work directly with customers during incidents, service reviews and technical workshops. Perform root cause analysis (RCA) and contribute to problem management activities. Develop and implement changes while ensuring technical governance and service continuity. Provide technical leadership and recommendations during service-impacting incidents and customer escalations. About ...

Service Systems Engineer

Hiring Organisation
ARM
Location
Twickenham, London, St. Margarets and North Twickenham, United Kingdom
Employment Type
Permanent
responsibilities Provide advanced troubleshooting and support for live CAD systems in mission-critical public safety environments. Investigate and resolve complex CAD application issues affecting incident management, mobilisation, mapping, geospatial services and external interfaces. Lead technical investigations during major incidents and act as an escalation point for operational issues. … third-party systems. Work directly with customers during incidents, service reviews and technical workshops. Perform root cause analysis (RCA) and contribute to problem management activities. Develop and implement changes while ensuring technical governance and service continuity. Provide technical leadership and recommendations during service-impacting incidents and customer escalations. About ...

Senior Software Engineer I

Location
Greater London, England, United Kingdom
Product Vision: Partner with Product Managers to prioritize features, translate client requirements into technical specifications, and ensure solutions align with our overarching product vision. Incident Management: Confidently own the resolution of incidents, coordinating mitigation, fixes, and post-mortem artifacts. Collaboration & Mentorship Team Growth: Mentor and support less experienced … engineers, helping them level up technically and adapt to Thought Machine’s engineering philosophy. Stakeholder Management: Translate highly technical problems and architectural decisions into clear concepts for non-technical stakeholders. What We’re Looking For Technical Expertise: Strong proficiency in modern backend languages (ideally Go or Python) and experience ...

Mandarin speaking IT Support

Hiring Organisation
People First
Location
London, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
What You'll be Doing: Carry out the Help Desk first line support duty for IT Centre (Europe) including Help Desk cases, user ID management, production access and Head Office’s service requests and manage the workflow of these requests Carry out machine room environment and devices first line … maintenance, support second line engineer to perform system changes and maintenance Monitor the production systems and applications and follow the incident management procedure to deal with critical messages. Perform first line handling of system and network alerts Carry out daily operational duty and run batch jobs Assist with ...

Forward Deployed Infrastructure Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
permissioning, networking, and infrastructure provisioning. Manage the full infrastructure lifecycle for some of Sierra's most critical customers: initial deployment, ongoing upgrades, scaling, and incident support. Develop and maintain deployment runbooks, upgrade procedures, and automation tooling to make operations repeatable and reliable. Partner with engineering teams to ensure Sierra … VPCs, IAM, DNS, load balancing). Track record of managing deployments in customer-owned cloud environments, including navigating security reviews, compliance requirements, and change management processes. Ability to navigate ambiguity with enterprise stakeholders - from platform engineers to CISOs - and turn infrastructure conversations into actionable deployment plans. Strong operational instincts ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
auto-remediate global SaaS infrastructure. What you'll do Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated … quality guardrails. Reliability & Scalability Engineering: Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews (PIRs). Mentorship & Collaboration: Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams. ...

Forward Deployed Infrastructure Engineer - Spanish speaking

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
permissioning, networking, and infrastructure provisioning. Manage the full infrastructure lifecycle for some of Sierra's most critical customers: initial deployment, ongoing upgrades, scaling, and incident support. Develop and maintain deployment runbooks, upgrade procedures, and automation tooling to make operations repeatable and reliable. Partner with engineering teams to ensure Sierra … VPCs, IAM, DNS, load balancing). Track record of managing deployments in customer-owned cloud environments, including navigating security reviews, compliance requirements, and change management processes. Ability to navigate ambiguity with enterprise stakeholders - from platform engineers to CISOs - and turn infrastructure conversations into actionable deployment plans. Strong operational instincts ...

Forward Deployed Infrastructure Engineer (Spanish speaking)

Location
Greater London, England, United Kingdom
networking, and infrastructure provisioning. Manage the full infrastructure lifecycle for some of the company’s most critical customers: initial deployment, ongoing upgrades, scaling, and incident support. Develop and maintain deployment runbooks, upgrade procedures, and automation tooling to make operations repeatable and reliable. Partner with engineering teams to ensure … VPCs, IAM, DNS, load balancing). Track record of managing deployments in customer-owned cloud environments, including navigating security reviews, compliance requirements, and change management processes. Ability to navigate ambiguity with enterprise stakeholders — from platform engineers to CISOs — and turn infrastructure conversations into actionable deployment plans. Strong operational instincts ...

Service Operations Lead

Location
Greater London, England, United Kingdom
behaviours, while deliberately and consistently driving high performance as you coach and nurture talent. With practical, hands on expertise in Dynatrace, you'll drive incident management providing leadership and communication to resolution. What will you bring to the role? Strong, visible people leader who role models, cares, connects … role, leading alerts and monitoring for several services, using Dynatrace. Experience in deploying, monitoring and troubleshooting large scale distributed and cloud solutions. ITIL Service Management Framework within complex IT environments. Self starter, with strong stakeholder management skills and the ability to influence at all levels. In return ...

Shift Team Leader

Hiring Organisation
Dynamic Resourcing
Location
Wembley, London, United Kingdom
Employment Type
Permanent
Salary
£55,000
function Likely to have 5 years business experience and/or be fully qualified within a mechanical or electrical trade Significant M&E Operations Management experience within a business critical client environment HV or LV Authorised Person Requirement to complete the critical environment training, and to carry out scenario … generation and control systems Experience with working within a critical and secure environment Shift Leader Responsibilities Establish good working relationships with the client Facilities Management Team, to ensure proactive, open and co-operative partnership Monitor component, system and building performance against target profiles via Building Management System ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Appcast
Location
London, UK
environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software products. … like Loki, Mimir, and TempoBackground in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or RundeckExperience using configuration management platforms like Ansible, Puppet, or ChefProfessional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software products. … Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using configuration management platforms like Ansible, Puppet, or Chef Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer ...

ML Ops Engineer

Hiring Organisation
Appcast
Location
London, UK
engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response.What you’ll work onML lifecycle and platform engineeringBuild repeatable workflows for model training, validation, promotion, deployment and retraining.Productionise models through packaging, versioning, model registry … efficiency through automation, observability and infrastructure as code.Write production-grade Python for long-running services, deployment tooling and ML workflows.Establish testing, validation, release and incident-management practices for ML systems.Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance.Make explicit ...