326 to 350 of 484 Incident Management Jobs in London

Azure Data Support Engineer

Hiring Organisation
Avanade
Location
London, UK
Employment Type
Full-time
solutions. Develop and implement self-healing automation for recurring failures. Service Operations & Support (Managed Services)Provide L2/L3 support aligned with ITIL practices (incident, problem, change management).Participate in on-call rotations and handle critical incident response. Maintain detailed SOPs, runbooks, knowledge base articles, and client … Data WarehousingSQL skills for debugging, data validation, and optimizationExperience with Azure SQL DB, or SQL ServerFamiliarity with data modeling concepts and warehouse performance tuningSupport & Incident ManagementStrong troubleshooting and analytical skills for root cause analysisExposure to ITSM tools (e.g., ServiceNow, Jira)Preferred Qualifications: Microsoft Certifications (e.g. ...

Senior AI Quality Engineer

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
board to solve quality problems togetherExperience shaping CI/CD pipelines for test reliability (GitHub Actions or similar), including flakiness reduction, parallelisation, artefact management, and runtime controlProactive, low-ego, and clear communicator, able to chase things down when blockedAble to operate independently and drive improvements end-to-end (frameworks … test data, reporting)Can point to 3-4 projects where you moved the needle on quality - for example improving SLO coverage or reducing incident rate - and can speak to the measurable impact of eachNice to haveHave worked with an incident management process to root cause issues with ...

Application Security Engineer

Hiring Organisation
The Pokémon Company
Location
London, UK
Employment Type
Full-time
know The Pokémon Company InternationalThe Pokémon Company International manages the Pokémon property outside of Asia and is responsible for brand management, licensing and marketing, the Pokémon Trading Card Game, the animated TV series, home entertainment, and the official Pokémon website. Pokémon was launched in Japan in 1996 and today … facing vulnerabilities. Collaborate with Information Security partners to advance roadmaps and shared initiatives. Integrate security testing and controls into development lifecycle phases. Provide management with insights on business impact related to data compromise or system unavailability. Work within an agile framework to deliver iterative improvements toward team and organizational ...

Integration Engineering Manager

Location
Greater London, England, United Kingdom
copilots, automation platforms, and intelligent workflows to securely discover, access, and act on enterprise systems and data. MuleSoft AI Gateway and secure AI traffic management Model Context Protocol (MCP) enablement for enterprise APIs and tools MuleSoft Agent Fabric adoption, including tool discovery, agent governance, and agent-to-system orchestration … Anypoint Monitoring, Exchange, and CI/CD practices for AI-enabled integration delivery. Drive best practices for error handling, retries, idempotency, throttling, resiliency, traffic management, and graceful failure modes. Partner with SRE and DevOps teams on observability, capacity planning, incident management, and operational runbooks for agent-driven ...

Senior/Lead SAP FICO Consultant

Location
Greater London, England, United Kingdom
other teams. Provide expert guidance on S/4HANA conversion, Universal Journal, New Asset Accounting, and Margin Analysis setups. 2. Client Advisory & Stakeholder Management Act as the trusted advisor to CFOs, Finance Directors, and business leads. Translate complex business requirements into strategic SAP solutions. Present solution options, impact assessments … ensure adherence to consulting methodologies. Review functional specs, design documents, solution gaps, and enhancement proposals. 5. Support & Continuous Improvement Provide leadership for L3 incident management and complex issue resolution. Recommend continuous process improvements and roadmap enhancements using SAP innovations. Lead post-implementation reviews to ensure business adoption ...

Support Engineer - £50K - Remote

Location
City Of London, England, United Kingdom
operational health of the Digital team’s products and end user support and project work will cover security and IAM, incident management and resolution and even automation. Requirements: Good, broad IT support and service management experience Experience of operational support, DevOps and/or platform AWS/ ...

Service Systems Engineer

Hiring Organisation
ARM
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
responsibilities Provide advanced troubleshooting and support for live CAD systems in mission-critical public safety environments. Investigate and resolve complex CAD application issues affecting incident management, mobilisation, mapping, geospatial services and external interfaces. Lead technical investigations during major incidents and act as an escalation point for operational issues. … third-party systems. Work directly with customers during incidents, service reviews and technical workshops. Perform root cause analysis (RCA) and contribute to problem management activities. Develop and implement changes while ensuring technical governance and service continuity. Provide technical leadership and recommendations during service-impacting incidents and customer escalations. About ...

Service Systems Engineer

Hiring Organisation
ARM
Location
Twickenham, London, St. Margarets and North Twickenham, United Kingdom
Employment Type
Permanent
responsibilities Provide advanced troubleshooting and support for live CAD systems in mission-critical public safety environments. Investigate and resolve complex CAD application issues affecting incident management, mobilisation, mapping, geospatial services and external interfaces. Lead technical investigations during major incidents and act as an escalation point for operational issues. … third-party systems. Work directly with customers during incidents, service reviews and technical workshops. Perform root cause analysis (RCA) and contribute to problem management activities. Develop and implement changes while ensuring technical governance and service continuity. Provide technical leadership and recommendations during service-impacting incidents and customer escalations. About ...

Senior Software Engineer I

Location
Greater London, England, United Kingdom
Product Vision: Partner with Product Managers to prioritize features, translate client requirements into technical specifications, and ensure solutions align with our overarching product vision. Incident Management: Confidently own the resolution of incidents, coordinating mitigation, fixes, and post-mortem artifacts. Collaboration & Mentorship Team Growth: Mentor and support less experienced … engineers, helping them level up technically and adapt to Thought Machine’s engineering philosophy. Stakeholder Management: Translate highly technical problems and architectural decisions into clear concepts for non-technical stakeholders. What We’re Looking For Technical Expertise: Strong proficiency in modern backend languages (ideally Go or Python) and experience ...

Mandarin speaking IT Support

Hiring Organisation
People First
Location
London, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
What You'll be Doing: Carry out the Help Desk first line support duty for IT Centre (Europe) including Help Desk cases, user ID management, production access and Head Office’s service requests and manage the workflow of these requests Carry out machine room environment and devices first line … maintenance, support second line engineer to perform system changes and maintenance Monitor the production systems and applications and follow the incident management procedure to deal with critical messages. Perform first line handling of system and network alerts Carry out daily operational duty and run batch jobs Assist with ...

Forward Deployed Infrastructure Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
permissioning, networking, and infrastructure provisioning. Manage the full infrastructure lifecycle for some of Sierra's most critical customers: initial deployment, ongoing upgrades, scaling, and incident support. Develop and maintain deployment runbooks, upgrade procedures, and automation tooling to make operations repeatable and reliable. Partner with engineering teams to ensure Sierra … VPCs, IAM, DNS, load balancing). Track record of managing deployments in customer-owned cloud environments, including navigating security reviews, compliance requirements, and change management processes. Ability to navigate ambiguity with enterprise stakeholders - from platform engineers to CISOs - and turn infrastructure conversations into actionable deployment plans. Strong operational instincts ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
auto-remediate global SaaS infrastructure. What you'll do Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated … quality guardrails. Reliability & Scalability Engineering: Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews (PIRs). Mentorship & Collaboration: Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams. ...

Forward Deployed Infrastructure Engineer - Spanish speaking

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
permissioning, networking, and infrastructure provisioning. Manage the full infrastructure lifecycle for some of Sierra's most critical customers: initial deployment, ongoing upgrades, scaling, and incident support. Develop and maintain deployment runbooks, upgrade procedures, and automation tooling to make operations repeatable and reliable. Partner with engineering teams to ensure Sierra … VPCs, IAM, DNS, load balancing). Track record of managing deployments in customer-owned cloud environments, including navigating security reviews, compliance requirements, and change management processes. Ability to navigate ambiguity with enterprise stakeholders - from platform engineers to CISOs - and turn infrastructure conversations into actionable deployment plans. Strong operational instincts ...

Forward Deployed Infrastructure Engineer (Spanish speaking)

Location
Greater London, England, United Kingdom
networking, and infrastructure provisioning. Manage the full infrastructure lifecycle for some of the company’s most critical customers: initial deployment, ongoing upgrades, scaling, and incident support. Develop and maintain deployment runbooks, upgrade procedures, and automation tooling to make operations repeatable and reliable. Partner with engineering teams to ensure … VPCs, IAM, DNS, load balancing). Track record of managing deployments in customer-owned cloud environments, including navigating security reviews, compliance requirements, and change management processes. Ability to navigate ambiguity with enterprise stakeholders — from platform engineers to CISOs — and turn infrastructure conversations into actionable deployment plans. Strong operational instincts ...

Service Operations Lead

Location
Greater London, England, United Kingdom
behaviours, while deliberately and consistently driving high performance as you coach and nurture talent. With practical, hands on expertise in Dynatrace, you'll drive incident management providing leadership and communication to resolution. What will you bring to the role? Strong, visible people leader who role models, cares, connects … role, leading alerts and monitoring for several services, using Dynatrace. Experience in deploying, monitoring and troubleshooting large scale distributed and cloud solutions. ITIL Service Management Framework within complex IT environments. Self starter, with strong stakeholder management skills and the ability to influence at all levels. In return ...

Shift Team Leader

Hiring Organisation
Dynamic Resourcing
Location
Wembley, London, United Kingdom
Employment Type
Permanent
Salary
£55,000
function Likely to have 5 years business experience and/or be fully qualified within a mechanical or electrical trade Significant M&E Operations Management experience within a business critical client environment HV or LV Authorised Person Requirement to complete the critical environment training, and to carry out scenario … generation and control systems Experience with working within a critical and secure environment Shift Leader Responsibilities Establish good working relationships with the client Facilities Management Team, to ensure proactive, open and co-operative partnership Monitor component, system and building performance against target profiles via Building Management System ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software products. … Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using configuration management platforms like Ansible, Puppet, or Chef Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer ...

Lead Software Engineer - Platform

Location
Greater London, England, United Kingdom
others to improve our platform infrastructure and shape how we build reliable, scalable systems at Trint. This is a technical leadership role, not a management one. You won't have direct reports, but you will have significant influence. You'll make decisions that affect how the whole engineering team … infrastructure. You’ll be responsible for the reliability, scalability, and operability of our systems. That means CI/CD pipelines, infrastructure‐as‐code, observability, incident response, and the day‐to‐day health of production. You’ll make sure we can ship with confidence and sleep at night. Shape technical ...

Mandarin speaking Job - IT Support - iw

Hiring Organisation
People First Recruitment
Location
London, UK
Employment Type
Full-time
What You'll be Doing: Carry out the Help Desk first line support duty for IT Centre (Europe) including Help Desk cases, user ID management, production access and Head Office's service requests and manage the workflow of these requestsCarry out machine room environment and devices first line maintenance … support second line engineer to perform system changes and maintenanceMonitor the production systems and applications and follow the incident management procedure to deal with critical messages. Perform first line handling of system and network alertsCarry out daily operational duty and run batch jobsAssist with producing system and network ...

Machine Learning Operations Engineer

Location
Greater London, England, United Kingdom
engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response ML lifecycle and platform engineering Build repeatable workflows for model training, validation, promotion, deployment and retraining Productionise models through packaging, versioning, model registry integration … automation, observability and infrastructure as code Write production-grade Python for long-running services, deployment tooling and ML workflows Establish testing, validation, release and incident-management practices for ML systems Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance ...

Senior AV Technician - EMEA

Location
London, United Kingdom
users, stakeholders and other service partners. Monitor service performance and contribute to operational reviews. Produce documentation including SOPs, technical guidance and training material. Support incident management and ensure issues are followed through to resolution. Assist with onboarding, compliance and administration relating to technical resources. Support operational reporting, purchasing … events, alongside the ability to troubleshoot technical issues confidently. Previous experience supervising or leading AV Technicians is important, as the position includes direct line-management responsibilities. You will also need excellent communication and stakeholder-management skills, a proactive approach to service improvement and the ability to work effectively ...

Staff Software Engineer - Commercial Planning

Location
Greater London, England, United Kingdom
engineers at all levels, sharing knowledge and helping develop technical capability across the wider engineering community. Drive operational excellence through monitoring, observability, alerting and incident management practices, ensuring learnings from production environments are fed back into development. Stay informed of emerging technologies, AI capabilities and engineering trends, proposing … tooling. Experience designing and implementing scalable cloud-based applications with a strong focus on performance, security and operational excellence. Strong communication and stakeholder management skills, with the ability to explain complex technical concepts to both technical and non-technical audiences. A mentoring mindset with a genuine passion for developing ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
maintain and debug reproducible multi-language CI pipelines, and optimize CI performance across large compute clusters. Build and maintain infrastructure observability, alerting, runbooks, and incident response workflows for Fractile's infrastructure. IaC TODO Scale and maintain Fractile's Bazel monorepo as we continue growing across Python, C++, Rust, SystemVerilog … Pulumi) Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on‐call practice Strong proficiency in at least one systems or scripting language used for tooling (Python preferred; Go or Rust also ...

Platform Operations Lead - Hybrid Cloud & Resilience

Location
Greater London, England, United Kingdom
Gravitas Group, a leading financial services firm in London, is seeking a Senior IT Operations Manager to lead delivery, resilience and operational management of enterprise infrastructure across cloud and on-prem environments. The role requires proven leadership of large teams, strong background in hybrid cloud, automation and incident management, and experience with trading or asset management environments. #J-18808-Ljbffr ...

Senior Business Solutions Analyst - Intapp (VA771)

Location
Greater London, England, United Kingdom
solutions to address business challenges Work to Carey Olsen’s design standards and suggest relevant industry standards where appropriate Recommend best practices for release management, coding standards, and code reviews Collaborate with internal teams to ensure quality and scalability Train Business Solutions team members on Intapp so they … incorporating them into solutions and system configurations Mentor junior Technology Team members and provide support when needed Follow priorities set by Group requirements and management Troubleshoot technical issues and respond to support requests. Analyse recurring problems and propose permanent fixes Deliver excellent customer service and represent the Technology department ...