301 to 325 of 496 Incident Management Jobs in London

Cloud Platforms Engineer

Location
City Of London, England, United Kingdom
than a queue of requests. Keep production healthy. Help diagnose and resolve production issues when they arise, bring a clear view of what good incident management looks like, and fix the class of problem rather than just the instance. Make sure we can recover. Look after backup … Azure or Google Cloud Platform experience is fine. Practical experience with infrastructure as code in a team setting, Terraform preferred – modules, code review, state management and applying through a pipeline, rather than clicking in a console. Desirable: experience of cloud identity and access at organisation scale, and single sign ...

Forward Deployed Engineering Lead - Global Broking

Hiring Organisation
Liquidnet
Location
London, UK
Employment Type
Full-time
trusted relationships with desks and stakeholders, ensuring priorities, progress, and risks are clearly communicated. Deliver within TP ICAP's engineering, governance, security, and change management frameworks. Experience/CompetencesEssentialExperience leading people, engineering teams, or owning the delivery of an engineering squad. Strong stakeholder management, communication, and judgement. Ability … deliver in fast-moving environments with evolving requirements. Experience operating production services, including deployment, monitoring, incident management, and support. Strong hands-on engineering skills in Python and TypeScript, with the ability to review code, set technical direction, and guide implementation. Experience using modern software development tools and practices ...

Lead Software Engineer- Python / Java- (Cloud Data Platform — AWS/Databricks/Terraform)

Hiring Organisation
Appcast
Location
London, UK
practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across … platform engineering capability, including CI/CD design, automated security/quality gates, observability (logs/metrics/traces), SLOs/SLAs, and incident management/RCA in regulated environments.Proven end-to-end delivery leadership for large, cross-functional initiatives (roadmaps, dependency management, stakeholder alignment, budgeting/ ...

DevOps Engineer x 5

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £552.00 per day
capabilities Troubleshoot and resolve infrastructure, platform and service issues Monitor platform health, availability and performance Contribute to infrastructure improvements and operational resilience Support incident management and service recovery activities Produce and maintain technical documentation Collaborate with engineering teams to improve deployment processes and service reliability Participate in support ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. * Support live service operations, incident management, troubleshooting and platform resilience. * Collaborate with engineers, architects, product teams and stakeholders whilst coaching and mentoring colleagues. Essential Skills for the Senior ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. * Support live service operations, incident management, troubleshooting and platform resilience. * Collaborate with engineers, architects, product teams and stakeholders whilst coaching and mentoring colleagues. Essential Skills for the Senior ...

Senior Engineering Manager — Hotels.com Technology

Location
Greater London, England, United Kingdom
quality and reliability in the shopping funnel: set engineering quality and operational excellence standards, and ensure robust CI/CD, monitoring, alerting, and incident response practices are consistently applied. Move fast with incomplete information — make well‐reasoned decisions without waiting for certainty or broad consensus, and validate assumptions through … domain contexts. Minimum Qualifications: Bachelor's Degree or Equivalent Level; Technical Degree Preferred. 8+ years of relevant professional experience and 3+ years of people management experience. Experience managing engineering teams with end‐to‐end ownership of services or domains, including design, implementation, deployment, and ongoing operations for production services ...

AVP, Automation & AI Delivery

Hiring Organisation
Arch Capital Group
Location
London, UK
Employment Type
Full-time
success metrics and determine feasibilityOwn, Track and report on delivery scope, progress, risks and dependenciesEnsure delivery standards for requirements, documentation, testing, releases, monitoring and incident management are metImplement feedback loops and metrics to track and report on adoption, trust, and impact to SA leadership and senior stakeholdersCollaborate with … standards and approvals (permitted use, privacy, security, model governance and third-party risk).Role Requirements & SkillsSkills/CompetenciesHighly organised with strong programme/project management skills; able to manage multiple initiatives concurrently and deliver results. Exceptional stakeholder management with the ability to build trusting relationships and drive consensus ...

Data Operations Specialist

Location
Greater London, England, United Kingdom
initiatives that improve data quality, resilience, governance, and scalability, while acting as a trusted partner to business leaders and technology teams.**Responsibilities*** Ongoing Data Management: Own and optimize Data Operations pipelines, with deep expertise in cleansing, and maintaining corporate structures and company- data. Establish best practices and standards … ongoing data management.* Incident Management & Resolution: Serve as the escalation point for complex or high-impact data issues. Lead root cause analysis, post-mortems, and the design of long-term preventative solutions to reduce operational risk.* Knowledge Leadership: Act as a subject-matter expert in company-level data ...

Cyber Risk & Security Manager

Location
Greater London, England, United Kingdom
governance advisory services, security maturity programmes, and SOC strategy engagements. The role requires a highly capable security professional with experience across risk advisory, vulnerability management, security operations design, compliance frameworks, and executive reporting. The successful candidate will operate at both strategic and technical levels, supporting board-level decision‐making … resilience simulations. Produce executive‐level cyber risk reports and mitigation roadmaps. Provide strategic advisory support to senior stakeholders and underwriters. Security Operations & Vulnerability Management Design and implement Security Operations Centre (SOC) operating models. Conduct vulnerability management programmes, gap analyses, and remediation tracking. Lead penetration testing and application security ...

Environmental Manager

Location
Greater London, England, United Kingdom
Environmental Manager London, United Kingdom | EMEA | Approximately 50% travel Role Summary As Environmental Manager, you will be the regional voice for environmental management across STACK Infrastructure EMEA, helping embed a consistent, high-quality approach throughout the full lifecycle of our data centre facilities. Based in London and working across … portfolio, supporting the Senior Director EHS - EMEA in meeting company protocols and applicable legislative requirements. Establish, maintain and promote an ISO 14001-compliant Environmental Management System (EMS), including effective maintenance of environmental legal registers. Monitor environmental performance across projects and operational facilities, developing inspection schedules and benchmarking to support ...

Senior Product Engineer (Product Manager)

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
move faster and scale with confidence. What you'll be doingOwn platform stabilityOwn platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce CI/… answer "why are we spending six engineers on this for three months?" in terms leadership understands. Track platform impact in business terms: developer velocity, incident cost, time-to-market, reliability, capacity - e.g. "CXO rollout removes a 3-second API fan-out, reducing session abandonment" or "we can now handle ...

Technical Product Manager (Superapp)

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
move faster and scale with confidence. What you'll be doingOwn platform stabilityOwn platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce CI/… answer "why are we spending six engineers on this for three months?" in terms leadership understands. Track platform impact in business terms: developer velocity, incident cost, time-to-market, reliability, capacity - e.g. "CXO rollout removes a 3-second API fan-out, reducing session abandonment" or "we can now handle ...

Senior Product Engineer (Product Manager)

Location
Greater London, England, United Kingdom
faster and scale with confidence. What you'll be doing Own platform stability Own platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce … answer "why are we spending six engineers on this for three months?" in terms leadership understands. Track platform impact in business terms: developer velocity, incident cost, time-to-market, reliability, capacity - e.g. "CXO rollout removes a 3-second API fan-out, reducing session abandonment" or "we can now handle ...

Principal Engineer (AWS & Java)

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
business‐critical changes, setting the standard for quality, performance, and maintainability. Provide hands‐on expertise in Java‐based systems, including concurrency, performance tuning, memory management, and API design. Non‐Functional Ownership & System PerformanceTake a leading role in defining, implementing, and validating non‐functional requirements, including system latency, throughput, scalability … reliability. Drive solutions for cloud scaling and capacity management, ensuring the platform performs predictably under variable and peak loads. Provide deep expertise in working with large databases and data stores, including performance optimisation, data access patterns, and operational scaling. Problem Solving & Technical Decision SupportSolve the most complex and ambiguous ...

Head of Derivatives eTrading and SDP

Location
Greater London, England, United Kingdom
adoption is aligned with the Bank’s risk appetite, regulatory obligations and technology governance standards. As a member of the Market Pricing, Connectivity & Analytics Management Team, the role provides senior engineering and platform leadership across electronic trading platforms (spot and derivatives), leading high‐performing global teams to deliver scalable … adoption is aligned with the Bank’s risk appetite, regulatory obligations and technology governance standards. As a member of the Market Pricing, Connectivity & Analytics Management Team, the role provides senior engineering and platform leadership across electronic trading platforms (spot and derivatives), leading high‐performing global teams to deliver scalable ...

Payments Systems Manager

Location
Greater London, England, United Kingdom
Maintain operational procedures, technical documentation, and support models. Ensure compliance with industry standards, including PCI DSS and P2PE requirements. Team Leadership Project and Change Management Lead the delivery of payment-related projects and strategic initiatives. Coordinate internal stakeholders, suppliers, and business teams to deliver successful outcomes. Manage project plans … architectures. Evaluate and recommend technology solutions that align with business strategy. Ensure solutions meet operational, security, compliance, and scalability requirements. Mobile and Loyalty Platform Management Own and manage two customer-facing mobile applications from a payment and technology perspective. Manage technical requirements and integrations for two loyalty platforms. Ensure ...

Customer Experience Engineering Manager

Hiring Organisation
Appcast
Location
London, UK
complex issues but also invests in engineering practices such as daily scrums and triage to deeply understand platform gaps from customer insights and incident signals. • Collaborate with Azure engineering teams using a prioritized set of opportunities to eliminate top issues impacting customer experience and improve Azure quality and security … continuously improve diagnostics and supportability. • Lead operational excellence by reinforcing ACE accountability for complex cases, improving Time to Mitigate (TTM), and maturing the ACE Incident-Management function to ensure high‐quality, engineering‐driven problem resolution. People Leadership: • Attract and build a diverse, high‐performing team with the capabilities ...

Customer Experience Engineering Manager

Location
City of Westminster, England, United Kingdom
complex issues but also invests in engineering practices such as daily scrums and triage to deeply understand platform gaps from customer insights and incident signals. Collaborate with Azure engineering teams using a prioritized set of opportunities to eliminate top issues impacting customer experience and improve Azure quality and security … continuously improve diagnostics and supportability. Lead operational excellence by reinforcing ACE accountability for complex cases, improving Time to Mitigate (TTM), and maturing the ACE Incident-Management function to ensure high-quality, engineering-driven problem resolution. Attract and build a diverse, high-performing team with the capabilities needed ...

Infrastructure Security Engineer

Hiring Organisation
Blockchain
Location
London, UK
Employment Type
Full-time
modeling, design reviews, and architectural assessments for new and existing systems. Contribute to internal security documentation, best practices, and developer guidance. Participate in security incident response when engineering expertise or automation support is needed. AI-Driven Innovation: Research, architect, and safely embed cutting-edge AI utilities and Large Language … continuously improve the security posture of complex systems. Familiarity with some of the following: Cloudflare (DDoS protection, WAF), OSS SIEM tools (Splunk, Elastic, etc), Incident management platforms (e.g. Incident.io, PagerDuty)Familiarity with at least one of the following CI/CD systems (Github Actions, Concourse, CircleCI)Familiarity with ...

Software Engineer III- JPM Personal Investing- Mid Level

Location
Greater London, England, United Kingdom
investing, and the trust that 150 years of J.P. Morgan heritage brings. J.P. Morgan Personal Investing offers award-winning investments, products and digital wealth management services to over 275,000 investors in the UK. We built the business withinnovation as a core part of our ethos to give consumers … Kotlin Write unit, component, integration, end-to-end & performance tests Support the products you've built through their entire life cycle, including production and incident management Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables ...

AI LLM Engineer - Autonomous Network

Location
Greater London, England, United Kingdom
vector search, semantic retrieval, graph-enhanced retrieval, and hybrid search patterns.* Develop AI agents for fault diagnosis, root-cause analysis, KPI analysis, configuration recommendation, incident summarisation, and operational decision support.* Build token-efficient prompting, context optimisation, caching, and response generation techniques.* Integrate LLM solutions with OSS, AIOps, inventory, graph … with telecom network data including RAN, Core, IP/MPLS, SD-WAN, OSS, alarms, KPIs, and inventory.* Experience developing LLM agents for network operations, incident management, or service assurance.* Experience with AI model optimisation, inference cost reduction, latency optimisation, and scalable AI serving.If you're excited about this ...

Technical Supervisor

Location
Northolt, England, United Kingdom
CRITICAL ENVIRONMENT Shift Pattern: Monday - Friday, Days Employment Type: Permanent We are currently recruiting for an experienced Technical Supervisor to join a leading facilities management team operating within a highly critical technical environment in Northolt. This is an excellent opportunity for an experienced engineering professional with a background … LOTO procedures. Monitor maintenance schedules and ensure PPM completion within required timescales. Conduct regular plant inspections and engineering audits. Monitor BMS and associated building management systems. Maintain accurate maintenance records, compliance documentation and operational reports. Support incident management and emergency response procedures. Liaise with the client ...

Field Tech Analyst -Windows support experience

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
Windows, troubleshooting, diagnosing, imaging/deployment and software installation. Serves as company liaison with customer on administrative and technical matters. Provide technical support and incident management service desk functions (Service Now) Reviews, troubleshoots, and approves operational quality desktops, notebooks, and associated peripherals (Windows ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
Pulumi, Kubernetes, Docker Strong knowledge of DevOps principles, CI/CD pipelines, and engineering best practices Experience driving operational excellence including reliability, monitoring, and incident management Ability to influence and communicate effectively with senior stakeholders and non-technical audiences Commercial awareness with experience managing budgets and optimising platform ...