301 to 325 of 484 Incident Management Jobs in London

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. * Support live service operations, incident management, troubleshooting and platform resilience. * Collaborate with engineers, architects, product teams and stakeholders whilst coaching and mentoring colleagues. Essential Skills for the Senior ...

Senior Engineering Manager — Hotels.com Technology

Location
Greater London, England, United Kingdom
quality and reliability in the shopping funnel: set engineering quality and operational excellence standards, and ensure robust CI/CD, monitoring, alerting, and incident response practices are consistently applied. Move fast with incomplete information — make well‐reasoned decisions without waiting for certainty or broad consensus, and validate assumptions through … domain contexts. Minimum Qualifications: Bachelor's Degree or Equivalent Level; Technical Degree Preferred. 8+ years of relevant professional experience and 3+ years of people management experience. Experience managing engineering teams with end‐to‐end ownership of services or domains, including design, implementation, deployment, and ongoing operations for production services ...

AVP, Automation & AI Delivery

Hiring Organisation
Arch Capital Group
Location
London, UK
Employment Type
Full-time
success metrics and determine feasibilityOwn, Track and report on delivery scope, progress, risks and dependenciesEnsure delivery standards for requirements, documentation, testing, releases, monitoring and incident management are metImplement feedback loops and metrics to track and report on adoption, trust, and impact to SA leadership and senior stakeholdersCollaborate with … standards and approvals (permitted use, privacy, security, model governance and third-party risk).Role Requirements & SkillsSkills/CompetenciesHighly organised with strong programme/project management skills; able to manage multiple initiatives concurrently and deliver results. Exceptional stakeholder management with the ability to build trusting relationships and drive consensus ...

Data Operations Specialist

Location
Greater London, England, United Kingdom
initiatives that improve data quality, resilience, governance, and scalability, while acting as a trusted partner to business leaders and technology teams.**Responsibilities*** Ongoing Data Management: Own and optimize Data Operations pipelines, with deep expertise in cleansing, and maintaining corporate structures and company- data. Establish best practices and standards … ongoing data management.* Incident Management & Resolution: Serve as the escalation point for complex or high-impact data issues. Lead root cause analysis, post-mortems, and the design of long-term preventative solutions to reduce operational risk.* Knowledge Leadership: Act as a subject-matter expert in company-level data ...

Cyber Risk & Security Manager

Location
Greater London, England, United Kingdom
governance advisory services, security maturity programmes, and SOC strategy engagements. The role requires a highly capable security professional with experience across risk advisory, vulnerability management, security operations design, compliance frameworks, and executive reporting. The successful candidate will operate at both strategic and technical levels, supporting board-level decision‐making … resilience simulations. Produce executive‐level cyber risk reports and mitigation roadmaps. Provide strategic advisory support to senior stakeholders and underwriters. Security Operations & Vulnerability Management Design and implement Security Operations Centre (SOC) operating models. Conduct vulnerability management programmes, gap analyses, and remediation tracking. Lead penetration testing and application security ...

Environmental Manager

Location
Greater London, England, United Kingdom
Environmental Manager London, United Kingdom | EMEA | Approximately 50% travel Role Summary As Environmental Manager, you will be the regional voice for environmental management across STACK Infrastructure EMEA, helping embed a consistent, high-quality approach throughout the full lifecycle of our data centre facilities. Based in London and working across … portfolio, supporting the Senior Director EHS - EMEA in meeting company protocols and applicable legislative requirements. Establish, maintain and promote an ISO 14001-compliant Environmental Management System (EMS), including effective maintenance of environmental legal registers. Monitor environmental performance across projects and operational facilities, developing inspection schedules and benchmarking to support ...

Senior Product Engineer (Product Manager)

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
move faster and scale with confidence. What you'll be doingOwn platform stabilityOwn platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce CI/… answer "why are we spending six engineers on this for three months?" in terms leadership understands. Track platform impact in business terms: developer velocity, incident cost, time-to-market, reliability, capacity - e.g. "CXO rollout removes a 3-second API fan-out, reducing session abandonment" or "we can now handle ...

Technical Product Manager (Superapp)

Hiring Organisation
Lendable
Location
London, UK
Employment Type
Full-time
move faster and scale with confidence. What you'll be doingOwn platform stabilityOwn platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce CI/… answer "why are we spending six engineers on this for three months?" in terms leadership understands. Track platform impact in business terms: developer velocity, incident cost, time-to-market, reliability, capacity - e.g. "CXO rollout removes a 3-second API fan-out, reducing session abandonment" or "we can now handle ...

Senior Product Engineer (Product Manager)

Location
Greater London, England, United Kingdom
faster and scale with confidence. What you'll be doing Own platform stability Own platform stability as a primary accountability - SLA/SLO metrics, incident rates, release quality. Hold release gating authority: block features from production that don't meet testing standards or introduce unacceptable regression risk. Enforce … answer "why are we spending six engineers on this for three months?" in terms leadership understands. Track platform impact in business terms: developer velocity, incident cost, time-to-market, reliability, capacity - e.g. "CXO rollout removes a 3-second API fan-out, reducing session abandonment" or "we can now handle ...

Principal Engineer (AWS & Java)

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
business‐critical changes, setting the standard for quality, performance, and maintainability. Provide hands‐on expertise in Java‐based systems, including concurrency, performance tuning, memory management, and API design. Non‐Functional Ownership & System PerformanceTake a leading role in defining, implementing, and validating non‐functional requirements, including system latency, throughput, scalability … reliability. Drive solutions for cloud scaling and capacity management, ensuring the platform performs predictably under variable and peak loads. Provide deep expertise in working with large databases and data stores, including performance optimisation, data access patterns, and operational scaling. Problem Solving & Technical Decision SupportSolve the most complex and ambiguous ...

Head of Derivatives eTrading and SDP

Location
Greater London, England, United Kingdom
adoption is aligned with the Bank’s risk appetite, regulatory obligations and technology governance standards. As a member of the Market Pricing, Connectivity & Analytics Management Team, the role provides senior engineering and platform leadership across electronic trading platforms (spot and derivatives), leading high‐performing global teams to deliver scalable … adoption is aligned with the Bank’s risk appetite, regulatory obligations and technology governance standards. As a member of the Market Pricing, Connectivity & Analytics Management Team, the role provides senior engineering and platform leadership across electronic trading platforms (spot and derivatives), leading high‐performing global teams to deliver scalable ...

Payments Systems Manager

Location
Greater London, England, United Kingdom
Maintain operational procedures, technical documentation, and support models. Ensure compliance with industry standards, including PCI DSS and P2PE requirements. Team Leadership Project and Change Management Lead the delivery of payment-related projects and strategic initiatives. Coordinate internal stakeholders, suppliers, and business teams to deliver successful outcomes. Manage project plans … architectures. Evaluate and recommend technology solutions that align with business strategy. Ensure solutions meet operational, security, compliance, and scalability requirements. Mobile and Loyalty Platform Management Own and manage two customer-facing mobile applications from a payment and technology perspective. Manage technical requirements and integrations for two loyalty platforms. Ensure ...

Customer Experience Engineering Manager

Location
City of Westminster, England, United Kingdom
complex issues but also invests in engineering practices such as daily scrums and triage to deeply understand platform gaps from customer insights and incident signals. Collaborate with Azure engineering teams using a prioritized set of opportunities to eliminate top issues impacting customer experience and improve Azure quality and security … continuously improve diagnostics and supportability. Lead operational excellence by reinforcing ACE accountability for complex cases, improving Time to Mitigate (TTM), and maturing the ACE Incident-Management function to ensure high-quality, engineering-driven problem resolution. Attract and build a diverse, high-performing team with the capabilities needed ...

Infrastructure Security Engineer

Hiring Organisation
Blockchain
Location
London, UK
Employment Type
Full-time
modeling, design reviews, and architectural assessments for new and existing systems. Contribute to internal security documentation, best practices, and developer guidance. Participate in security incident response when engineering expertise or automation support is needed. AI-Driven Innovation: Research, architect, and safely embed cutting-edge AI utilities and Large Language … continuously improve the security posture of complex systems. Familiarity with some of the following: Cloudflare (DDoS protection, WAF), OSS SIEM tools (Splunk, Elastic, etc), Incident management platforms (e.g. Incident.io, PagerDuty)Familiarity with at least one of the following CI/CD systems (Github Actions, Concourse, CircleCI)Familiarity with ...

Software Engineer III- JPM Personal Investing- Mid Level

Location
Greater London, England, United Kingdom
investing, and the trust that 150 years of J.P. Morgan heritage brings. J.P. Morgan Personal Investing offers award-winning investments, products and digital wealth management services to over 275,000 investors in the UK. We built the business withinnovation as a core part of our ethos to give consumers … Kotlin Write unit, component, integration, end-to-end & performance tests Support the products you've built through their entire life cycle, including production and incident management Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables ...

AI LLM Engineer - Autonomous Network

Location
Greater London, England, United Kingdom
vector search, semantic retrieval, graph-enhanced retrieval, and hybrid search patterns.* Develop AI agents for fault diagnosis, root-cause analysis, KPI analysis, configuration recommendation, incident summarisation, and operational decision support.* Build token-efficient prompting, context optimisation, caching, and response generation techniques.* Integrate LLM solutions with OSS, AIOps, inventory, graph … with telecom network data including RAN, Core, IP/MPLS, SD-WAN, OSS, alarms, KPIs, and inventory.* Experience developing LLM agents for network operations, incident management, or service assurance.* Experience with AI model optimisation, inference cost reduction, latency optimisation, and scalable AI serving.If you're excited about this ...

Technical Supervisor

Location
Northolt, England, United Kingdom
CRITICAL ENVIRONMENT Shift Pattern: Monday - Friday, Days Employment Type: Permanent We are currently recruiting for an experienced Technical Supervisor to join a leading facilities management team operating within a highly critical technical environment in Northolt. This is an excellent opportunity for an experienced engineering professional with a background … LOTO procedures. Monitor maintenance schedules and ensure PPM completion within required timescales. Conduct regular plant inspections and engineering audits. Monitor BMS and associated building management systems. Maintain accurate maintenance records, compliance documentation and operational reports. Support incident management and emergency response procedures. Liaise with the client ...

Field Tech Analyst -Windows support experience

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
Windows, troubleshooting, diagnosing, imaging/deployment and software installation. Serves as company liaison with customer on administrative and technical matters. Provide technical support and incident management service desk functions (Service Now) Reviews, troubleshoots, and approves operational quality desktops, notebooks, and associated peripherals (Windows ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
Pulumi, Kubernetes, Docker Strong knowledge of DevOps principles, CI/CD pipelines, and engineering best practices Experience driving operational excellence including reliability, monitoring, and incident management Ability to influence and communicate effectively with senior stakeholders and non-technical audiences Commercial awareness with experience managing budgets and optimising platform ...

Senior Backend Engineer - Payments

Location
Greater London, England, United Kingdom
users Take ownership of product development, from feature discovery, to the breakdown of work, and its implementation End-to-end application support, including production incident management Embrace agile methodologies and user-centred thinking Engage in a culture of continuous improvement by attending events such as blameless post-mortems ...

Site Reliability Engineer

Hiring Organisation
GoCardless
Location
London, UK
Employment Type
Full-time
Building new platform components, from initial design to deployment. Operational Support: Ensuring the health and stability of our systems through on-call rotations and incident management. Business as Usual (BAU): Continuous maintenance, improvements, and support for audits and compliance. Tech stack and tools: Python, Ruby, Golang;Terraform;Atlantis ...

Staff Engineer (Architect) - Researcher Operations

Location
Greater London, England, United Kingdom
pairing using tools like Git and GitHub. Experience of operationally managing software components once live, including observability, logging, metrics, error reporting, debugging and live incident management. Experience of working with sensitive personal data. Competencies Experience working in/with cross‐functional teams consisting of engineers, product ...

PROJECT LEAD L1(CONTRACT)

Location
Greater London, England, United Kingdom
Lead and oversee complex IT projects focused on Data Center and Infrastructure Migrations aligned with business goals. Manage project governance including budget control, risk management, scope maintenance, and delivery timelines. Coordinate project deliverables and proactively identify and resolve risks and issues. Engage with diverse stakeholders, facilitating decision-making … securing approvals at critical milestones. Optimize resource allocation and enhance team productivity. Track progress using advanced project management tools and report outcomes to senior management. Ensure compliance with organizational policies and industry best practices. Requirements Experience: 5–10 years in IT project management, specifically in Infrastructure and Data ...

Desktop Support Engineer - Quant Trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Desktop Engineer will: Manage raised requests and incidents in a timely manner following well-rehearsed internal request runbooks/knowledge based and incident management and request fulfillment processes;Partner with Level 2 & 3 internal and external teams/vendors as appropriate to help troubleshoot more complex incidents;Provide … alert management to pro-active resolve failure alerts, degraded services and unavailable services through internal monitoring tools;You will need: Detailed knowledge and experience in a Microsoft environment supporting Windows 10, Windows Server 2012+, Office Suite including Outlook, Office 365 and Azure, including OneDrive, DLP and SharePointDetailed knowledge ...

Principal Data Engineer

Location
Greater London, England, United Kingdom
authority for our clients, ensuring their data platforms are resilient, scalable, and aligned with their long-term business goals. Key Responsibilities Customer Excellence & Stakeholder Management: Act as the primary strategic partner to client business owners, aligning technical data solutions with corporate goals while championing reliability and \"clean data … Execution: Oversee \"team-of-teams\" delivery performance, ensuring the quality and resilience of high-impact data pipelines while managing capacity planning and team velocity. Incident Management & Resilience: Lead the triage and resolution of critical platform incidents, taking full ownership of root-cause prevention and maintaining transparent communication with ...