26 to 40 of 40 Incident Management Jobs in Central London

Sr Director, Platform Engineering Data Platform & Agentic Platform

Hiring Organisation
Hackajob Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
responsible for delivering and operating shared platform capabilities that enable trusted data services and production-grade AI-driven workflows. Working closely with Platform Product Management, you will help ensure platforms are scalable, reliable, secure, and widely adopted across product domains while supporting innovation through high-quality engineering execution. Responsibilities … mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC, vulnerability management, and change management compliance Requirements 12+ years leading platform engineering teams delivering shared services or large-scale SaaS systems with production operations accountability ...

Application Support Analyst | Enterprise Technology

Location
City Of London, England, United Kingdom
Enterprise platforms offered by Marex to both internal and external client base. Support business users offering second- and third-line support. Provide incident management per ITIL standards. Support applications and infrastructure hosted within AWS. Support REST APIs, API gateways, microservices and distributed services. Support application releases and deployments … through CI/CD pipelines, ensuring changes are managed and approved in accordance with Marex change management and release governance processes. Manage new system analysis and implementation. Liaison between technology departments to communicate system changes. Manage process and system documentation in existing template; produce and regularly maintain ...

Lead GCP Engineer

Location
City Of London, England, United Kingdom
platform direction across the product lifecycle. Remain hands‐on in the delivery of complex engineering work rather than operating solely in an architectural or management capacity. Lead solution design activities, code reviews and technical assurance across the platform. Cloud Architecture & Engineering Design, implement and operate secure, scalable, cloud‐native … platform components and services. Ensure solutions align with organisational governance standards, security policies and regulatory obligations. Balance innovation and rapid delivery with appropriate risk management and control frameworks. Design secure data integration patterns incorporating authentication, authorisation and access management technologies. Support delivery of solutions that are suitable ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve performance, availability, and reliability issues Lead and contribute to incident response, troubleshooting, and root cause analysis Automate operational processes and eliminate repetitive manual tasks Work closely with software engineers to improve deployment processes, system … understanding of metrics, logging, tracing, alerting, and system health Experience troubleshooting complex production environments Understanding of SLIs, SLOs, SLAs, and error budgets Experience with incident management and root cause analysis Good understanding of cloud networking, security, and infrastructure fundamentals Strong scripting/automation skills A strong understanding ...

Onshoring - Capacity Manager

Hiring Organisation
Hays Technology
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600/day Up to £600pd inside IR35 via umbrella
Identify potential capacity constraints and performance risks before they impact users or services. Coordinate capacity planning activities with technical and operational teams. Support major incident management by assessing capacity and resilience impacts. Contribute to service recovery planning where capacity and performance are affected. Produce demand forecasts and growth … related risks and issues to operational and leadership stakeholders. Maintain capacity dashboards, reports, and continuous improvement plans. Key Skills & Experience Proven experience in Capacity Management within a complex IT environment. Strong understanding of infrastructure, application performance, and service management principles. Experience producing capacity forecasts, utilisation reports, and risk ...

Cloud Platforms Engineer

Location
City Of London, England, United Kingdom
than a queue of requests. Keep production healthy. Help diagnose and resolve production issues when they arise, bring a clear view of what good incident management looks like, and fix the class of problem rather than just the instance. Make sure we can recover. Look after backup … Azure or Google Cloud Platform experience is fine. Practical experience with infrastructure as code in a team setting, Terraform preferred – modules, code review, state management and applying through a pipeline, rather than clicking in a console. Desirable: experience of cloud identity and access at organisation scale, and single sign ...

Customer Experience Engineering Manager

Location
City of Westminster, England, United Kingdom
complex issues but also invests in engineering practices such as daily scrums and triage to deeply understand platform gaps from customer insights and incident signals. Collaborate with Azure engineering teams using a prioritized set of opportunities to eliminate top issues impacting customer experience and improve Azure quality and security … continuously improve diagnostics and supportability. Lead operational excellence by reinforcing ACE accountability for complex cases, improving Time to Mitigate (TTM), and maturing the ACE Incident-Management function to ensure high-quality, engineering-driven problem resolution. Attract and build a diverse, high-performing team with the capabilities needed ...

Support Engineer - £50K - Remote

Location
City Of London, England, United Kingdom
operational health of the Digital team’s products and end user support and project work will cover security and IAM, incident management and resolution and even automation. Requirements: Good, broad IT support and service management experience Experience of operational support, DevOps and/or platform AWS/ ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software products. … Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using configuration management platforms like Ansible, Puppet, or Chef Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer ...

Software Engineer - (Typescript & Node)

Location
City Of London, England, United Kingdom
pairing using tools like Git and GitHub Experience of operationally managing software components once live, including; observability, logging, metrics, error reporting, debugging and live incident management Experience of working with sensitive personal data Competencies Experience working in/with cross-functional teams consisting of e.g. engineers, product, security ...

Technical Lead Infrastructure

Hiring Organisation
Saga Group
Location
Central London, London, United Kingdom
Employment Type
Permanent
Salary
£70,000
infrastructure expert when critical technology systems fail or degrade and will act as the primary technical advisor to the major incident management team. Provide assurance that the infrastructure remains fully supported, adopting best practice and working with the partner to ensure all systems are maintained. Work with … partner to ensure that they are fully competent in the management and maintenance of the infrastructure. Be the technical gatekeeper and critical reviewer for our partners and actively interrogate their engineering choices and designs using best practice and company standards. Be the bridge between infrastructure teams and corporate governance ...

Business Analyst / PM (Payments Systems)

Location
City Of London, England, United Kingdom
Ensure seamless interaction between payment, loyalty, and customer engagement solutions. Drive enhancements to improve customer experience and business value. Lead the implementation and ongoing management of point-to-point encryption (P2PE) across the estate. Support security audits, assessments, and remediation activities. Maintain awareness of emerging threats and regulatory developments. … success delivering technology and business change projects. Experience translating business requirements into technical solutions. Experience managing mobile applications and customer-facing digital platforms. Strong incident management and problem-solving capabilities. Experience operating within PCI DSS-regulated environments. Desirable Experience P2PE implementation experience. Knowledge of loyalty platforms and customer ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning and performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure … team, sharing learnings and best practices with the broader production engineering organizationContribute hands-on to technical work including code, system design reviews, and incident response, using AI tooling to expand personal and team reach across disciplinesPartner cross-functionally with software engineering, data science, and product teams to unblock dependencies ...

Lead Software Engineer - Backend Engineer - Chase UK

Location
City Of London, England, United Kingdom
date by continuously updating our technologies and patterns Support the products you've built through their entire lifecycle, including in production and during incident management Requirements Formal training or certification on software engineering concepts and applied experience Recent hands-on professional experience as a back-end software engineer ...

Senior IT Service Desk Engineer – AI-Driven Support

Location
City Of London, England, United Kingdom
Europe, and global needs. You will act as a technical escalation point, troubleshoot Windows/macOS, identity and access, endpoint lifecycle, and incident management, while driving automation and knowledge base improvements. You will partner with security, networks, CRM, facilities, and regional IT teams, and support on-call rotations ...