701 to 725 of 1,254 Incident Management Jobs in the UK

IT Innovation Operations Specialist (AI) - IT Services - 107877 - Grade 7

Hiring Organisation
University of Birmingham
Location
United Kingdom
Salary
£ 70 K
innovation initiatives as part of the university’s strategic framework and digital strategy. The role will focus on the deployment, administration, monitoring, and operational management of innovative technical solutions and platforms that help to advance the University, improve its efficiency and enhance user experience.A primary focus for this role … term will be the comprehensive administration and operational rollout of the NebulaOne AI platform, ensuring its stability, user access management, and performance. Additionally, the post holder will manage and maintain operational views of Service Innovation and Automaton (SIAP) delivered platforms and applications, establishing clear monitoring, dashboards, and health checks ...

Business Applications Integration Analyst TLNT1 NI

Hiring Organisation
Dale Farm Ltd
Location
Belfast, UK
including ERP, WMS, HR, Payroll, Manufacturing, Logistics, and Reporting solutions. Success in the role requires strong application support, integration, SQL, problem-solving, and stakeholder management skills, with a focus on delivering practical technology solutions that improve business performance. Key Responsibilities Provide 1st, 2nd and 3rd line support for business … critical applications, including incident management, troubleshooting and root cause analysis. Develop and enhance business applications to meet operational and business requirements. Design, implement and support integrations between systems using APIs, file transfers, EDI and web services. Develop automation solutions using Power Automate, Power Apps and AI technologies ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
shape and drive how the firm builds and operates reliable, observable, secure, and cost-efficient systems on AWS. Working closely with development, platform, and incident management teams, you will define reliability in measurable terms and build the tooling and processes to achieve it, improving platform speed, stability … chaos engineering experiments to strengthen resilience and recovery. Automate operational processes to reduce manual intervention and toil across the stack. Support major incident response, root-cause analysis, and continual improvement actions. Collaborate cross-functionally to raise standards for stability, security, performance, and compliance. Required skills & experience 3+ years' experience ...

Principal Systems Architect - M365 and EUC

Hiring Organisation
Finastra
Location
London, UK
Employment Type
Full-time
Role OverviewThe Microsoft 365 Architect at Finastra is a senior, hands-on engineering and operations role responsible for the design and day-to-day management of the global Microsoft 365 environment. This role combines strategic platform architecture with practical administration and L3/L4 support, ensuring services are secure … Establish tenant-wide policies for data protection, retention, and regulatory compliance. Platform Administration & Monitoring: Administer and troubleshoot M365 services to ensure performance and availability. Incident Management & Escalation: Act as SME and escalation point for complex issues, performing root cause analysis. Migrations & Projects: Lead tenant-to-tenant migrations ...

Azure Databricks Engineer

Hiring Organisation
Capgemini
Location
Manchester, United Kingdom
Employment Type
Full Time
Databricks Platform Engineering Hands-on implementation and troubleshooting of Azure Databricks Databricks Serverless configuration and workload optimisation Databricks SQL and Delta Lake development Cluster management, compute policies, autoscaling, and workload monitoring Data Engineering Strong Python, PySpark, and SQL development skills Design and build ingestion and transformation pipelines Full … reconciliation, and error handling Integration Development REST API integration experience Integration with databases, SaaS applications, cloud storage, and enterprise systems Secure authentication and credential management Azure Platform Services Azure Identity and Access Management Azure Networking and Private Endpoints Azure Security and Key Vault Azure Monitoring and Operational Support ...

Software Engineering Specialist

Location
Belfast City District, Northern Ireland, United Kingdom
knowledge** and associated billing platform integrations.* **Event-Driven Architecture & Domain-Driven Design (DDD)** experience.* **CI/CD Automation & DevOps tooling** expertise.* **Production Support & Incident Management** including root cause analysis and resilience planning.* **Cross-Functional Stakeholder Management** across Product, Architecture, Security, QA, and Operations teams.* **AI technologies within ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
architect, scale and secure our cloud infrastructure, enabling world‐class software and data engineering teams to deliver at pace with confidence. Responsibilities Strict IaC Management : Define, version, and manage multi‐cloud GPU and CPU infrastructure strictly through modular, reusable Terraform. GitOps & GitFlow : Enforce rigorous GitFlow branching strategies, code reviews … Azure DevOps Pipelines : Design, build, and maintain secure, high‐speed CI/CD pipelines within Azure DevOps to automate code promotion and MLOps workflows. Incident Management : Lead the response, containment, and resolution of critical platform outages, serving as a primary escalation point for high‐priority cloud and infrastructure ...

Principal Software Engineer - HR

Hiring Organisation
Marks & Spencer
Location
United Kingdom
Salary
£ 60 K
solutions.Aligned to the People & Labour sub-domains, this role leads the modernisation of the Azure microservices and integration estate supporting colleague, HR and workforce management capabilities. The ecosystem includes SaaS platforms such as Oracle HCM and Blue Yonder Workforce Management, along with Azure services that integrate with external … approaches.• Mentor engineering managers, staff engineers and key engineering talent, fostering technical excellence, continuous improvement and high hiring standards.• Partner with HR, Payroll, Workforce Management, Product, Architecture, Security and Data stakeholders to align technology and business objectives and communicate technical trade-offs effectively.• Drive operational excellence through improved observability ...

AI Platform & Site Reliability Engineering Managing Consultant

Location
United Kingdom
This will include: AI Platform Strategy & Architecture: Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite … operational requirements. AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. Reliability Engineering & SRE: Establish SRE practices ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
capability.This will include:• AI Platform Strategy & Architecture: Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices ...

Expert Service Delivery Manager

Location
Greater London, England, United Kingdom
tower estate. This is a senior, client-facing role for someone who is equally comfortable in an ITIL governance forum, a Sev-1 mainframe incident bridge, and a boardroom QBR. Key responsibilities Acts as a client advocate and a point of escalations for client service delivery needs, including … well as exceeding expectations. Establishes and leads operational meetings focused on ITSM governance and SLA adherence. Required Qualifications 8+ years of IT Service Management experience in a client-facing role, with meaningful exposure to mainframe (z/OS) managed service environments. Operational ability in diverse, large-scale, multi-platform ...

Service Desk Operator

Hiring Organisation
Akkodis
Location
Stevenage, England, United Kingdom
scalable, data-driven recruitment ecosystem. Through redesigning, building, and rolling out a sophisticated Big Data system, our diverse roles span across architecture, project management, data analytics, development, and technical support, giving you the chance to shape a dynamic, next-gen digital infrastructure. As a 1st Line IT Service Desk … responsible for providing essential IT support, ensuring effective communication with users, and maintaining an organised and efficient service desk operation. Your key duties include incident management, problem-solving, and escalation where necessary, to maintain a high level of service and user satisfaction. Act as the first point ...

IT & Infrastructure Architec

Location
United Kingdom
Zabbix, Grafana). Work closely with developers to support microservices environments using Docker and Kubernetes. Own IT operations policies, processes, and ITIL‐aligned service management practices. Professional English proficiency (written and spoken). Qualifications 5+ years of practical experience in IT infrastructure and operations, with strong Azure expertise. Bachelor … Kubernetes in production. Solid database administration experience: SQL Server, PostgreSQL, Azure SQL, Cosmos DB. Experience managing monitoring tools (Zabbix, Grafana). Proven troubleshooting and incident management skills focused on uptime and reliability. Practical experience with CI/CD pipelines (Azure DevOps, GitHub Actions, Jenkins). Strong understanding ...

Cyber Security Operations Specialist

Hiring Organisation
Tank Recruitment
Location
Bath, Somerset, United Kingdom
Employment Type
Contract
Contract Rate
£450 - £550/day
will play a key role in detecting, investigating and responding to cyber threats, while helping to improve the organisation's overall security monitoring and incident response capabilities. The role combines hands-on security operations, tooling optimisation and continuous improvement across a varied technology estate. Key responsibilities: Monitor, triage … investigate security alerts and incidents across IT, cloud and OT environments. Lead or support incident response, including containment, eradication, recovery and escalation. Develop and optimise SIEM detection rules, EDR policies and security monitoring capabilities. Reduce false positives and improve detection coverage using threat intelligence and incident learnings. Investigate ...

Senior Cloud Engineer, Cloud COE

Hiring Organisation
Janus Henderson
Location
London, United Kingdom
Salary
£ 100 K
Ensure infrastructure deployments are consistent, version-controlled, and policy-compliantManage drift detection, environment consistency, and release governanceAzure Core Platform SkillsHands-on engineering across: Subscriptions, Management Groups, RBAC, PolicyVirtual Machines, Storage, Backup, DRAzure Monitor, Log AnalyticsPrivate Endpoints and secure service exposureImplement secure-by-design configurations aligned to enterprise cloud controlsAzure … Microsoft Entra ID (Azure AD) configurations and integrationsImplement RBAC, SSO, and least privilege access modelsAssist with security controls, compliance policies, and remediation activitiesMonitoring, BAU & Incident ManagementSupport business-as-usual (BAU) cloud operations across environmentsMonitor cloud services using Azure Monitor, Log Analytics, and alerting toolsInvestigate incidents, perform root cause analysis ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
will collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one.While this is not a people‐management or shift‐based role, you will work closely with global teams (UK and US) and may occasionally be called upon for major incidents …/SLOs.Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.Partner with Security teams to ensure our platforms meet compliance, security, and risk‐management expectations.Lead seamless handovers ...

Full Stack Engineer, Platform Reliability

Location
East Midlands, England, United Kingdom
removes recurring failure Be a first point of contact for the people who use our systems, triaging what comes in and working to defined incident severities and response times Investigate problems through to root cause and drive fixes to closure, including the design changes that stop them happening again … culture of continuous improvement, knowledge sharing, and technical excellence Required Skills & Qualifications Strong proficiency in Python for backend development Experience with React (hooks, state management, component architecture) Experience with AWS cloud services and cloud-native architecture Solid understanding of SQL and relational and non-relational databases Experience building ...

Platform Support Lead

Location
Greater London, England, United Kingdom
week) Employment: Fulltime Role Summary: The lead will act as the primary onshore contact for IDMC and TIBCO EBX support, coordinating platform operations, incident management, deployments, service transition, offshore delivery, and client communication. Key Responsibilities: Lead daily monitoring of IDMC and EBX platform health, Secure Agents, metadata scans … data-quality workflows, integrations, and scheduled jobs. Own high-priority incident coordination, perform technical triage, drive restoration, manage escalations, and support root-cause analysis. Coordinate Dev, UAT, and Production activities, including Secure Agent configuration, SSO/access issues, deployments, rollback planning, and post-deployment validation. Provide functional and technical ...

Salesforce.com Commercial Application Manager

Location
Thurso, Scotland, United Kingdom
prioritization, and delivery of enhancements. Stakeholder Engagement: Serve as liaisonbetween business and IT; lead discovery sessions,requirements gathering and solution design. Project/Enhancement Management and Execution Lead implementation activities through all phases of the project lifecycle. Support project planning, scope definition, timelines, resource coordination, and delivery management. Collaborate … operational excellence. Identifyproject risks, dependencies, and mitigation strategies to ensure successful project outcomes. Testing & QA: Ensure quality through testing, UAT, and release. Change Management: Drive user adoption through training, communication, and support. Leadership& Communication Serve as the primary interface between business stakeholders and IT, acting as the trusted advisor ...

Senior Platform Owner - Customer Engagement

Location
United Kingdom
Society. As our Senior Platform Owner , you will lead the evolution of the Dynamics 365 and Power Platform ecosystem, spanning CRM, marketing automation, case management, colleague engagement tools, workflow orchestration, and low‐code applications. This platform is a central enabler of Skipton’s purpose and transformation goals, powering journeys … Platform is the engine room of how Skipton understands, supports and communicates with its members. Powering Dynamics 365, Power Platform solutions, marketing automation, case management and workflow tooling, CEP is the central nervous system that ensures colleagues have the insight, context and capability to deliver human, meaningful interactions ...

Senior Engineering Manager- Payment Experience

Location
United Kingdom
evolve the full-stack systems behind both squads, guiding Cards decisions across PSP integrations, card token vaulting with 3DS, 3DS strategy, traffic routing management, fraud integration boundaries, and decline-recovery flows. Guide Cashier decisions across the Player Cashier frontend platform, the “golden development path”, web performance and TTI optimisation … reviewed, and reversible where possible. Own production excellence across the domain, including observability with Prometheus, Grafana, and Sentry; end-to-end payment-journey traceability; incident management and post-incident reviews; and reporting pipelines to Snowflake/DWH for approval-rate, cost, and conversion reporting. Champion safe delivery ...

SRE Engineer with .Net C# - Glasgow, UK

Hiring Organisation
Capgemini
Location
Glasgow City, United Kingdom
Employment Type
Full Time
Skills: Site Reliability Operational Support Monitor maintain and support businesscritical applications and cloud infrastructure Ensure high system availability performance scalability and reliability Participate in incident management root cause analysis and problem resolution Implement proactive monitoring alerting and observability solutions Reduce operational overhead through automation and selfhealing mechanisms Support … scripts and tools to improve operational efficiency Partner with development teams to streamline software delivery processes Drive continuous improvement initiatives across DevOps and release management practices Software Development Engineering Contribute to application enhancements and operational tooling using C and NET technologies Develop utilities automation scripts and support tools Assist ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve performance, availability, and reliability issues Lead and contribute to incident response, troubleshooting, and root cause analysis Automate operational processes and eliminate repetitive manual tasks Work closely with software engineers to improve deployment processes, system … understanding of metrics, logging, tracing, alerting, and system health Experience troubleshooting complex production environments Understanding of SLIs, SLOs, SLAs, and error budgets Experience with incident management and root cause analysis Good understanding of cloud networking, security, and infrastructure fundamentals Strong scripting/automation skills A strong understanding ...

Production Engineer – Trading & Electronic Trading Systems

Location
Greater London, England, United Kingdom
live trading and electronic trading platforms. These environments are production‐critical, operate in real time and require strong ownership of system stability, performance and incident management. The role involves close interaction with traders, developers, IT support and infrastructure teams. Role Overview We are looking for a Front Office Production … Responsibilities Production Ownership & Reliability Operate and support Front Office trading systems in production Ensure high availability, stability and performance during market hours Own incident management from detection to resolution and root cause analysis Participate in on‐call or production support rotations when required Contribute to post‐incident ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
distributed and event-driven systems using messaging and asynchronous processing patterns. Drive engineering excellence through architecture, coding standards, testing, and automation. Support production systems, incident management, troubleshooting, and operational resilience. Balance long-term architecture goals with short-term delivery priorities. Collaborate across engineering teams, product owners, architects … production support. Non-Technical Skills Strong leadership, ownership, and problem-solving mindset. Ability to thrive in ambiguous and rapidly changing environments. Excellent stakeholder management, communication, and collaboration skills. Proven ability to influence engineering practices beyond immediate teams. Experience coaching and developing engineers while driving continuous improvement. If you like ...