1,226 to 1,250 of 2,364 Incident Response Jobs in the UK

DevOps Engineer, Veeqo

Hiring Organisation
Veeqo
Location
Swansea, West Glamorgan, United Kingdom
Salary
£ 55 K
environments (EKS, Aurora, ElastiCache/Valkey, SQS, Amazon MQ, S3)- Drive security and compliance delivery: ASR certification, Shepherd risk remediation, environment separation, CVE response and patching- Execute infrastructure cost optimisation against S-team goals (DRAM reduction, right-sizing, VPA adoption)- Build and maintain CI/CD pipelines (Jenkins/… VQJenkins migration, deployment automation, Terraform)- Participate in on-call rotation (primary and secondary) for production incident response, triage and resolution- Deliver infrastructure changes via MCMs (Managed Change Management) with full rollback planning and operational rigour- Contribute to platform reliability: multi-AZ resilience, monitoring (Prometheus/CloudWatch), alerting ...

Infrastructure Software Engineer, Apps Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 80 K
debugging skills and the ability to navigate performance/security tradeoffs in production systemsComfort with ambiguity, and the ability to context-switch between reactive incident work and proactive product developmentNice to haves:Experience as a founder or early engineer at an infrastructure-focused startup, owning a product … multi-tenant or untrusted environments (e.g., FaaS, CI sandboxes, remote notebooks)Open-source contributions to systems or developer-tools projectsHistory of on-call/incident response for production systemsPLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows ...

Senior System Specialist (SuccessFactors)

Location
Farnborough, England, United Kingdom
compliance, including role-based access controls, change management, audit readiness, data protection, and adherence to organisational policies and best practices Lead system support, incident resolution, and service delivery, acting as an escalation point for complex issues, coordinating with stakeholders, and ensuring system availability, performance, and service excellence Oversee integrations … align technology solutions with business objectives across complex, multi-site organisations Strong governance, security, and service management knowledge, including access controls, compliance, audit requirements, incident, change, and problem management, and adherence to regulatory standards such as GDPR and ISO 27001 Excellent stakeholder, vendor, and communication skills, able to influence ...

Interim Cyber Security Officer

Location
Greater London, England, United Kingdom
CrowdStrike and Splunk platforms. The successful candidate will ensure the effective integration, configuration, and operational use of security tools to enhance threat detection, incident response, and overall security maturity. Additionally, the role involves providing technical leadership, mentoring, and knowledge transfer to bolster internal cyber capabilities during a period … high‐severity security incidents, facilitating rapid investigation, containment, and remediation using EDR and SIEM tools. Develop and implement SOAR workflows to automate detection, response, and security operations processes. Conduct proactive threat hunting using SIEM/EDR data and MITRE ATT&CK‐aligned techniques. Support vulnerability assessment and security scanning ...

security engineer in cloud security

Location
Greater London, England, United Kingdom
environments Build and maintain protective and detective controls for cloud threat visibility Develop detection engineering capabilities to identify attacker behaviour and reduce detection and response times Lead cloud security investigations and incident response with engineering teams Participate in a shared on-call rotation and coordinate containment, remediation … security operations capabilities at scale Experience with cloud-native security platforms, CNAPP technologies, or large-scale cloud security monitoring Experience applying AI to detection, response, threat analysis, or security operations workflows Familiarity with AWS networking, identity, and cloud security patterns Experience working across large-scale SaaS environments and high ...

ML Ops Engineer

Hiring Organisation
CMC Markets
Location
London, United Kingdom
Salary
£ 80 K
engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response.What you’ll work onML lifecycle and platform engineeringBuild repeatable workflows for model training, validation, promotion, deployment and retraining.Productionise models through packaging, versioning, model registry … efficiency through automation, observability and infrastructure as code.Write production-grade Python for long-running services, deployment tooling and ML workflows.Establish testing, validation, release and incident-management practices for ML systems.Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance.Make explicit trade ...

Cyber Security Engineer (Identity & Cloud)

Location
Warrington, England, United Kingdom
secure configuration and operation of Microsoft Azure and Microsoft 365. Implement and maintain Microsoft Defender security controls. Support security monitoring, threat investigation and incident response activities. Support vulnerability management, including identifying, prioritising and remediating security weaknesses. Support Power Platform governance, including Data Loss Prevention (DLP) policies, environment management … Experience implementing or supporting Conditional Access and access governance. Knowledge of RBAC and privileged access management. Experience investigating security alerts and supporting cyber security incident response. Good understanding of Azure security fundamentals and Microsoft security best practices. Experience supporting Microsoft 365 security and cloud-based services. Experience producing technical ...

Blockchain Cyber Intelligence Vice President

Hiring Organisation
JP Morgan Chase
Location
London, United Kingdom
Salary
> £ 150 K
vulnerabilities and compromises, utilizing advanced tools and techniques to detect anomalies and contribute to the development of strategies for security investigation, threat mitigation, and incident responseProduce actionable threat intelligence products including IOC development, exploit analysis reports, and threat actor tracking to inform detection engineering and security operationsCollaborate with cross … Computer Science, Cybersecurity, Data Science, or related disciplines5+ years of experience in cybersecurity threat intelligence or operations, with a focus on threat detection, incident response, and security infrastructure managementDemonstrated expertise in multiple security domains, including network security, malware analysis, threat hunting, threat intelligence production, and security architecture ...

Site Reliability Engineer

Hiring Organisation
Bristow Holland Ltd
Location
Manchester, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum - Offering 100% Work from home
scalability, performance and security Responding to production incidents, diagnosing complex technical issues and restoring services Carrying out root cause analysis and contributing to post-incident reviews Building automation to reduce manual and repetitive operational tasks Using Infrastructure as Code to improve the consistency and scalability of environments Working closely … solid background within Site Reliability Engineering, DevOps, Platform Engineering or a similar environment Terraform and Infrastructure as Code experience Strong production troubleshooting and incident response experience Experience supporting highly available production environments A strong understanding of reliability, scalability and automation Experience with scripting or programming, ideally PowerShell ...

Site Reliability Engineer

Hiring Organisation
Bristow Holland Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum - Offering 100% Work from home
scalability, performance and security Responding to production incidents, diagnosing complex technical issues and restoring services Carrying out root cause analysis and contributing to post-incident reviews Building automation to reduce manual and repetitive operational tasks Using Infrastructure as Code to improve the consistency and scalability of environments Working closely … solid background within Site Reliability Engineering, DevOps, Platform Engineering or a similar environment Terraform and Infrastructure as Code experience Strong production troubleshooting and incident response experience Experience supporting highly available production environments A strong understanding of reliability, scalability and automation Experience with scripting or programming, ideally PowerShell ...

Site Reliability Engineer

Hiring Organisation
Bristow Holland Ltd
Location
Edinburgh, City of Edinburgh, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum - Offering 100% Work from home
scalability, performance and security Responding to production incidents, diagnosing complex technical issues and restoring services Carrying out root cause analysis and contributing to post-incident reviews Building automation to reduce manual and repetitive operational tasks Using Infrastructure as Code to improve the consistency and scalability of environments Working closely … solid background within Site Reliability Engineering, DevOps, Platform Engineering or a similar environment Terraform and Infrastructure as Code experience Strong production troubleshooting and incident response experience Experience supporting highly available production environments A strong understanding of reliability, scalability and automation Experience with scripting or programming, ideally PowerShell ...

Site Relaibility Engineer

Hiring Organisation
Bristow Holland
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£55,000 - £60,000 per annum
scalability, performance and security Responding to production incidents, diagnosing complex technical issues and restoring services Carrying out root cause analysis and contributing to post-incident reviews Building automation to reduce manual and repetitive operational tasks Using Infrastructure as Code to improve the consistency and scalability of environments Working closely … solid background within Site Reliability Engineering, DevOps, Platform Engineering or a similar environment Terraform and Infrastructure as Code experience Strong production troubleshooting and incident response experience Experience supporting highly available production environments A strong understanding of reliability, scalability and automation Experience with scripting or programming, ideally PowerShell ...

ML Ops Engineer

Location
City of Westminster, England, United Kingdom
engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response. Build repeatable workflows for model training, validation, promotion, deployment and retraining. Productionise models through packaging, versioning, model registry integration, deployment automation and safe rollback. … automation, observability and infrastructure as code. Write production‐grade Python for long-running services, deployment tooling and ML workflows. Establish testing, validation, release and incident-management practices for ML systems. Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance. Make ...

Workday Global HCM Support Lead

Hiring Organisation
WPP
Location
London, United Kingdom
Salary
£ 80 K
role will be responsible for maintaining relationships with key stakeholders and managing a team of Workday specialists.What you'll be doing:Lead or coordinate incident managementrelated to system issues, ensuring incidents are assessed, escalated, communicated, and resolved in a timely and controlled mannerProvide clear communication to stakeholders during incidents … including business impact, status updates, actions being taken, and next stepsSupport post-incident review, root cause analysis, and follow-up actions to improve release quality, system reliability, and operational processesLead the planning and coordination of monthly local changes and bi-annual global releases, ensuring activity is prioritised, governed ...

Systems Administrator

Location
Greater London, England, United Kingdom
triage, prioritise, resolve and keep tickets updated. This is a core part of the role, not an occasional overflow duty. Work to agreed response and resolution SLAs and flag any ticket at risk of breach in advance. Handle access requests directly: password and MFA resets, account lockouts, licence … audits and surveillance visits, and drive remediation of findings to closure. Lead on client and prospect security questionnaires and due diligence requests. Participate in incident response, including out of hours where the severity requires it. Essential 5 or more years in systems administration, infrastructure or IT engineering, including ...

Lead Software Engineer - Cloud Platform Engineering

Location
Greater London, England, United Kingdom
overall operational stability of software applications and systems Provide operational support for production systems in a “you-build-it-you-run-it” culture, including incident response and continuous reliability improvements. Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs … practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across ...

Oracle Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, United Kingdom
Salary
£ 60 K
DescriptionPurpose of the roleTo apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them. AccountabilitiesAvailability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.Resolution, analysis and response … Drive – the operating manual for how we behave. Join Barclays as an Oracle Site Reliability Engineer, where you'll leverage automation, engineering excellence, and incident management best practices to ensure our critical systems remain resilient, available, and ready to support millions of customers every day.To be successful ...

Network Event Monitoring Analyst

Hiring Organisation
Pontoon
Location
Chester, Cheshire, United Kingdom
Employment Type
Contract
cutting-edge technologies and contribute to the stability and performance of their network infrastructure. Key Responsibilities: Monitoring and Management: Oversee Monitoring Event Management and Incident Management, ensuring timely Level 1 triage of networking issues. Network Operations: Provide operational support for network environments, including switching, routing, perimeter security, traffic management … advanced tools such as Splunk, SevOne, IBM Watson AI Ops, Wireshark, NetScout, and Gigamon to identify service impacts and interpret monitors, dashboards, and logs. Incident Response: React effectively to production failures, managing triage efforts, engaging appropriate teams, and communicating with management and technical escalation. Documentation and Reporting: Maintain ...

Software Engineer

Hiring Organisation
Topcoder
Location
United Kingdom
establishing best practices and ensuring code quality across mission-critical data platform initiatives. You will drive continuous improvement in data governance, infrastructure optimization, and incident response while managing the full lifecycle of data warehouse operations. Responsibilities Lead technical design and architecture of Snowflake and AWS data warehouse platforms … services including data lake storage, compute, and related data platform services Proficiency in data modeling, SQL, and ETL/ELT processes Proven experience with incident management and troubleshooting complex data platform issues Knowledge of data governance, security, and compliance best practices Client-facing experience with ability to communicate technical ...

Director, Cloud Infrastructure

Location
Greater London, England, United Kingdom
This person will own the platform foundations Sanity engineers build on every day: cloud infrastructure, Kubernetes, networking, routing, observability, CI/CD, deployment paths, incident response, and the standards that make production ownership work across product teams. The scale is real. Content Lake alone handles around … line between Platform and SRE. Platform should own the shared foundations. SRE should help product teams run services well, with strong tooling, standards, and incident support. Raise the reliability bar across Sanity's production systems, including dashboards, alert severity, paging standards, service ownership, on‐call readiness, and incident ...

Senior Platform Engineer (Infrastructure)

Hiring Organisation
Fresha
Location
London, United Kingdom
Salary
£ 80 K
Argo Workflows for legacy pipelinesRedis for key/value storageTerraform for infrastructure as codePython and GoLang for our Platform toolsDatadog, Sentry for observability and incident responseWhat you will be doing:Platform Architecture & ScaleDefine and evolve infrastructure architecture to support multi-region deploymentsDesign systems for resilience, scalability, and operational simplicityExplore … scalabilityBring hands-on experience with AI/ML or LLM-based tools in a platform or operational context (e.g. automation, observability, developer experience, or incident response)Demonstrate strong programming and automation skills in Python, Go, or Bash, and use Infrastructure as Code (Terraform) to build repeatable, low-risk ...

Oracle Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, UK
Employment Type
Full-time
DescriptionPurpose of the roleTo apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them. AccountabilitiesAvailability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning. Resolution, analysis … response to system outages and disruptions, and implement measures to prevent similar incidents from recurring. Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience. Monitoring and optimisation of system performance and resource usage, identify and address bottlenecks, and implement best ...

AI Factory Lead Architect

Hiring Organisation
Advania UK
Location
London, United Kingdom
Salary
£ 80 K
Studio, Power Platform, Azure DevOps, and our Azure subscriptions used for hostingDefine and run the support model for the AI estate – monitoring and alerting, incident response, change and release management, documentation and knowledge transfer, so solutions remain reliable as usage growsEnable the wider maker community – publish agent templates … business can move quickly, safely and within known guardrailsA supportable estate: solutions are monitored, maintained and improved after launch, with clear ownership and response expectationsSenior stakeholders across the business who see the AI Factory as the natural route to solving their problems with AIQualification & ExperienceDemonstrable experience designing and delivering ...

Senior Firewall / Security Network Engineer

Location
Greater London, England, United Kingdom
hands‐on senior role: you'll set the firewall architecture — especially for hybrid and cloud connectivity — while running the day‐to‐day change and incident work that keeps the business connected and secure. You'll be the technical anchor as we migrate a multi‐vendor estate onto a standardized … Plan and execute quarterly firewall software upgrades and hardware refreshes with minimal disruption. Manage IPsec/site-to-site migrations and firewall decommissioning. Operations, incident response & availability Triage and resolve connectivity incidents — blocked/dropped traffic, TLS handshake failures, VPN outages, firewall failovers. Own the out-of-band ...

Technical Support Manager

Location
Perth, Scotland, United Kingdom
Technical Support Manager to lead the operational support, maintenance and resilience of high-availability operational technology (OT) systems, with a focus on incident response, service delivery, patching and vulnerability remediation. This is a hands‐on operational leadership role covering the day‐to‐day support and operational management … core OT platforms, including DMS, SCADA, data historians, telemetry and communications systems. The successful candidate will own incident management, fault diagnosis and recovery within a high‐availability environment, lead OT patching and vulnerability remediation activity, and manage resourcing, capability and development across a technical team. This role suits ...