1 to 25 of 57 Permanent Incident Response Jobs in South Wales

Security Engineer - Systems Integrator

Location
Cardiff, Wales, United Kingdom
growing Security Operations function, with a strong focus on Microsoft Defender XDR and Microsoft Sentinel. The role covers security engineering, detection development, vulnerability management, incident response, and Microsoft 365 security, with direct client exposure. This is a primarily remote position with occasional attendance at the London or Cardiff … associated risks. Produce structured security reports and recommendations. Present technical findings to clients and stakeholders. Support client remediation and security improvement activities. Handle escalated Incident Response investigations. Support advanced security investigations from the triage team. Develop and improve Incident Response playbooks. Automate processes using Azure Logic ...

Remote Staff Software Engineer - Databases SRE | UK | Remote

Hiring Organisation
Grafana Labs
Location
Monmouthshire, United Kingdom
models Proactively reduce SLO burn to prevent repeat incidents Serving as a primary escalation point and on-call for relevant incidents Lead customer-impacting incident response and post-incident reviews Contribute to design docs and code reviews Influence feature design to ensure production scalability and operability Build … some knowledge of networking, cloud storage, and scaling. Excellent problem-solving and troubleshooting skills. Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents) Ability to reason about performance, scaling ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Cowbridge, Glamorgan, United Kingdom
models Proactively reduce SLO burn to prevent repeat incidents Serving as a primary escalation point and on-call for relevant incidents Lead customer-impacting incident response and post-incident reviews Contribute to design docs and code reviews Influence feature design to ensure production scalability and operability Build … some knowledge of networking, cloud storage, and scaling. Excellent problem-solving and troubleshooting skills. Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents) Ability to reason about performance, scaling ...

Remote Head of DevOps

Location
Cardiff, Glamorgan, United Kingdom
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design … problem-solving. Skills: Platform Engineering: Building developer-centric tools to increase engineering velocity. Cloud Governance: Managing multi-provider cloud footprints and resource optimisation. Incident Response: Leading high-pressure incident resolution and system recovery. Mentorship: Coaching senior engineers and developing future technical leaders. Process Design: Crafting standardised, scalable ...

Security Architect

Location
Cardiff, Wales, United Kingdom
Defining, implementing, and maintaining corporate security policies, standards, and procedures to ensure compliance with industry regulations, legal requirements (e.g., GDPR, HIPAA), and best practices. Incident Response and Management: Playing a key role in developing incident response plans and coordinating efforts to detect, analyse, and respond ...

Remote DevSecOps Engineer

Hiring Organisation
Mbr Partners Inc
Location
Blaenau gwent, United Kingdom
make active use of AI coding assistants to accelerate infrastructure and automation work. While ownership of the security posture, disaster recovery, business continuity, and incident response sits with the Head of Technology Delivery, you will play a central role in implementing and maintaining the controls, backups, and monitoring … that underpin these, and in supporting incident resolution when issues arise. The reliability, performance, and efficiency of the platform depend directly on the quality of the work you do, making this a role of significant technical responsibility. The business is a well-funded business with excellent revenues ...

Remote Senior Director, Engineering- X-Ops Platform

Location
Caldicot, Monmouthshire, United Kingdom
leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform reliability, resilience, and operational excellence: incident management, on-call standards, SLIs/SLOs, post-incident learning, and continuous improvement Partner with product engineering and threat research leaders to improve … focused environment Strong track record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution ...

Cyber Security Architect

Location
Newport, Wales, United Kingdom
proposed technology changes and provide recommendations to stakeholders. Promote security awareness and influence teams to adopt security best practices. Assist with cyber security incident response by identifying architectural weaknesses and recommending improvements. Continuously monitor emerging cyber threats, vulnerabilities, technologies, and regulatory changes to improve the organisation's security ...

Remote Senior Platform Engineer

Location
Pontypool, Monmouthshire, United Kingdom
troubleshooting techniques, analyzing logs and metrics, and collaborating with development teams to identify and resolve root causes. You will also contribute to creating comprehensive incident response plans and post-mortem analyses. Drive Automation Initiatives to Streamline Operational Tasks and Enhance System Reliability: You will champion automation initiatives ...

Remote Senior Platform Engineer

Location
Pontyclun, Glamorgan, United Kingdom
troubleshooting techniques, analyzing logs and metrics, and collaborating with development teams to identify and resolve root causes. You will also contribute to creating comprehensive incident response plans and post-mortem analyses. Drive Automation Initiatives to Streamline Operational Tasks and Enhance System Reliability: You will champion automation initiatives ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Bridgend, Glamorgan, United Kingdom
Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling. Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents) Ability to reason about performance ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Ebbw Vale, Monmouthshire, United Kingdom
Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling. Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents) Ability to reason about performance ...

Remote Senior Software Engineer

Location
Bridgend, Glamorgan, United Kingdom
thousand happens somewhere every day. Mentor a Software Engineer, and make them better rather than doing their work for them. Take part in incident response, on-call, and post-incident review for the systems your team owns, where the clock is measured against shops that cannot serve ...

Remote Senior Software Engineer

Location
Ebbw Vale, Monmouthshire, United Kingdom
thousand happens somewhere every day. Mentor a Software Engineer, and make them better rather than doing their work for them. Take part in incident response, on-call, and post-incident review for the systems your team owns, where the clock is measured against shops that cannot serve ...

Remote DevOps Engineer

Location
Ebbw Vale, Monmouthshire, United Kingdom
tooling for application and infrastructure metrics, e.g. Prometheus, Datadog, Open Telemetry Candidates with the following background will be of particular interest: Experience contributing to incident response across a complex microservice-based application Application Security best practice including identifying potential threats and vulnerabilities in applications and cloud infrastructure, designing ...

Remote Engineer

Location
Glamorgan, United Kingdom
contact for diagnosing and resolving platform-related issues, including performance bottlenecks, scalability challenges, and security vulnerabilities. You will also contribute to creating comprehensive incident response plans and post-mortem analyses. Drive Automation Initiatives to Streamline Operational Tasks and Enhance System Reliability: You will champion automation initiatives to eliminate ...

Senior Engineer H/F

Location
Pontyclun, Glamorgan, United Kingdom
contact for diagnosing and resolving platform-related issues, including performance bottlenecks, scalability challenges, and security vulnerabilities. You will also contribute to creating comprehensive incident response plans and post-mortem analyses. Drive Automation Initiatives to Streamline Operational Tasks and Enhance System Reliability: You will champion automation initiatives to eliminate ...

Engineer / Senior Engineer M/F

Location
Swansea, Glamorgan, United Kingdom
contact for diagnosing and resolving platform-related issues, including performance bottlenecks, scalability challenges, and security vulnerabilities. You will also contribute to creating comprehensive incident response plans and post-mortem analyses. Drive Automation Initiatives to Streamline Operational Tasks and Enhance System Reliability: You will champion automation initiatives to eliminate ...

Senior Engineer H/F

Location
Pontypool, Monmouthshire, United Kingdom
contact for diagnosing and resolving platform-related issues, including performance bottlenecks, scalability challenges, and security vulnerabilities. You will also contribute to creating comprehensive incident response plans and post-mortem analyses. Drive Automation Initiatives to Streamline Operational Tasks and Enhance System Reliability: You will champion automation initiatives to eliminate ...

Site Reliability Engineer

Location
Cardiff, Wales, United Kingdom
hosted workloads Demonstrable experience having worked on AWS builds (standing up new infrastructure, IaC, migrations) as well as ongoing operational/support work (troubleshooting, incident response, day-to-day maintenance) Solid working knowledge of observability: metrics, logs, traces, and how to design or tune alerting and dashboards Comfortable ...

Remote Senior Platform Engineer, Fanatics Markets - UK

Location
Bridgend, Glamorgan, United Kingdom
high operational reliability. Partner with product and engineering teams to translate technical requirements into scalable platform solutions. Mentor junior and mid-level engineers, lead incident response for team-owned services, and advocate for modern engineering best practices. Experience and Skills 5 plus years of experience in platform engineering ...

Senior DevOps Engineer - Remote UK

Location
Cardiff, Wales, United Kingdom
facilitate seamless product feature releases, adhering to established TechOps peer‐review processes. Monitor application-level health within the cluster, collaborating with TechOps on incident response and driving post‐mortems for engineering‐led deployments. Maintain clear documentation, deployment runbooks, and developer self‐service guides to streamline the engineering onboarding ...

Senior DevOps Engineer

Location
Cardiff, Wales, United Kingdom
facilitate seamless product feature releases, adhering to established TechOps peer-review processes. Monitor application-level health within the cluster, collaborating with TechOps on incident response and driving post-mortems for engineering-led deployments. Maintain clear documentation, deployment runbooks, and developer self-service guides to streamline the engineering onboarding ...

Senior Infrastructure Engineer (UK Gov)

Hiring Organisation
NonStop Consulting Ltd
Location
Cardiff, South Glamorgan, United Kingdom
Employment Type
Full-Time
Salary
£550.00 - £600.00 per day
maintain secure infrastructure and platforms Drive automation, resilience and continuous improvement Support infrastructure hardening, vulnerability management and monitoring Work across IAM, configuration management and incident response Provide technical expertise on secure and compliant environments Essential Skills Linux AWS Terraform Docker Bash Python Strong security and infrastructure principles ...

Infrastructure Engineer

Hiring Organisation
Circle Recruitment
Location
Cardiff, South Glamorgan, United Kingdom
Employment Type
Full-Time
Salary
£550.00 - £600.00 per day
with teams across the organisation to provide expert advice on secure environments, configuration management, identity and access management, infrastructure hardening, vulnerability management, monitoring, and incident response You will help ensure that critical services remain secure, available, and resilient against evolving cyber threats Linux/Terraform/Docker/ ...