651 to 675 of 1,062 Incident Management Jobs in England

Software Engineer

Location
Greater London, England, United Kingdom
auto-remediate global SaaS infrastructure. What You'll Do Technical Design & Architecture: Design and implement high-resilience software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design, build, and maintain production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines … quality guardrails. Reliability & Scalability Engineering: Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews (PIRs). Mentorship & Best Practices: Mentor mid-level and junior engineers, conduct thorough code reviews, establish engineering guidelines, and drive operational excellence across ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
planning, operational readiness reviews, and launch reviews. Scale and evolve systems by driving changes that improve reliability, capacity, performance, and operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. … financial markets, technology, and continuous learning. ABOUT GOLDMAN SACHS The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
planning, operational readiness reviews, and launch reviews. Scale and evolve systems by driving changes that improve reliability, capacity, performance, and operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. … financial markets, technology, and continuous learning. ABOUT GOLDMAN SACHS The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded ...

Lead Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, UK
Employment Type
Full-time
collaborate with Architecture, Engineering, Security, and Platform teams to ensure reliability is built into systems from day one. While this is not a people‐management or shift‐based role, you will work closely with global teams and may occasionally be called upon for major incidents or critical issues. This … Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health. Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns. Partner with Security teams to ensure our platforms meet compliance, security, and risk‐management expectations. Lead ...

Application Support Analyst | Energy Trading Operations

Location
Greater London, England, United Kingdom
Support business users offering second‐ and third‐line support. Knowledge of scripting language (PowerShell, Python, etc.). Manage new system analysis and implementation. Provide incident management per ITIL standards. Liaison between technology department and business groups to communicate system changes. Manage process and trading system documentation in existing … responsibility. Report any breaches of policy to Compliance and/or your supervisor as required. Escalate risk events immediately. Provide input to risk management processes, as required. Competencies, Skills and Experience Competencies: A collaborative team player, approachable, self‐efficient and influences a positive work environment. Demonstrates curiosity. Resilient ...

Senior Platform Owner - Customer Engagement

Location
Skipton, England, United Kingdom
Society. As our Senior Platform Owner , you will lead the evolution of the Dynamics 365 and Power Platform ecosystem, spanning CRM, marketing automation, case management, colleague engagement tools, workflow orchestration, and low‐code applications. This platform is a central enabler of Skipton’s purpose and transformation goals, powering journeys … Platform is the engine room of how Skipton understands, supports and communicates with its members. Powering Dynamics 365, Power Platform solutions, marketing automation, case management and workflow tooling, CEP is the central nervous system that ensures colleagues have the insight, context and capability to deliver human, meaningful interactions ...

Platform Engineer

Location
Leeds, England, United Kingdom
pipelines to enable frequent, automated and reliable software delivery.Improve platform reliability through observability, monitoring, automated testing and disaster recovery practices.Support production systems, participate in incident response and contribute to continuous service improvements.Collaborate with engineers, architects and stakeholders to deliver scalable platform solutions that meet business needs.Contribute to platform standards … Terraform.Knowledge of CI/CD technologies, GitLab CI, GitHub Actions, ArgoCD and modern software delivery practices.Familiarity with observability, security and operational excellence, including monitoring, incident management and platform reliability.Experience working in agile environments with a passion for automation, continuous learning and emerging technologies, including AI-powered engineering tools.What ...

Network Engineer

Location
Greater London, England, United Kingdom
Python, Ansible, Jinja2 and Git Supporting and extending Infrastructure-as-Code and CI/CD practices for network provisioning Owning break/fix and incident response for complex enterprise and campus network issues Driving incident management practices, including on-call rotation, escalation procedures, postmortems and root-cause ...

Head of Platform Engineering

Location
Manchester, England, United Kingdom
improvement Identify capability gaps and support hiring, development, and succession planning Own resource planning, including capacity, skills mix, and on-call structure Operational Resilience & Incident Ownership Establish and run a two-tier resolver model: Ops on-call for standard support and runbook-driven issues, Product Engineering (Buyer/Seller … minimum 3 engineers per domain, or full team where smaller) so resolution doesn't depend on any one individual Drive a genuinely blameless Post-Incident Review culture, building on the existing Incident Management SOP Security & Risk Embed secure coding practice, SAST/DAST scanning, and OWASP-aligned ...

Cloud Platforms Engineer

Location
Greater London, England, United Kingdom
than a queue of requests. Keep production healthy. Help diagnose and resolve production issues when they arise, bring a clear view of what good incident management looks like, and fix the class of problem rather than just the instance. Make sure we can recover. Look after backup … Azure or Google Cloud Platform experience is fine. Practical experience with infrastructure as code in a team setting, Terraform preferred – modules, code review, state management and applying through a pipeline, rather than clicking in a console. Desirable: experience of cloud identity and access at organisation scale, and single sign ...

Lead Software Engineer- Python / Java- (Cloud Data Platform — AWS/Databricks/Terraform)

Location
Greater London, England, United Kingdom
practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across … platform engineering capability, including CI/CD design, automated security/quality gates, observability (logs/metrics/traces), SLOs/SLAs, and incident management/RCA in regulated environments. Proven end-to-end delivery leadershipfor large, cross-functional initiatives (roadmaps, dependency management, stakeholder alignment, budgeting/ ...

Associate Application Analyst

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
obligations during any inclement weather/disaster event. Ability to track, update and resolve all assigned incidents, changes and problem tickets via the Service management system, ensuring that documentation is thorough, accurate and meets a standard of high quality. Proactively monitor, recognize, and alert to a variety of system … more of Linux, Unix, MVS, VM, SuperVision, HMC, ServiceNow, CoPilot, ChatGPT, Kubernetes, Kafka, Opera, DB2, ClickHouse/Splunk or CassandraStrong troubleshooting, critical-thinking and incident management skillsExcellent verbal/written communication, organizational skills; ability to prioritize in a constantly changing workload. Strong interpersonal skills and ability to excel ...

Security Incident Governance Analyst

Location
Preston, England, United Kingdom
CBSbutler is seeking Security Incident Management Analysts to join a secure programme on a contract basis in the Preston/Inverness area. The role focuses on governance and coordination across multiple suppliers to track, report and evidence security incidents, rather than hands-on SOC activity. You will work … with security stakeholders to ensure effective incident governance, escalation, and reporting, with emphasis on evidence and audit readiness. Active SC clearance is required. #J-18808-Ljbffr ...

Senior Systems Engineer (Server and Cloud Technologies Specialist)

Hiring Organisation
Trusted Technology Partnership
Location
Ringwood, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£55,000
technologies. You will support in the technical delivery of complex projects and collaborate with stakeholders across the business, while also ensuring operational excellence through incident management, preventative maintenance, and ongoing service improvement. Skills and Experience Essential: Extensive experience in the design, deployment, and ongoing support of Microsoft server … Best Place to Work award and overall winner of the Ringwood Business Awards 2024. Our core services include support desk, on-site engineering, project management and delivery, storage and logistics, and technical consultancy. We encourage progression within Trusted Technology Partnership for our colleagues, offering opportunities in other teams ...

Lead Software Engineer- Python / Java- (Cloud Data Platform — AWS/Databricks/Terraform)

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across … platform engineering capability, including CI/CD design, automated security/quality gates, observability (logs/metrics/traces), SLOs/SLAs, and incident management/RCA in regulated environments. Proven end-to-end delivery leadership for large, cross-functional initiatives (roadmaps, dependency management, stakeholder alignment, budgeting ...

Python Lead Software Engineer - (Cloud Data Platform — AWS/Databricks/Terraform)

Location
Greater London, England, United Kingdom
practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across … platform engineering capability, including CI/CD design, automated security/quality gates, observability (logs/metrics/traces), SLOs/SLAs, and incident management/RCA in regulated environments. Proven end-to-end delivery leadership for large, cross-functional initiatives (roadmaps, dependency management, stakeholder alignment, budgeting ...

Technical Supervisor

Location
Northolt, England, United Kingdom
CRITICAL ENVIRONMENT Shift Pattern: Monday - Friday, Days Employment Type: Permanent We are currently recruiting for an experienced Technical Supervisor to join a leading facilities management team operating within a highly critical technical environment in Northolt. This is an excellent opportunity for an experienced engineering professional with a background … LOTO procedures. Monitor maintenance schedules and ensure PPM completion within required timescales. Conduct regular plant inspections and engineering audits. Monitor BMS and associated building management systems. Maintain accurate maintenance records, compliance documentation and operational reports. Support incident management and emergency response procedures. Liaise with the client ...

DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions
Location
Newcastle upon Tyne, Shiremoor, Tyne & Wear, United Kingdom
Employment Type
Permanent
Salary
£45000 - £58000/annum
solutions. Promote DevSecOps practices, security controls, and operational excellence. Collaborate with engineers, architects, testers, and product teams to deliver high-quality software. Lead troubleshooting, incident management, and continuous improvement initiatives. Skills & Experience Experience with AWS, Azure, or GCP. Strong knowledge of CI/CD tools such as GitLab ...

IT Support Technician

Hiring Organisation
Syntax Consultancy Ltd
Location
Derby, Derbyshire, United Kingdom
Employment Type
Permanent
Salary
£26000 - £28000/annum £26-28k (DOE)
+ Pension + Health Care + Training & Career Development: Providing 1st/2nd line IT support to end users including: trouble-shooting, fixes, IT incident management, call logging, prioritisation, managing tickets + escalating complex IT incidents. Providing IT support across a range of customer IT environments including: Desktop ...

Infrastructure Engineer

Location
Southampton, England, United Kingdom
Infrastructure Engineer be doing? Maintaining, securing, and enhancing multi-classification virtualised environments across public, private, and hybrid cloud systems. The role covers monitoring, incident management, controlled maintenance, and ensuring platforms remain secure and compliant, alongside continuous skills development, stakeholder support, and knowledge sharing to improve performance and team ...

DevOps Engineer x 5

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £552.00 per day
capabilities Troubleshoot and resolve infrastructure, platform and service issues Monitor platform health, availability and performance Contribute to infrastructure improvements and operational resilience Support incident management and service recovery activities Produce and maintain technical documentation Collaborate with engineering teams to improve deployment processes and service reliability Participate in support ...

Security Analyst

Location
Reading, England, United Kingdom
clients’ infrastructure, applications and digital environments secure and running smoothly. You’ll deliver high‐quality technical support, lead on day‐to‐day security incident management, and ensure we meet our SLAs across multiple managed service customers. Responsibilities Provide 2nd line support via our Service Desk, resolving incidents ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
Pulumi, Kubernetes, Docker Strong knowledge of DevOps principles, CI/CD pipelines, and engineering best practices Experience driving operational excellence including reliability, monitoring, and incident management Ability to influence and communicate effectively with senior stakeholders and non-technical audiences Commercial awareness with experience managing budgets and optimising platform ...

Senior Machine Learning Engineer (MLOps)

Location
Greater London, England, United Kingdom
frameworks that enable data scientists and ML engineers to deploy models safely and efficiently. Own production services, infrastructure and operational excellence practices, including incident management and root cause analysis. Optimise distributed compute workloads and resource utilisation across cloud environments. Drive Infrastructure-as-Code adoption and platform standardisation across ...

Digital Technology Engineer - MES

Location
Templecombe, England, United Kingdom
report test results, defects, and improvement opportunities. Ongoing Support: Provide technical assistance to super users and functional teams, addressing application issues or queries. Support incident management and troubleshooting for MES related problems during and after deployment. Assist in maintaining and updating MES documentation, user guides, and training materials. ...