1 to 25 of 463 Permanent High Availability Jobs in London

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
Automate infrastructure provisioning, deployment, configuration, database schema creation, and operational workflows across cloud and on-prem platforms.* Support production systems with a focus on high availability, disaster recovery, incident response, vulnerability management, patching, and change management.* Containerize applications and support deployments using container orchestration platforms.* Partner with development … management with Terraform, Ansible, Chef, or similar.* **Infrastructure & Operations** + Strong understanding of networking, routing, and traffic flow across distributed systems. + Knowledge of high availability, resiliency, capacity, and latency considerations. + Experience supporting customer-facing, low-latency, high-availability, or other business-critical platforms. + ...

devops engineer in low-latency trading infrastructure

Location
Greater London, England, United Kingdom
Automate infrastructure provisioning, deployment, configuration, database schema creation, and operational workflows across cloud and on-premises platforms; Support production systems with a focus on high availability, disaster recovery, incident response, vulnerability management, patching, and change management; Containerize applications and support deployments using container orchestration platforms; Partner with development … code and configuration management using Terraform, Ansible, Chef, or similar; Strong understanding of networking, routing, and traffic flow across distributed systems; Knowledge of high availability, resiliency, capacity, and latency considerations; Experience supporting customer-facing, low-latency, high-availability, or other business-critical platforms; Monitoring, logging ...

Technical Account Manager

Location
Greater London, England, United Kingdom
experience. End-to-End Assessment: Evaluate business application dependencies, integrating cloud and on-premises risk assessments; provide recommendations and risk alerts for critical pathways. High-Availability Drills: Conduct high-availability drills with customers under extreme scenarios, simulating failures, practicing failover procedures, and participating in post-drill … cloud migration journey. Well-Architected Support: Support comprehensive optimization aligned with Well-Architected Framework principles, covering security & compliance, operational stability, cost optimization, and high-performance architecture design. Service Assurance Service Management: Establish service and communication channels during cloud adoption; provide tailored support across online, on-site, and multi-project ...

Global Banking & Markets - Software Engineer - Vice President - London London · United Kingdom [...]

Location
Greater London, England, United Kingdom
production idioms, while meeting the diverse and often complex business requirements of a global 24x7 trading operation, while balancing stringent non‐functional demands for availability, latency, and resilience. Critically, you will lean heavily into AI‐driven development. Using Goldman Sachs' AI tooling and agentic coding assistants, you will govern … magnitude higher volumes at lower operational costs, directly contributing to the firm's competitive and commercial edge. Build for the future : Design and implement high‐availability, multi‐region, event‐driven services on a modern cloud‐native platform, setting the architectural standard for years to come. What You Will ...

Global Banking & Markets - Software Engineer - Vice President - London

Hiring Organisation
Goldman Sachs
Location
London, UK
Employment Type
Full-time
production idioms, while meeting the diverse and often complex business requirements of a global 247 trading operation, while balancing stringent non-functional demands for availability, latency, and resilience. Critically, you will lean heavily into AI-driven development. Using Goldman Sachs' AI tooling and agentic coding assistants, you will govern … magnitude higher volumes at lower operational costs, directly contributing to the firm's competitive and commercial edge. Build for the future: Design and implement high-availability, multi-region, event-driven services on a modern cloud-native platform, setting the architectural standard for years to come. What You Will ...

DB2 Database Administrator (DBA)

Location
Greater London, England, United Kingdom
Database Administrator (DBA) , the primary responsibility is to design, administer, and provide support for database technologies, ensuring customer satisfaction through timely, high-quality delivery aligned to defined service levels, project schedules, and organizational standards. The role requires maintaining and improving SLAs while ensuring compliance with organizational policies, standards … efficiency and service excellence. As part of a global support organization, the individual will leverage industry best practices to support diverse customer environments, maintain high service quality, and contribute to continuous improvement initiatives. The role demands strong technical expertise in IBM z/DB2 infrastructure, collaboration with stakeholders ...

CICS Mainframe Systems Programmer

Location
Greater London, England, United Kingdom
growing team in the UK and rest of Europe. About the role The CICS Mainframe Systems Programmer is responsible for ensuring the stability, availability, performance, and continuous improvement of CICS services within the IBM Mainframe environment. The role requires strong technical expertise, domain knowledge, and specialist skills to operate … solution design, deployment planning, configuration, testing, and ongoing operational support. The successful candidate will oversee the management and support of all CICS services, ensure high availability of business-critical applications, provide technical leadership during incident resolution, and collaborate with development and application support teams to deliver robust solutions. ...

Network and Voice Engineer

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
Network and Voice Engineer to design, support and enhance resilient enterprise networks and unified communications within a regulated financial services environment. You will ensure high availability, performance, security and service quality across LAN/WAN, data centre connectivity, cloud networking and voice platforms, working closely with infrastructure, security … service management teams. Key responsibilities Design, implement and maintain routing and switching, including VLANs, STP, QoS, OSPF/BGP and high-availability patterns. Support WAN and remote connectivity (MPLS/SD-WAN, VPN, IPSec/SSL), optimising performance and resilience. Administer network security controls such as firewalls ...

Infrastructure Engineer

Location
London, United Kingdom
commissioning and live operations, for next-generation Building Management Systems (BMS), Central Monitoring Platforms (CMS), and CCTV environments. Key Responsibilities: Architect, build, and deploy high-availability BMS, CMS, and CCTV environments. Drive an automation-first approach using Infrastructure as Code tools (Terraform, Ansible, PowerShell). Solve complex platform … across compute, storage, and networking layers. Integrate multi-vendor systems into cohesive, scalable, and secure CNI platforms. Enhance platform security, uptime, and resilience within high-availability environments. Develop reusable runbooks, technical documentation, and engineering standards. What You Need to Succeed: Active UK DV Clearance (Must have resided ...

z/OS Mainframe Systems Programmer

Location
Greater London, England, United Kingdom
Plan and execute operating system upgrades and software maintenance activities using SMP/E. Administer and support large-scale Parallel Sysplex environments to ensure high availability and performance. Perform system performance monitoring, capacity planning, and resource optimization. Analyse and resolve complex system-level issues involving z/… security, and operations teams to ensure platform stability. Implement software maintenance, fixes, service packs, and vendor-recommended updates. Support disaster recovery, business continuity, and high-availability initiatives. Maintain system security controls and ensure compliance with organizational standards. Create and maintain technical documentation, procedures, and operational runbooks. Participate ...

Database Administrator (short term contract)

Location
Greater London, England, United Kingdom
team on an initial three‐month contract to take ownership of the database infrastructure that underpins SilverRail’s rail distribution platforms. These systems process high volumes of search, booking and ticketing transactions for rail operators, retailers and travellers across the UK, Europe, North America and Australia/APAC … schema design, query and database optimisation, and cloud‐hosted database infrastructure from day one. We are looking for someone who has supported systems with high transaction volumes and strict availability requirements, such as travel, e‐commerce or financial services. Key Responsibilities Database Administration & Design Design, implement and maintain ...

Principal Engineer I — Prepurchase Platform

Location
City Of London, England, United Kingdom
Prepurchase Platform group owns the backend services that power every fan's path to a ticket. We handle event discovery, real-time seat availability, and interactive venue experiences at massive scale, processing billions of API calls during peakon-sales. Our engineers build systems that must be fast, resilient … reaching a fan is a problem, not a data point; treat it accordingly and act with urgency regardless of volume. Build end-to-end, high-availability systems that handle extreme traffic spikes during high-demand on-sales without degradation. Lead platform modernization initiatives - migrating legacy services ...

DevOps Engineer, Studios

Location
Greater London, England, United Kingdom
pipelines across our digital and broadcast platforms. The successful candidate will play a key role in enabling continuous delivery, improving system reliability, and supporting high-profile clients and live event services, ensuring optimal performance and resilience across all environments. Key Responsibilities and Accountabilities Design, build, and maintain scalable cloud … efficient, reliable software delivery across multiple teams. Automate infrastructure provisioning using Infrastructure as Code tools such as Terraform, CloudFormation, or similar. Monitor system performance, availability, and reliability using observability tools such as Prometheus, Grafana, and ELK stack. Ensure high availability and disaster recovery strategies are in place ...

Senior Fullstack Engineer (Python + React.js)

Location
Greater London, England, United Kingdom
maintain scalable, secure, and efficient server-side applications. The ideal candidate should have experience in microservices architecture, API development, database management, and frontend ensuring high availability and performance of backend and frontend services. In this role, you will collaborate closely with frontend engineers, product managers, and other stakeholders … Redux or React Query Collaborate with frontend developers to ensure efficient API integration and a seamless user experience. Troubleshoot and resolve production issues, ensuring high availability and minimal downtime. Write unit and integration tests to maintain code reliability and ensure high- quality releases. Continuously monitor and optimize ...

Database Administrator

Location
Greater London, England, United Kingdom
solutions under the guidance of the DBA Team Participate in team meetings and proactively contribute to database strategy discussions Other requirements raised by management High Availability & Scalability Support high-availability and disaster recovery (HA/DR) solutions, including replication and failover strategies Optimize databases for scalability ...

Site Reliability Engineer, Studios

Location
Uxbridge, England, United Kingdom
project stakeholders. Key Responsibilities And Accountabilities Design, build, and maintain reliable, scalable infrastructure and platform services across on‐premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering‐led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards … Drive root cause analysis and corrective actions following incidents, with a focus on prevention and continuous improvement. Support the design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements. Help enforce security, access control, patching, and operational best practices across infrastructure and services. Optimise system ...

Site Reliability Engineer, Studios

Location
Greater London, England, United Kingdom
project stakeholders. Key Responsibilities and Accountabilities Design, build, and maintain reliable, scalable infrastructure and platform services across on-premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering-led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards … Drive root cause analysis and corrective actions following incidents, with a focus on prevention and continuous improvement. Support the design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements. Help enforce security, access control, patching, and operational best practices across infrastructure and services. Optimise system ...

platform engineer in hybrid cloud

Location
Greater London, England, United Kingdom
deployment models Integrate applications with DevOps toolchains and CI/CD pipelines Administer hybrid cloud environments using Storage, Compute, and Security Services Configure High Availability and Disaster Recovery solutions Manage Kubernetes clusters Contribute to innovative AWS infrastructure solutions that drive business success. Требования: Hands‐on experience designing … cloud‐native architecture patterns Experience integrating CI/CD pipelines and DevOps toolchains such as GitLab CI, Jenkins, or GitHub Actions Experience implementing High Availability and Disaster Recovery strategies in AWS environments Strong Linux systems knowledge and troubleshooting capability in cloud and container environments Ability to work collaboratively ...

Engineer - Site Reliability

Location
Greater London, England, United Kingdom
support model for its US Global Trading Hours (GTH) markets, providing critical overnight and early‐session coverage from London that ensures continuous, high‐availability operations across Cboe's real‐time low‐latency trading platforms. The London‐based SRE provides technical support to Cboe Trade Desk and Operations Support … timely, precise communication to stakeholders during active incidents and contribute to post‐incident reviews and remediation tracking to drive long‐term platform stability. System Availability & Technical Support: Provide technical support and operational oversight to sustain resiliency and high availability of critical business operations. Monitor production, disaster recovery ...

DevSecOps Engineer — Data-Sovereign, High-Availability Platform

Location
Greater London, England, United Kingdom
tangible impact on our world. Whether you join our London HQ or the wider global organisation, you’ll be a part of collaborative, high-performing teams, creating cutting-edge software, platforms, and infrastructure. The Role Join us as a DevSecOps and help us build the future of data sovereignty … seeking an DevSecOps passionate about creating high-performance, secure, scalable, and reliable services for our production infrastructure and security. You'll have a direct impact, improving existing systems and developing innovative solutions to complex challenges. Our small, collaborative engineering teams own the full lifecycle of their services, from development ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
team responsible for network architecture, deployment, and operational readiness. You will play a key role in ensuring connectivity solutions align with performance, security, and availability requirements. The ideal candidate combines deep networking expertise with practical experience in cloud networking, compute platforms, and enterprise infrastructure. Success in this role requires … network technologies including routing, switching, wireless, firewalls, and secure remote access solutions. Lead network architecture decisions with a focus on scalability, resiliency, and security (high availability, redundancy, segmentation). Implement and enforce network security controls including segmentation, zero trust principles, and secure access patterns. Support hybrid connectivity models ...

Java Software Developer

Location
Greater London, England, United Kingdom
Fasanara Digital is a quantitative investment team applying a scientific, high frequency investment style in digital assets, seeking to achieve exceptional risk-adjusted returns for our investors. We were founded in 2018 and have grown to a 30-person strong team, managing over $500m USD in a basket … delta-neutral trading strategies. Our team members come from diverse backgrounds. We are fully dedicated to building out our globally deployed, 24/7 availability trading platform, which allows us to capture trading opportunities on >15 crypto liquidity venues, and maintain our position as one of the top trading ...

Senior DevOps Analyst

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
/CD, cloud platforms, automation, containerization, and infrastructure management. The ideal candidate will drive DevOps best practices, enable scalable deployments, and ensure high availability and performance of applications. Key ResponsibilitiesDesign, implement, and maintain CI/CD pipelines for automated build, test, and deployment. Manage and optimize cloud infrastructure … Qualifications8+ years of experience in DevOps/Site Reliability/Infrastructure Engineering. Strong understanding of DevOps and CI/CD best practices. Experience supporting high-availability, production systems. Experience in Agile/Scrum environments. Bachelor's degree in Computer Science, Engineering, or equivalent. Good to HaveExperience with DevSecOps ...

Platform Engineer

Location
Greater London, England, United Kingdom
work on afford you the ability to deliver real impact. Role As a DevOps Engineer, you bring extensive expertise in designing, deploying, and managing high-availability systems on AWS and Azure. With a strong foundation in infrastructure-as-code tools like Terraform, you have successfully built scalable, resilient …/CD pipelines to accelerate delivery. Strong communication skills allow you to bridge technical and non-technical stakeholders effectively. Responsibilities Build and maintain high-availability infrastructure for innovative organisations. Automate infrastructure management using Terraform and Ansible. Optimise containerised applications with Kubernetes and Docker. Design robust CI/ ...

Enterprise Architect

Location
Greater London, England, United Kingdom
business with a strong global reputation, an impressive client base, and ambitious growth plans. We deliver deep insights and domain-led digital transformation to high-growth and heavily regulated organisations. To our customers, we bring a partnership that provides the talent, technology, and capability to enhance performance and operational … code, automated quality gates, release orchestration, environment promotion, vulnerability scanning, and audit-ready change management. Ability to define resilience and operational readiness requirements including high availability, disaster recovery, failover, observability, monitoring, logging, tracing, SLOs, incident handling, and support transition. Good appreciation of enterprise data platforms like Databricks , governed ...