26 to 50 of 902 High Availability Jobs

Software Engineer

Location
Milton Keynes, England, United Kingdom
build distributed systems on AWS using serverless-first principles Liaise with 3rd Party Engineering teams on all aspects of the engineering lifecycle Architect for high availability, fault tolerance, and resilience by default Validate 3rd Party deliverables, workshop solutions, troubleshoot issues Design scalable data approaches and reporting structures aligned … secure, PCI-aware designs aligned with payments industry expectations Data & Infrastructure Design and maintain data pipelines and reporting structures using AWS native tooling Apply high availability and fault tolerance patterns across all infrastructure Ensure codebases and infrastructure are maintainable and transferable across engineers Technology Environment AWS (primary platform ...

IMS Mainframe Systems Programmer

Location
Knutsford, England, United Kingdom
Barclays, where you’ll maintain and optimise IMS environments to ensure the performance, stability, and resilience of critical mainframe systems. You’ll support high‐availability operations, resolve multi‐faceted incidents, and collaborate with teams across the bank while contributing to ongoing modernisation and automation initiatives. Responsibilities Maintain … optimise IMS environments to ensure performance, stability, and resilience of critical mainframe systems. Support high‐availability operations and resolve multi‐faceted incidents. Collaborate with teams across the bank while contributing to ongoing modernisation and automation initiatives. Qualifications Ample knowledge of IMS DB/DC, z/ ...

Database Administrator (short term contract)

Location
Greater London, England, United Kingdom
team on an initial three‐month contract to take ownership of the database infrastructure that underpins SilverRail’s rail distribution platforms. These systems process high volumes of search, booking and ticketing transactions for rail operators, retailers and travellers across the UK, Europe, North America and Australia/APAC … schema design, query and database optimisation, and cloud‐hosted database infrastructure from day one. We are looking for someone who has supported systems with high transaction volumes and strict availability requirements, such as travel, e‐commerce or financial services. Key Responsibilities Database Administration & Design Design, implement and maintain ...

Senior SQL Database Administrator

Hiring Organisation
F5 consultants
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
Senior SQL Database Administrator to join the team in Cardiff. You'll work across development, test and production environments, taking ownership of the performance, availability, security and reliability of mission-critical SQL Server estates. This is a senior technical position where you'll act as a SQL Server subject … managed and automated. What you'll do: Manage SQL Server databases across development, test and production environments. Configure, maintain and monitor database servers, ensuring high levels of performance, availability and security. Lead the investigation, diagnosis and resolution of complex database issues. Act as an escalation point for production ...

Principal Engineer I — Prepurchase Platform

Location
City Of London, England, United Kingdom
Prepurchase Platform group owns the backend services that power every fan's path to a ticket. We handle event discovery, real-time seat availability, and interactive venue experiences at massive scale, processing billions of API calls during peakon-sales. Our engineers build systems that must be fast, resilient … reaching a fan is a problem, not a data point; treat it accordingly and act with urgency regardless of volume. Build end-to-end, high-availability systems that handle extreme traffic spikes during high-demand on-sales without degradation. Lead platform modernization initiatives - migrating legacy services ...

L2/L3 Support Engineer| Project & Resource Management COTS| Stevenage, UK

Location
Stevenage, England, United Kingdom
across enterprise IT systems and business-critical applications. The role is responsible for incident resolution, service improvement, platform administration, and technical troubleshooting while ensuring high availability, performance, and user satisfaction. The successful candidate will act as an escalation point for complex technical issues, working closely with business stakeholders … across enterprise IT systems and business-critical applications. The role is responsible for incident resolution, service improvement, platform administration, and technical troubleshooting while ensuring high availability, performance, and user satisfaction. The successful candidate will act as an escalation point for complex technical issues, working closely with business stakeholders ...

DevOps Engineer- Night Shift(10:00 PM - 6:00 AM)

Location
United Kingdom
production environments healthy and performant, while simultaneously designing and maintaining the CI/CD pipelines, infrastructure‐as‐code frameworks, and tooling that enable rapid, high‐quality software delivery. You are the connective tissue between engineering, platform, and operations — someone who is equally comfortable in an incident bridge call … production downtime, performance degradation, and security‐related incidents in a timely, structured manner. Perform end‐to‐end operational duties covering application server health, service availability, and platform integrity in accordance with documented processes and runbooks. Review and manage client service request tickets in adherence to defined SLAs, ensuring accountability ...

Senior Data Engineer

Location
West of England, England, United Kingdom
resilience, and availability.* Troubleshoot and resolve database-related issues in development and production environments.* Implement and maintain database security controls, backup strategies, replication, and high-availability solutions.* Support change management activities and continuous improvement of data platform services.* Develop tools for data-driven insights and task automation* Collaborate … Spark, Hadoop, NiFi* Exposure to infrastructure scaling e.g. Kubernetes* Exposure to Ansible and Terraform* Understanding of ML algorithms and deployment* Experience with SQL Server high-availability technologies, including replication, mirroring, or clustering.* Experience with Windows Server administration.* Exposure to PowerShell scripting and automation.* Experience working within Agile delivery ...

DevOps Engineer, Studios

Location
Greater London, England, United Kingdom
pipelines across our digital and broadcast platforms. The successful candidate will play a key role in enabling continuous delivery, improving system reliability, and supporting high-profile clients and live event services, ensuring optimal performance and resilience across all environments. Key Responsibilities and Accountabilities Design, build, and maintain scalable cloud … efficient, reliable software delivery across multiple teams. Automate infrastructure provisioning using Infrastructure as Code tools such as Terraform, CloudFormation, or similar. Monitor system performance, availability, and reliability using observability tools such as Prometheus, Grafana, and ELK stack. Ensure high availability and disaster recovery strategies are in place ...

Senior Fullstack Engineer (Python + React.js)

Location
Greater London, England, United Kingdom
maintain scalable, secure, and efficient server-side applications. The ideal candidate should have experience in microservices architecture, API development, database management, and frontend ensuring high availability and performance of backend and frontend services. In this role, you will collaborate closely with frontend engineers, product managers, and other stakeholders … Redux or React Query Collaborate with frontend developers to ensure efficient API integration and a seamless user experience. Troubleshoot and resolve production issues, ensuring high availability and minimal downtime. Write unit and integration tests to maintain code reliability and ensure high- quality releases. Continuously monitor and optimize ...

AWS DevOps Platform Engineer - eSC/eDV Clearance

Location
Leicester, England, United Kingdom
gain hands on experience with cutting edge technologies, and deliver solutions that create real business impact. From day one, you’ll work on meaningful, high profile programmes that stretch your skills and accelerate your growth. We invest heavily in you—supporting continuous learning, in demand skills development, and long … seamless integration with DevOps toolchains and CI/CD pipelines Apply knowledge of Storage, Compute, and Security Services to administer hybrid cloud environments Configure High Availability (HA) and Disaster Recovery (DR) solutions and manage Kubernetes clusters Join our team and contribute to the development of innovative AWS infrastructure ...

Database Administrator

Location
Greater London, England, United Kingdom
solutions under the guidance of the DBA Team Participate in team meetings and proactively contribute to database strategy discussions Other requirements raised by management High Availability & Scalability Support high-availability and disaster recovery (HA/DR) solutions, including replication and failover strategies Optimize databases for scalability ...

Site Reliability Engineer, Studios

Location
Uxbridge, England, United Kingdom
project stakeholders. Key Responsibilities And Accountabilities Design, build, and maintain reliable, scalable infrastructure and platform services across on‐premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering‐led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards … Drive root cause analysis and corrective actions following incidents, with a focus on prevention and continuous improvement. Support the design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements. Help enforce security, access control, patching, and operational best practices across infrastructure and services. Optimise system ...

Site Reliability Engineer, Studios

Location
Greater London, England, United Kingdom
project stakeholders. Key Responsibilities and Accountabilities Design, build, and maintain reliable, scalable infrastructure and platform services across on-premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering-led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards … Drive root cause analysis and corrective actions following incidents, with a focus on prevention and continuous improvement. Support the design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements. Help enforce security, access control, patching, and operational best practices across infrastructure and services. Optimise system ...

Cloud Infrastructure Consultant

Hiring Organisation
Hackajob Ltd
Location
York, North Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
alongside them. Working knowledge of Microsoft 365 is desirable. You will work closely with pre-sales, Principal Consultants, Architects, and operational teams to deliver high quality, supportable solutions aligned to agreed standards. Provide hands-on design and delivery expertise across datacentre, hybrid, and Azure environments, acting as the customer … Services, Group Policy, DNS, DHCP, certificate services, and hybrid identity through Entra Connect. Design and deploy resilient compute and storage platforms, including failover clustering, high availability, shared and software-defined storage, and the backup and replication that underpins them. Lead and support migration of customer server estates, including ...

platform engineer in hybrid cloud

Location
Greater London, England, United Kingdom
deployment models Integrate applications with DevOps toolchains and CI/CD pipelines Administer hybrid cloud environments using Storage, Compute, and Security Services Configure High Availability and Disaster Recovery solutions Manage Kubernetes clusters Contribute to innovative AWS infrastructure solutions that drive business success. Требования: Hands‐on experience designing … cloud‐native architecture patterns Experience integrating CI/CD pipelines and DevOps toolchains such as GitLab CI, Jenkins, or GitHub Actions Experience implementing High Availability and Disaster Recovery strategies in AWS environments Strong Linux systems knowledge and troubleshooting capability in cloud and container environments Ability to work collaboratively ...

Cloud Infrastructure Consultant

Location
Birmingham, England, United Kingdom
alongside them. Working knowledge of Microsoft 365 is desirable. You will work closely with pre-sales, Principal Consultants, Architects, and operational teams to deliver high quality, supportable solutions aligned to agreed standards. What You Will Do Provide hands‐on design and delivery expertise across datacentre, hybrid, and Azure environments … Services, Group Policy, DNS, DHCP, certificate services, and hybrid identity through Entra Connect. Design and deploy resilient compute and storage platforms, including failover clustering, high availability, shared and software‐defined storage, and the backup and replication that underpins them. Lead and support migration of customer server estates, including ...

Sr. Database Architect

Location
United Kingdom
opportunities for improvement, and establish database engineering best practices across our products. A primary responsibility of this role is to optimize and maintain our high-volume OLTP databases, with a strong emphasis on PostgreSQL. You will assess existing environments, identify performance and scalability bottlenecks, and drive improvements across schema … database-backed features. You will help strengthen production readiness by establishing and improving practices around performance testing, release validation, monitoring, alerting, backup and recovery, high availability, access controls, and incident response. You will play a key role in troubleshooting complex production issues, performing root-cause analysis, and ensuring ...

Engineer - Site Reliability

Location
Greater London, England, United Kingdom
support model for its US Global Trading Hours (GTH) markets, providing critical overnight and early‐session coverage from London that ensures continuous, high‐availability operations across Cboe's real‐time low‐latency trading platforms. The London‐based SRE provides technical support to Cboe Trade Desk and Operations Support … timely, precise communication to stakeholders during active incidents and contribute to post‐incident reviews and remediation tracking to drive long‐term platform stability. System Availability & Technical Support: Provide technical support and operational oversight to sustain resiliency and high availability of critical business operations. Monitor production, disaster recovery ...

DBA Lead

Location
City Of London, England, United Kingdom
https://risk.lexisnexis.com/insurance About our Team The Database Administration team provides enterprise-wide support for mission-critical database platforms, ensuring high availability, performance, security, and operational stability. The team manages MySQL, SQL Server, PostgreSQL, Oracle, and cloud database technologies across hybrid environments while driving automation … Additional Skills Database Technologies 7+ years of MySQL DBA or Database Engineering experience. Strong experience with MySQL replication design, implementation, and support. Experience with High Availability (HA) and Disaster Recovery (DR) architectures. Extensive experience with backup and recovery technologies including Percona XtraBackup and snapshot-based recovery solutions. Experience ...

Sr. IT Systems Engineer (On-site, Lawrenceville GA)

Hiring Organisation
Lendmark Financial Services, LLC
Location
Lawrenceville, Georgia, United States
Employment Type
Permanent
Salary
USD Annual
Technologies Site-to-Site VPN Connectivity Cloud-Managed Network Monitoring and Administration Manage and support physical and virtual network infrastructure, ensuring security, high availability, performance, and reliability. Monitor, troubleshoot, and optimize enterprise infrastructure across cloud, server, identity, networking, and virtualization platforms. Collaborate with Information Security teams to implement … engineering staff. Oversee installation, configuration, patching, lifecycle management, and performance optimization of enterprise infrastructure platforms. Support Azure-based virtualization technologies. Maintain accountability for infrastructure availability, disaster recovery readiness, business continuity planning, and operational excellence. Manage relationships with technology vendors, consultants, cloud providers, and managed service partners. Support Citrix technologies ...

Data Operations Engineer

Hiring Organisation
Wilson Elser - Business & Legal Professionals
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
ideal candidate will have strong expertise in Microsoft SQL Server (On-Premises), Azure SQL, and Azure Databricks, with a focus on ensuring data platform availability, performance, security, and operational excellence. This role is responsible for maintaining critical database infrastructure, supporting data engineering workloads, automating operational processes, monitoring platform health … Managed Instance environments. Monitor database health, performance, storage utilization, and capacity planning. Perform database installation, configuration, patching, upgrades, and migrations. Configure and manage high availability and disaster recovery solutions, including SQL Server Always On Availability Groups, failover clustering, backup, and recovery. Develop and maintain backup, restore ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
team responsible for network architecture, deployment, and operational readiness. You will play a key role in ensuring connectivity solutions align with performance, security, and availability requirements. The ideal candidate combines deep networking expertise with practical experience in cloud networking, compute platforms, and enterprise infrastructure. Success in this role requires … network technologies including routing, switching, wireless, firewalls, and secure remote access solutions. Lead network architecture decisions with a focus on scalability, resiliency, and security (high availability, redundancy, segmentation). Implement and enforce network security controls including segmentation, zero trust principles, and secure access patterns. Support hybrid connectivity models ...

Java Software Developer

Location
Greater London, England, United Kingdom
Fasanara Digital is a quantitative investment team applying a scientific, high frequency investment style in digital assets, seeking to achieve exceptional risk-adjusted returns for our investors. We were founded in 2018 and have grown to a 30-person strong team, managing over $500m USD in a basket … delta-neutral trading strategies. Our team members come from diverse backgrounds. We are fully dedicated to building out our globally deployed, 24/7 availability trading platform, which allows us to capture trading opportunities on >15 crypto liquidity venues, and maintain our position as one of the top trading ...

Senior / Lead Site Reliability Engineer

Location
Watford, England, United Kingdom
role At Allwyn, the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate, ensuring high availability, performance, and resilience of customer-facing systems during both normal operation and peak lottery events. The role combines hands-on engineering, incident leadership … reporting, working across platform, product, and operational teams. Objectives of the role Own reliability outcomes across services using SLOs, SLIs, and error budgets Improve availability, latency, and scalability across Instant-Win and Draw-based platforms Lead incident response and operational readiness, including peak jackpot events Drive automation and platform ...