151 to 175 of 475 High Availability Jobs

Sr Lead Software Engineer - KDB+ / Q

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems Develops secure high‐quality production code, and reviews and debugs code written by others Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall … experience developing/running large datasets and optimizing query performance. Practical experience scaling and load‐balancing of KDB applications. Practical experience building resilient and highavailability KDB applications. Preferred qualifications, capabilities, and skills Experience with market data venue and vendor data platforms. AWS Experience. Experience in Terraform ...

DevOps Engineer

Hiring Organisation
ISR Recruitment
Location
United Kingdom
excellent opportunity to work within a modern AWS cloud environment, supporting the deployment, monitoring and continuous improvement of live services used across a high-profile government programme. The successful candidate will be a hands-on DevOps Engineer with strong AWS expertise, excellent troubleshooting skills and experience supporting production environments. … other highly regulated environments Role and Responsibilities: Support the deployment, operation and continuous improvement of cloud-based services hosted within AWS. Monitor the health, availability and performance of live production environments. Investigate, diagnose and resolve infrastructure, application and performance issues. Collaborate with development teams to identify root causes ...

DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
continuous improvement of our deployment pipelines, cloud and bare metal environments, and Kubernetes‐based orchestration layer — ensuring our platform scales reliably to support high‐performance AI and GPU workloads. You'll set the standard for engineering excellence, partnering with Platform, SRE and Product Engineering teams to align on strategy … Ansible, ensuring infrastructure is version‐controlled, peer‐reviewed, and consistently provisioned across dev, staging, and production Manage and maintain hosted application environments, ensuring high availability, scalability, and smooth deployment of services into production Partner closely with Go engineers to deeply understand their development workflows and release processes, identifying ...

Principal Platform Security Engineer

Hiring Organisation
Jobleads-UK
Location
York and North Yorkshire, England, United Kingdom
DevOps/Platform Engineering experience delivering solutions in Azure and/or GCP. Full‐stack application and infrastructure solution design with robust security controls, high availability, and operational resilience. Working knowledge of vulnerability and compliance management (scanning to remediation), patch management, endpoint protection/anti‐malware, and access … delivery focus, capable of prioritizing effectively and delivering outcomes in a fast‐paced environment with shifting demands. Ability to operate effectively in a small, high‐impact team while collaborating across a wider product/engineering organisation. Excellent communication and stakeholder‐management skills, able to influence at all levels ...

Platform Engineer - Hashi Vault IRC298441

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
functional stakeholders. Job responsibilities Architecture & Infrastructure Management Deploy and Maintain: Manage the lifecycle of HashiCorp Vault clusters across multiple environments (Dev, Stage, Prod). High Availability: Ensure 99.9% uptime through proper configuration of storage backends (like Raft or Consul) and unseal workflows (Auto‐unseal via KMS/Cloud … offer Our goal is to build an inclusive positive culture where everyone can feel comfortable being themselves, empowering our people to create their own high standards and therefore more value. We work together to promote fairness while recognising, valuing and embracing differences – providing a transparent support structure and generous ...

Site Reliability Engineer

Hiring Organisation
ISR Recruitment
Location
United Kingdom
excellent opportunity to work within a modern AWS cloud environment, supporting the deployment, monitoring and continuous improvement of live services used across a high-profile government programme. The successful candidate will be a hands-on engineer with strong AWS expertise, excellent troubleshooting skills and experience supporting production environments. … call back in the strictest confidence. If you're an experienced Site Reliability Engineer (DevOps) with strong AWS expertise and a passion for supporting high-availability cloud services; please contact Edward Laing here at ISR to learn more about our client and how they are leading ...

Principal Platform Security Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
DevOps/Platform Engineering experience delivering solutions in Azure and/or GCP.* Full‐stack application and infrastructure solution design, ensuring robust security controls, high availability, and operational resilience throughout.* Working knowledge of vulnerability and compliance management (scanning through to remediation), patch management, endpoint protection/anti-malware … delivery focus, able to prioritise effectively and deliver outcomes in a fast-paced environment with shifting demands.* Able to operate effectively in a small, high-impact team while collaborating across a wider product/engineering organisation.* Excellent communication and stakeholder-management skills, able to influence at all levels ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
commercial and residential installation companies. Access to our unique on-line portal and design tool, plus our emphasis on product quality, consistency and availability sets us apart in the market. Segen is a fast-moving business that responds quickly to any market changes supported by its bespoke ERP system … guidance. Promote cloud security best practices and ensure compliance within infrastructure design. Contribute to the evolution of our monitoring and alerting systems to maintain high availability and performance. What We’re Looking For Deep experience in Microsoft Azure, including Hub and Spoke networking. Strong proficiency with Azure DevOps ...

Senior Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
exchange. As a key part of the Robinhood family, Exchange Platform team owns the full service lifecycle. We are the architects of a modern, high-velocity ecosystem that enables our global expansion, ensuring the world’s longest-running exchange remains unshakeable. The Role As a Senior Backend Engineer integrated … essential part of building the new generation infrastructure for low-latency services in the cloud, adopting cutting-edge technologies to drive our high-velocity ecosystem. This role is based in our London office(s), with in-person attendance expected at least 3 days per week. At Robinhood, we believe ...

Senior Cloud Platform Engineer (GCP | Kubernetes | DevSecOps)

Hiring Organisation
Jobleads-UK
Location
Bolsterstone, England, United Kingdom
service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability ...

Senior Cloud Platform Engineer (GCP | Kubernetes | DevSecOps)

Hiring Organisation
GCS
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £620/day
service mesh technologies and secure service-to-service communication. Manage ingress, API gateways and network routing components. Implement platform resilience, backup, disaster recovery and high availability capabilities. Support workload identity and platform authentication mechanisms. Ensure platform compliance with enterprise security, governance and regulatory requirements. Optimise platform performance, scalability ...

Senior IT Technician

Hiring Organisation
IT Talent Solutions
Location
Walsall, West Midlands (County), United Kingdom
Employment Type
Permanent
Salary
£40000 - £47000/annum + Bens
cloud services, vendor management, and continuous improvement. Key Responsibilities Own and manage the Group's IT infrastructure across on-premise and cloud environments. Ensure high availability, performance, resilience, and security of all IT systems. Lead cybersecurity initiatives including MFA, endpoint protection, patching, and vulnerability management. Manage Microsoft ...

Kafka Admin Lead-6months-London

Hiring Organisation
Kirtana Consulting
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
Kafka KRaft or ZooKeeper-backed if Legacy), Confluent Platform components (Schema Registry, Kafka Connect, ksqlDB, REST Proxy, Control Center), and Confluent Cloud resources. Ensure high availability (HA) and resilience across multi AZ/region deployments; manage partition replication, min.insync.replicas (ISR), and rack awareness. Implement and maintain client quotas ...

Frontend Architect - React

Hiring Organisation
HCLTech
Location
London Area, United Kingdom
Developer (React, TypeScript and Low-Latency) Role Overview- We are seeking a highly skilled UI Developer with strong expertise in building time-critical, high-availability, and high-performance applications. The ideal candidate will have hands-on experience across the full UI development lifecycle, including design, development, testing ...

DevOps Engineer

Hiring Organisation
Oscar Associates (UK) Limited
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£70,000
opportunity to take ownership of a clean, containerised environment, shape how it's operated, and play a key role in scaling a high-traffic platform as the business continues to grow. If you enjoy building reliable, automated cloud infrastructure without the overhead of managing Kubernetes clusters, this could … moves into full production. Working closely with engineering teams, you'll drive automation, improve deployment pipelines, strengthen observability and ensure the platform performs under high-volume, real-time workloads. This is a hands-on position with genuine ownership and plenty of opportunity to influence how the platform evolves. What ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while partnering closely with application development teams. We believe in operations, infrastructure … scale, highly available distributed systems that allow the ThousandEyes platform to process a constantly growing volume of telemetry data. Use AI Tooling to write high‐quality code and automated solutions that enable fast, reliable releases, reduce operational expense, and allow our infrastructure and platforms to scale thoughtfully across regions. ...

Senior / Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Watford, England, United Kingdom
role At Allwyn, the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate, ensuring high availability, performance, and resilience of customer-facing systems during both normal operation and peak lottery events. The role combines hands‐on engineering, incident leadership … reporting, working across platform, product, and operational teams. Objectives of the role Own reliability outcomes across services using SLOs, SLIs, and error budgets Improve availability, latency, and scalability across Instant-Win and Draw-based platforms Lead incident response and operational readiness, including peak jackpot events Drive automation and platform ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while partnering closely with application development teams. We believe in operations, infrastructure … scale, highly available distributed systems that allow the ThousandEyes platform to process a constantly growing volume of telemetry data.Use AI Tooling to w rite high-quality code and automated solutions that enable fast, reliable releases, reduce operational expense, and allow our infrastructure and platforms to scale thoughtfully across regions. ...

DevOps Engineer

Hiring Organisation
DGH Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£80000 - £100000/annum
MLOps Engineer (Platform/DevOps) - AI Platform - FULL TIME ONSITE We're partnering with a leading organisation on a high-profile AI initiative, seeking an experienced MLOps/Platform Engineer to build and scale infrastructure supporting advanced AI solutions. The Role You'll play a key role in designing … platform for deploying AI workloads and agents. Working closely with Data Science and Engineering teams, you'll ensure robust infrastructure, seamless deployment pipelines, and high platform reliability. Key Responsibilities - Design, deploy, and manage AI platforms and agent infrastructure - Build and maintain CI/CD pipelines and DevOps workflows - Implement ...

Senior Site Reliability Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
real‐time low‐latency trading platforms around the clock and provides direct platform support for Cboe's European operations. Responsibilities Technical Leadership & System Availability: Provide technical leadership, support and operational oversight to sustain resiliency and high availability of critical business operations across European and GTH market sessions ...

Principal Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
strategy as we scale our global fintech platform. You will be the primary technical authority for our system design, responsible for ensuring that our high-throughput transaction engines and complex microservices ecosystem remain resilient, performant, and secure under massive load. You will partner with Engineering, Product, and Security leadership … Doing Own the Architectural Vision: Define and drive the long-term technical architecture for KAST, ensuring our platform evolves to meet the demands of high-throughput transactional volume. Design for Scale & Reliability: Lead the design and review of our distributed systems, ensuring our microservices architecture is built for fault ...

Lead Software Engineer - GLCM Postings

Hiring Organisation
Jobleads-UK
Location
Bournemouth, England, United Kingdom
industry best practices Lead and mentor a team of software engineers, providing guidance, support, and professional development opportunities Drive the design and implementation of high-performance backend services using Java and cloud-native technologies Champion the development of highly resilient, share-nothing, multi-region architectures to achieve fault tolerance … high availability, and disaster recovery Drive the adoption and optimization of MongoDB (including change streams), Kubernetes, and other modern platforms Set and enforce standards for API design and development (REST/gRPC), including best practices for versioning, documentation, and error handling to ensure robust, secure, and scalable interfaces ...

System Engineer

Hiring Organisation
Venturi
Location
Manchester Area, United Kingdom
maturing Nice to have: Background in games, VFX, animation or another technical creative studio environment Exposure to Perforce or large-scale build-pipeline infrastructure High-availability and resilience design experience Comfortable using AI tooling to support infrastructure-as-code work If you're a systems generalist who enjoys … building structure into a high-change environment and wants the autonomy of an outside-IR35 contract, get in touch. ...

Staff Software Engineer - Commercial Trading

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
journey, creating solutions for the business that are robust and scalable, with good observability and metrics, following best-in-class engineering practice. Due to high interest, this role may close earlier than advertised. We recommend applying as soon as possible. Your key accountabilities will include: Lead the architecture, design … delivery of scalable software solutions, distributed systems and modernisation initiatives that support critical store operations across M&S. Provide technical leadership on complex, high-impact projects, driving engineering excellence through code reviews, deployment optimisation and adoption of engineering best practices. Collaborate with stakeholders across Product, Delivery, Architecture, Infrastructure ...

Member of Technical Staff (AI Infrastructure Engineer)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands Qualifications Strong expertise in Kubernetes … resource allocation, and cluster optimization Experience with deploying and managing distributed training systems at scale Deep understanding of container orchestration and distributed systems architecture High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies) Experience managing GPU clusters and optimizing ...