76 to 100 of 262 Site Reliability Engineering Jobs in London

Production Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance. This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements … approach to the day-to-day, with the resilience to handle high-pressures production incidents Desired Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

Principal Solutions Engineer - Observe by Snowflake

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
sales cycle, from initial contact to final contract. Feedback and Product Collaboration: Serve as the voice of the customer to the Product Management and Engineering teams, providing crucial feedback on product capabilities, market needs, and competitive landscape. Content and Training: Develop and maintain technical sales collateral, including demo environments … articulate complex technical concepts to both technical and non-technical audiences. Preferred Experience selling to Developer, DevOps, and Site Reliability Engineering (SRE) personas. Prior experience in a fast-paced, high-growth Observability or Application Performance Monitoring (APM) company. Snowflake is growing fast, and we’re scaling ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI‐powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … call rotation. Hands‐on experience with infrastructure‐as‐code tooling and codebases, preferably Terraform. Hands‐on experience leveraging AI as a force multiplier of SRE activities, such as automating toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems, networking ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … operational on-call rotation. Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform. Hands-on experienceleveraging AIas a force multiplier of SRE activities, such as automati ng toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems ...

AMBG - Cloud Security & Exposure Management Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
recovery sequencing. Identify risks, vulnerabilities, and single points of failure across workloads and operational processes. Recommend improvements aligned with Azure Well-Architected Framework, SRE principles, and ITIL practices. Engage customer stakeholders to understand RTO/RPO objectives and recovery workflows. Produce professional documentation outlining findings, risks, and recommended improvements. About … Azure architecture including availability zones, backup, recovery, and monitoring services. Familiarity with cloud-native resiliency patterns and site reliability engineering (SRE) practices. Experience designing and assessing Major Incident Response Plans (MIRPs). Experience in business continuity planning and operational resilience. Strong communication and documentation skills across technical ...

Hybrid SRE Manager: Scale Reliability & Platforms

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
Holland & Barrett is seeking a Site Reliability Engineering Manager to lead a high … performing team and drive reliability, scalability and security across cloud platforms and digital services. You will champion modern engineering practices, shape the SRE roadmap, and partner with Engineering, Security, Data and Product teams to embed operational excellence from design through production. This role combines technical leadership with ...

Systems Operations Lead

Hiring Organisation
Hays Technology
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£750 - £800/day Up to £800pd inside ir35 via umbrella
team of technical SMEs, ensuring workloads are prioritised and delivered effectively. Act as an escalation point for operational and infrastructure-related issues. Drive service reliability, operational excellence and continuous improvement across the environment. Required Experience Strong infrastructure background with … experience across Linux and Windows server environments. Good understanding of storage, backup and wider infrastructure technologies. Experience in Site Reliability Engineering (SRE), Infrastructure Operations, or Production Support environments. Proven experience leading and developing technical teams. Comfortable remaining hands-on and involved in technical delivery on a daily ...

Systems Operations Lead

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East, England, United Kingdom
Employment Type
Contractor
Contract Rate
£750 - £800 per day
team of technical SMEs, ensuring workloads are prioritised and delivered effectively. Act as an escalation point for operational and infrastructure-related issues. Drive service reliability, operational excellence and continuous improvement across the environment. Required Experience Strong infrastructure background with … experience across Linux and Windows server environments. Good understanding of storage, backup and wider infrastructure technologies. Experience in Site Reliability Engineering (SRE), Infrastructure Operations, or Production Support environments. Proven experience leading and developing technical teams. Comfortable remaining hands-on and involved in technical delivery on a daily ...

Tech Operations & SRE Leader | Reliability & Resilience

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Site Reliability Engineering to define the operating model for UK Tech. You will bring together service management, production support and SRE to improve stability and reliability across critical services and digital journeys. You will set SLIs/SLOs, govern incident and change practices, embed observability ...

Software Engineer-MUST HAVE CURRENT DV CLEARANCE

Hiring Organisation
Reed
Location
London, South East, England, United Kingdom
Employment Type
Temporary
Salary
£597 per day
Development Specialist/Professional Location: London Permanent | Full-Time | 37.5 Hours per Week Live DV Clearance Essential Our client are an expanding high-performing engineering team supporting mission-critical environments and are looking for talented Software Development Specialists and Engineering Professionals who are passionate about building, securing … within a software-defined, automation-first environment where infrastructure, security and applications are delivered through code. You'll join a team focused on modern engineering practices, CI/CD, Infrastructure as Code, cloud-native technologies and security-by-design principles. Please note: A current, live DV (Developed Vetting) Clearance ...

Principal SRE (AWS, Azure, Terraforms, Kubernetes)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
China, Australia, and UAE. Interested in joining our smart, fun, and talented team? Position Overview Fourth is actively seeking an experienced and pragmatic Principal SRE to join our worldwide team. We are progressing rapidly in developing automated, highly reliable, and zero-downtime infrastructure pipelines that are becoming the standard across … valuable and achievable chunks. You have excellent written and verbal communication skills, allowing you to work effectively with our worldwide development teams and SRE community to select the right patterns and practices. You understand the importance of standardisation of technology and practices and have experience of implementing these ...

SRE Architect: Cloud & Data Reliability Leader

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Hitachi Digital Services seeks a senior SRE/DevOps professional to lead the Site Reliability Engineering practice, transforming operations toward proactive, engineering-led reliability. You will define NFRs with FMEA-based methods, champion observability, self-healing automation, and automated incident management, while ensuring cost efficiency ...

Technical Lead

Hiring Organisation
17918
Location
London, United Kingdom
design and delivery of complex digital solutions for some of the world's most recognised organisations. This is an opportunity to combine hands-on engineering with technical leadership, helping clients solve difficult business challenges while mentoring high-performing teams. You'll work across a variety of industries, technologies … family Annual performance bonus Share ownership opportunities Generous pension scheme 25 days annual leave plus the option to purchase additional days Collaborative and supportive engineering community How You'll Succeed as a Technical Lead Lead multidisciplinary engineering teams through the delivery of complex technology solutions Partner with clients ...

Trainee DevOps Engineer | No experience needed (Ref: 7501)

Hiring Organisation
Qualify Nation Recruitment
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£28,000 - £38,000 per annum
Platforms (AWS, Microsoft Azure and Google Cloud) Configuration Management Monitoring and Logging Security Best Practices (DevSecOps) Networking Fundamentals Automation and Scripting Incident Management and Reliability Engineering Practical Experience You will work on realistic DevOps projects that may include: Building CI/CD pipelines Deploying applications to cloud environments … completion, learners may pursue roles such as: Junior DevOps Engineer DevOps Engineer Cloud Support Engineer Platform Engineer Infrastructure Engineer Site Reliability Engineer (SRE) Build and Release Engineer Cloud Operations Engineer Systems Administrator Cloud Infrastructure Engineer Apply Today If you are looking to start a career in DevOps ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
capital providers and quota share partners make us nimble. Our breadth of expertise and capabilities deliver outstanding market returns. The role The Analytics Product Engineering team at The Fidelis Partnership (TFP) builds and manages a bespoke analytical platform that powers our partnership-driven business model. We combine actuarial expertise … with advanced technology to deliver innovative, scalable solutions. The Platform Engineer partners with the Analytics product engineering team to design, build and operate platform capabilities that enable our market leading high-performance, distributed computing products to run reliably, securely and efficiently at scale. This includes infrastructure automation, runtime orchestration ...

Azure DevOps Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
organisation to recruit an experienced Azure DevOps/Platform Engineer for an initial 6-month contract. This role sits within a high-performing Cloud Engineering team, working on secure, scalable, cloud-native systems that support critical, real-world applications. This is an opportunity to contribute to meaningful, high-impact … experience with AKS Deep understanding of DevOps and platform engineering principles Strong knowledge of cloud security best practices Experience with observability, monitoring, and SRE concepts Why Apply Fully remote within the UK Work with a mission‐driven, highly respected health tech organisation Modern cloud environment with cutting‐edge tooling ...

Platform & Workplace Engineering Director

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Overview Leading and developing two engineering teams — Cloud Platform (cloud infrastructure, observability, SRE, incident response) and Enterprise Services (IT support, IAM, Okta, Slack and core SaaS) — setting technical direction, hiring and performance standards across both. Delivering the reliability, performance and cost-efficiency of the cloud platform underpinning production … post-incident learning, in line with the broader infrastructure strategy. Executing the roadmap for enterprise services and identity, treating internal IT as an engineering discipline — automation-first, self-service where possible, and measured on employee productivity and security outcomes. Tracking and reporting meaningful operational metrics — availability, MTTR, change failure ...

Staff Platform Engineer, UK

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
years leading engineering teams in startups, and that has always included being close to infrastructure teams - no matter what name they’ve worn (SRE, infrastructure, platform, etc). I’ve got my hands dirty building the initial infrastructure for startups and know the value a talented infrastructure engineer brings. … have to balance reliability with flexibility. Software and its availability are now mission critical to almost every working professional. To be in an SRE in today’s world, you have to be extremely comfortable evaluating risk, those you take and those others take. Why you should or shouldn ...

DevOps Engineer, Studios

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
infrastructure and delivery pipelines across our digital and broadcast platforms. The successful candidate will play a key role in enabling continuous delivery, improving system reliability, and supporting high-profile clients and live event services, ensuring optimal performance and resilience across all environments. Key Responsibilities and Accountabilities Design, build … software delivery across multiple teams. Automate infrastructure provisioning using Infrastructure as Code tools such as Terraform, CloudFormation, or similar. Monitor system performance, availability, and reliability using observability tools such as Prometheus, Grafana, and ELK stack. Ensure high availability and disaster recovery strategies are in place and tested regularly. Collaborate ...

Senior Site Reliability Engineer — Hybrid Kubernetes & GitOps

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Cisco’s Webex Engineering Group in London is seeking a Senior Site Reliability Engineer to own the design, deployment, and operation of Kubernetes-based microservices, delivering reliability and scalable deployments in a hybrid work environment. You will drive GitOps workflows with Argo CD, use Helm ...

Sr. Technical Support Engineer - Confluent

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
resolve complex Apache Kafka and Confluent Platform issues. They manage high-priority production cases, lead difficult escalations, build customer trust, and collaborate with Engineering, Product, and other internal teams What You’ll Do In this role you will: Own and manage complex customer cases across production and non-production … manage resolution plans, including priorities, dependencies, risks, mitigations, and timelines. Communicate clearly with customers and set expectations during high-pressure situations. Collaborate with Engineering and Product to identify bugs, improve product quality, and provide field feedback. Create and review technical documentation, knowledge articles, and training material. Mentor and coach ...

Global SRE Manager - Real-Time Trading Platform Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Cares, Inc is seeking a Sr. Manager of Site Reliability Engineering in London. The role entails overseeing a distributed team of SRE technologists to ensure high availability of trading platforms. Candidates should possess a Bachelor's Degree in a related field and deep expertise in technical operations ...

DevOps Engineer

Hiring Organisation
TRIA
Location
City of London, London, United Kingdom
DevOps Engineer Up to £85,000 + Bonus + Excellent Benefits We're supporting a growing technology consultancy that delivers complex engineering and cyber programmes across Defence and National Security. As demand for their services continues to grow, they're looking to recruit a DevOps Engineer to help build … work, you'll need to either have current DV Clearance. The Headlines A hands-on DevOps role focused on automation, cloud infrastructure and modern engineering practices, working alongside experienced engineers on secure, high-impact technology programmes. You'll be given real ownership, exposure to modern tooling and the opportunity ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
significant contributions to project goals, focusing on cloud platform engineering. This individual demonstrates a solid understanding of the principles and practices of platform engineering, especially within cloud environments, and shows proficiency in specific technical areas related to cloud infrastructure, automation, and scalability.Key responsibilities include at least … following: Cloud platforms:Community Involvement: Actively participating in the professional cloud platform engineering community, contributing insights, and staying abreast of the latest trends and best practices.Team Contribution: Making significant contributions to team objectives, particularly in designing, building, and maintaining cloud based platforms and infrastructure.Technical Proficiency: Exhibiting a good grasp ...

Senior Site Reliability Engineer, Scalable Infra

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Cisco ThousandEyes is seeking a seasoned Site Reliability Engineer to design, operate and scale large‐scale distributed systems that process telemetry data at high volumes. You will work across AWS, Kubernetes, and infrastructure‐as‐code, applying AI to automate toil … improve reliability. Collaboration with software engineers and on‐call incident management are core parts of the role. Ideal candidates have 5+ years in SRE/DevOps, strong coding skills in Python or Go, and demonstrable #J-18808-Ljbffr ...