201 to 225 of 1,104 Site Reliability Engineering Jobs in the UK

GenAI Platform Operations Lead

Location
United Kingdom
circa £65,000 - £74,000 Locations: Norwich/Leeds/Eastleigh/Bristol/Leatherhead/York A bit about the job The Oasis SRE Technical Lead is a senior hands-on technical leadership role responsible for the reliability, security, scalability, and operational excellence of the Oasis Platform. Working … automation, observability, DevOps practices, and AI-powered operations while leading and mentoring a high-performing team. This role is ideal for an experienced SRE, DevOps, Platform Engineering, or Cloud Operations professional who is passionate about solving complex challenges and enhancing platform resilience at scale. Skills and experience ...

Senior Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Senior Platform Engineer (12 Month FTC)

Location
London, United Kingdom
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Principal Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Principal Platform Engineer (12 Month FTC)

Location
Greater London, England, United Kingdom
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Site Reliability Engineer

Location
United Kingdom
role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight. Working across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across … hybrid work from home policy. Preferred Skills and Experience Knowledge of modern development practices, including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Location
United Kingdom
strong Senior and Lead Cloud Platform Engineers for future opportunities. Role Overview We are seeking an experienced Cloud Platform Engineer to join our platform engineering team. This is a hands‐on role focused on designing, building, operating, and continuously improving cloud‐native platforms that enable development teams to deliver … using Java, Kotlin, Python, or similar technologies Experience implementing cloud security controls, governance, and compliance requirements Exposure to site reliability engineering (SRE) practices and platform engineering frameworks What we’ll offer you: We trust people to do their best work. That means flexibility over rigid rules ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
Site Reliability Engineer (SRE) - Assistant Vice President is a technical professional responsible for the hands‐on execution, technical implementation, and deployment of SRE and observability principles in a complex, critical, and large-scale multi-disciplinary environment. In this role, you will apply a deep understanding of multiple technology … authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical onboarding support to application teams. We are seeking a passionate and experienced SRE to join our Production Management team. In this role, you will be instrumental in executing our strategy for end‐to‐end observability and resiliency, collaborating ...

AI Technical Platform Leader

Location
Greater London, England, United Kingdom
business stakeholders to mold and implement strategy. The role applies broad technical knowledge with depth in generative AI, agentic systems, enterprise platforms and engineering governance to ensure that AI platforms at WTW are optimally configured to achieve company vision and imperatives. It has end-to-end ownership of implementation … operations in a large, global enterprise: Software or platform engineering for global, enterprise-scaled solutions Management of Site reliability engineering (SRE) programs for mission-critical systems Creation and management of DevOps practices for automated, consistent, and secure solution deployment in regulated environments. Literacy in global compliance ...

Site Reliability Engineer

Location
Fenny Stratford, England, United Kingdom
seeking an experienced Site Reliability Engineer (SRE) to join our Group Technology Team in Milton Keynes. ConnellsX is Connells Group Technology’s internal developer platform, built on Microsoft Azure. It simplifies cloud hosting, embeds security and compliance by default, and enables a frictionless developer experience. As part … operating this platform, you will play a hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Reliability testing and capacity planning Toil reduction We offer a hybrid working arrangement with one day per week ...

Senior Software Development Engineer (SRE)

Location
Cambridge, England, United Kingdom
patterns that standardize and elevate the resilience of our SaaS products across multiple regions and environments. Cultivate a shared responsibility model where the SRE team collaborates with and educates engineering teams on reliability best practices. Contribute to incident response and management, ensuring rapid resolution, clear stakeholder communication … improve operational efficiency and scalability. Champion Service-Oriented Organization (SOO) principles to ensure accountability and clarity in service ownership. Qualifications 6+ years in SRE, DevOps or related role in a large-scale environment Software development experience(ideally working with and as a .NET developer) Strong understanding of SDLC, microservice ...

DevOps Engineer

Location
Knutsford, England, United Kingdom
Join us as a DevOps Engineer at Barclays, where you will design, build, and support secure, scalable infrastructure platforms, ensuring reliability, availability, and performance through automation and engineering best practices. To be successful as a DevOps Engineer, you should have: Proven experience designing, building and maintaining CI/… using DevOps toolchains, including GitLab, Maven, SonarQube and Wiz. Some other highly valued skills may include: Knowledge of Site Reliability Engineering (SRE) principles, including reliability, scalability and operational resilience. Experience embedding security, compliance and DevSecOps practices into software delivery pipelines. Expertise in monitoring, observability and performance ...

Senior SRE Engineer - ASE Traffic & Secure Services Network

Location
Greater London, England, United Kingdom
Senior SRE Engineer - ASE Traffic & Secure Services Network Shanghai, Shanghai, China Software and Services At Apple, we build systems that power services used by hundreds of millions of people around the world, and every second counts. The Services Engineering organization is at the heart of this mission, ensuring … platforms are performant, secure, and always available. We're seeking a technically strong Site Reliability Engineer (SRE) to join our growing London team, focused on the future of traffic management, load balancing, and secure networking infrastructure.You’ll play a key role in shaping the next generation ...

Site Reliability Engineer

Hiring Organisation
Connells Limited
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
seeking an experienced Site Reliability Engineer (SRE) to join our Group Technology Team in Milton Keynes. ConnellsX is Connells Group Technologys internal developer platform, built on Microsoft Azure. It simplifies cloud hosting, embeds security and compliance by default, and enables a frictionless developer experience. As part … operating this platform, you will play a hands-on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post-incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight. Working across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across … background with Python, Golang, JavaScript or similar language. Knowledge of modern development practices, including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight. Working across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across … background with Python, Golang, JavaScript or similar language. Knowledge of modern development practices, including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. ...

Sr Lead AI Platform Engineer

Location
Auchentibber, Scotland, United Kingdom
team moves from shipping individual use cases to running multiple production platforms and the production tail of new use cases, you will set the engineering standard for deployment, scalability, security, and reliability, and lead the practices that keep our services running. This is a Senior VP-level role … Computer Science, Engineering, or a related technical field (or equivalent applied experience) Preferred qualifications, capabilities, and skills Site Reliability Engineering (SRE) experience and familiarity with reliability practices (SLOs, error budgets) Experience within financial services technology Familiarity with JPM-internal platform, cloud, and AI/ ...

Sr Lead AI Platform Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
team moves from shipping individual use cases to running multiple production platforms and the production tail of new use cases, you will set the engineering standard for deployment, scalability, security, and reliability, and lead the practices that keep our services running. This is a Senior VP-level role … Computer Science, Engineering, or a related technical field (or equivalent applied experience) Preferred qualifications, capabilities, and skills Site Reliability Engineering (SRE) experience and familiarity with reliability practices (SLOs, error budgets) Experience within financial services technology Familiarity with JPM-internal platform, cloud, and AI/ ...

TechOps & Support Engineer, Amazon MGM Studios | Technology Operations & Support

Location
Greater London, England, United Kingdom
software and devices across Amazon MGM Studios global productions. TechOps teams support the Studios production personnel (cast and crew) and Studios business teams (development, engineering, programming and marketing). We work in a team environment and regularly interact with production personnel and studio executives at all levels.Regular activities include … Bachelor's degree in Systems Engineering, Computer Science, or related field or relevant work experience - Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration - Experience working with Linux - Experience in systems engineering - Experience ...

Site Reliability Engineer - Frontend

Hiring Organisation
Capital On Tap
Location
London, UK
Employment Type
Full-time
just getting started! ðLondon, Old Street | ð 2 Days in OfficeSRE at Capital On Tap ðAt Capital On Tap, we run a hybrid embedded SRE model - We aim to work closely with the teams to provide them the best support. As a Site Reliability Engineer (SRE) you will … apply. Interview process ðFirst stage: 30 minute intro and values call with Talent PartnerSecond stage: 60 minute CV overview and technical chat with the SRE team lead and the SRE & Platform Engineering ManagerThird stage: 75 minute technical exercise & questions with the SRE lead Final stage: 30 minute chat with ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
contribute to modern digital capabilities, drive continuous improvement, and support the delivery of future-ready solutions. This is an opportunity to shape high-quality engineering outcomes, embrace innovation and AI-enabled ways of working, and create lasting value in a complex, enterprise-scale environment.Hybrid working:The places that … principles, secure software development practices, and security-focused engineering approaches.Experience working within regulated Financial Services environments.Understanding of Site Reliability Engineering (SRE) concepts, operational resilience, and service reliability practices.Relevant cloud, DevOps, or engineering certifications.We are a Disability Confident Employer:Capgemini is proud ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
employ over 130 colleagues across Jersey, Geneva, London, Singapore, New York and Shanghai. Purpose and Overview of Role This is a hands‐on senior engineering role within the Platform Engineering team, which forms part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure … platforms and developer tooling that enable our engineering and quantitative research teams to deliver software reliably, securely and at scale. The role will contribute to the design, build, automation and operation of a hybrid production platform across AWS and on‐premises environments, with a particular focus on the HashiCorp ...

Devops SRE

Location
Greater London, England, United Kingdom
Cloud Engineering team is seeking a seasoned and passionate Senior Cloud Engineer with deep hands‐on development and cloud engineering expertise. In this role, you will serve as a key technical contributor within a cloud‐focused engineering team, working on one of the Group’s flagship initiatives … best practices and business goals. Required Skills & Experience Core Cloud & DevOps Competencies Extensive experience in DevOps or Site Reliability Engineering (SRE) roles across consumer or SaaS environments. Strong expertise in deploying and managing production‐grade Kubernetes clusters and containerised services. Hands‐on experience with Kubernetes ...

Platform Engineer

Location
City Of London, England, United Kingdom
Platform Engineer Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Segun Ikuesan Description This is a hands‐on engineering role within the Platform Engineering team, which forms part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure, platforms and developer … tooling that enable our engineering and quantitative research teams to deliver software reliably, securely and at scale. The role will contribute to the design, build, automation and operation of a hybrid production platform across AWS and on‐premises environments, with a particular focus on the HashiCorp platform, including Nomad ...

Venue & Studio Deployments System Engineer, Event Productions

Hiring Organisation
Amazon
Location
London, UK
Employment Type
Full-time
Bachelor's degree in Systems Engineering, Computer Science, or related field or relevant work experience- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...