76 to 100 of 426 Remote Site Reliability Engineering Jobs

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform services supporting critical digital products. * Develop and maintain platform tooling … automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. * Support live ...

Senior Backend Developer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform services supporting critical digital products. * Develop and maintain platform tooling … automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. * Support live ...

Senior Back End Engineer

Location
Cardiff, Wales, United Kingdom
platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: Design, build and operate reliable, secure and scalable cloud platform services supporting critical digital products. Develop and maintain platform tooling … automation, observability, monitoring and CI/CD capabilities. Build software solutions using Python and modern engineering practices. Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets and continual service improvement. Support live ...

Site Reliability Engineer / Production Support

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
fastest growing fintech in 2025. The momentum is real. THE OPPORTUNITY Monuments production environment is the heartbeat of a licensed bank, and the SRE role is the single point of ownership when incidents occur. You will directly oversee the offshore Production Support team, run on-call and incident response … eliminate it. Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source. WHAT YOU BRING Strong SRE or production support experience with accountability for incident response in a production environment. Deep understanding of observability tools, alerting, logging, and distributed systems debugging. Experience ...

Staff Technical Program Manager, Site Reliability Engineering

Hiring Organisation
MongoDB
Location
Lincoln, Lincolnshire, United Kingdom
Salary
£ 70 K
SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps … stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based remotely on the East CoastWhat You'll DoDrive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams ...

Staff Software Engineer - Databases SRE | UK | Remote

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
opportunity and we are looking for candidates from the UK, Sweden, Spain or Germany.About the role:We are looking for a Staff Software Engineer - SRE to help us support our highest value Grafana Cloud customers by increasing the reliability of our Cloud databases that are based on Mimir, Loki … Tempo, and Pyroscope. We provide these databases as a SaaS product from AWS, GCP, and Azure across all regions.The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
HIPAA compliance — encryption, access controls, audit logging, and network security Partner with engineering teams to optimize system performance, cost, and scalability Grow the SRE practice: runbooks, incident response playbooks, chaos engineering, and reliability reviews Location Remote (US). If you're in the Birmingham, AL area, this … infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site Reliability Engineering background or formal SRE experience Experience supporting real-time or high-throughput systems (voice, streaming, or similar) AWS certifications (Solutions Architect, DevOps Engineer, or SysOps) Experience with multi-region ...

Azure DevOps Engineer

Hiring Organisation
Anson Mccade
Location
South East London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £500 to £700 per day Inside IR35
programme work with strong extension potential High-impact role with significant technical ownership Enterprise-scale Azure cloud environment Opportunity to shape DevSecOps and platform engineering best practices Exposure to cloud security and infrastructure automation initiatives Work alongside experienced cloud, security and engineering teams Flexible hybrid working arrangements Opportunity … Azure DevOps Engineer Microsoft Certifications such as AZ-104, AZ-305 or AZ-400 Platform Engineering practices Site Reliability Engineering (SRE) Agile and Scrum methodologies Enterprise change and release management Mission-critical cloud platforms Disaster recovery and resilience planning Apply Today If you're an experienced ...

Senior Platform Engineer

Location
Warminster, England, United Kingdom
with integration engineers, customers and modelling engineers. Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This … pressure during outages or failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. ...

Site Reliability Engineer — Hybrid, Automation-Driven

Location
Swindon, England, United Kingdom
Edenred is hiring a Site Reliability Engineer (SRE) in Swindon on a hybrid basis. You will join the Infrastructure Engineering team to ensure reliability, scalability, and alignment with business priorities, driving automation and efficiency across our infrastructure. The role emphasizes engineering reliability, strategic planning ...

Principal Agentic Architect

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
systems to achieve their business objectives. At Lam Research, we are building the next generation of intelligent IT operations through Agentic AI, platform engineering, and autonomous service management. Aufgaben The impact you'll make As the Principal Agentic Architect, you will define the architecture that enables AI agents … Experience with agent frameworks such as Semantic Kernel, Microsoft Agent Framework, LangGraph, AutoGen, or equivalent technologies. Experience with Site Reliability Engineering (SRE), AIOps, Platform Engineering, or Internal Developer Platforms. Knowledge of AI governance, responsible AI, and enterprise security practices. Industry certifications in cloud architecture, AI, enterprise ...

Site Reliability Engineer

Location
United Kingdom
role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight. Working across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across … hybrid work from home policy. Preferred Skills and Experience Knowledge of modern development practices, including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Location
United Kingdom
strong Senior and Lead Cloud Platform Engineers for future opportunities. Role Overview We are seeking an experienced Cloud Platform Engineer to join our platform engineering team. This is a hands‐on role focused on designing, building, operating, and continuously improving cloud‐native platforms that enable development teams to deliver … using Java, Kotlin, Python, or similar technologies Experience implementing cloud security controls, governance, and compliance requirements Exposure to site reliability engineering (SRE) practices and platform engineering frameworks What we’ll offer you: We trust people to do their best work. That means flexibility over rigid rules ...

AI Technical Platform Leader

Hiring Organisation
Willis Towers Watson
Location
London, United Kingdom
Salary
£ 80 K
Compliance, and business stakeholders to mold and implement strategy.The role applies broad technical knowledge with depth in generative AI, agentic systems, enterprise platforms and engineering governance to ensure that AI platforms at WTW are optimally configured to achieve company vision and imperatives. It has end-to-end ownership … operations in a large, global enterprise: • Software or platform engineering for global, enterprise-scaled solutions• Management of Site reliability engineering (SRE) programs for mission-critical systems• Creation and management of DevOps practices for automated, consistent, and secure solution deployment in regulated environments.• Literacy in global compliance ...

Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
Description We are seeking an experienced Site Reliability Engineer (SRE) to join our Group Technology Team in Milton Keynes. ConnellsX is the company Technology's internal developer platform, built on Microsoft Azure. It simplifies cloud hosting, embeds security and compliance by default, and enables a frictionless developer experience. … operating this platform, you will play a hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post‐incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
contribute to modern digital capabilities, drive continuous improvement, and support the delivery of future-ready solutions. This is an opportunity to shape high-quality engineering outcomes, embrace innovation and AI-enabled ways of working, and create lasting value in a complex, enterprise-scale environment.Hybrid working:The places that … principles, secure software development practices, and security-focused engineering approaches.Experience working within regulated Financial Services environments.Understanding of Site Reliability Engineering (SRE) concepts, operational resilience, and service reliability practices.Relevant cloud, DevOps, or engineering certifications.We are a Disability Confident Employer:Capgemini is proud ...

Head of Cyber, Platforms & IT

Location
Chester, England, United Kingdom
enforce secure‐by‐design principles, including cybersecurity standards, cloud architecture guardrails and operational controls Lead DevOps and Site Reliability Engineering (SRE) maturity, embedding monitoring, observability, automated testing and structured incident response Drive adoption of automation and AI‐enabled tooling to improve anomaly detection, incident management, vulnerability management … , cybersecurity or reliability teams in complex SaaS environments Deep expertise in cloud‐native architecture (Azure, AWS or equivalent) Strong understanding of DevOps, SRE principles and SaaS production operations Experience implementing secure‐by‐design frameworks and managing cybersecurity governance Experience owning uptime, reliability and systemic operational performance Strong ...

Site Reliability Engineer - NS London

Hiring Organisation
BAE SYSTEMS
Location
London, United Kingdom
Salary
£ 70 K
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission.The role holder:As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

AWS Cloud Architect, Technology Consulting

Location
Greater London, England, United Kingdom
security and automation technologies. You will help design scalable, secure and resilient cloud solutions for our clients, enabling digital transformation through modern architecture, platform engineering and DevOps practices. You should bring practical experience of cloud technologies and infrastructure modernisation, alongside an understanding of how AI-enabled tooling and modern … capabilities Strong understanding of cloud networking, identity, security and resilience patterns Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) teams Cloud architecture certifications in additional public cloud platforms (Azure or GCP) Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

Site Reliability Engineer

Hiring Organisation
Twinstream Limited
Location
Cheltenham, Gloucestershire, United Kingdom
Employment Type
Permanent
Salary
GBP 95,000 Annual
Site Reliability Engineer Up to £95,000 DOE Fully Remote Initially Cheltenham Future Hybrid Working Build resilient systems. Solve complex challenges. Make a real impact. Are you a Site Reliability Engineer or DevOps professional who enjoys getting stuck into complex infrastructure challenges? Do you want … work with talented engineers on meaningful projects where reliability, security and t click apply for full job details ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Site Reliability Engineer Reports to: Labs Team Lead Looper Insights Remote-first, with regular visits to our Byfleet and Hounslow data centres The company Looper Insights builds analytics products that help the world’s leading media and entertainment companies understand how their content is performing across digital platforms … collaborate closely to produce a valuable service for an industry about which we are all passionate. The role We’re looking for a Site Reliability Engineer to keep the global LooperBox fleet running, the physical backbone behind every piece of data Looper Insights produces. LooperBoxes sit in front ...

Remote Senior Site Reliability Engineer Manager (Remote)

Location
Cambourne, England, United Kingdom
presence in London, Hong Kong, Amsterdam, and as well in Mumbai and now in New York in 2001. About the role : As the SRE Manager, you will play a critical role in ensuring the reliability, scalability, and performance of our infrastructure and services through both direct technical contribution along … streamline operational workflows and improve efficiency. Develop and maintain tools, scripts, and dashboards to monitor system health, performance, and reliability. Build a first class SRE team. Through a combination of leading by example, coaching and mentoring, mould the team would want to have around you. Provide leadership and guidance ...

Site Reliability Engineer with Python

Hiring Organisation
Nexus Jobs
Location
London, United Kingdom
Salary
£ 80 K
000Sector: I.T. & CommunicationsJob Type: PermanentWork Hours: Full TimeContact: Jas GujralEmail: cv@nexusjobs.comTelephone: 020 7488 6900Apply for this job nowJob DescriptionSite Reliability Engineer with PythonOur Client looking to bring on a site reliability engineer to help deploy, manage, troubleshoot, and enhance our complex cloud-based set of internal … variety of users across our wide-ranging organization.You will have at least 7 to 10 years hands-on expertise working as a Site Reliability Engineer.You will work closely with IT, product, and engineering to extend and maintain this set of tools and services and to help debug ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
toolchain and systems, code build and deployment, incident response, and 24x7 monitoring and support. The candidate will also have extensive experience operating within a SRE function within a complex, distributed environment. They will have a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor … learning through blameless post-mortems to improve the shared goal of reliability across services Transform operations teams by facilitating internal change to adopt SRE standard methodologies across the organization and driving strategic growth in this area within Global Technology Analyzes incidents impacting technology availability for high-level trends across ...

Senior Consultant, Platform Engineer, Solutions, Engineering, AI & Data

Hiring Organisation
Appcast
Location
Glasgow, UK
# 23178 Job description Connect to your IndustryDo you want to be at the heart of some of the biggest and most ambitious engineering projects Technology & Transformation: How do you make intelligent, future-proof decisions in a world where change is the one constant? That's exactly what … premise).Exposure to service mesh technologies (e.g., Istio, Linkerd) and configuration management tools (e.g., Chef, Puppet).Knowledge of site reliability engineering (SRE) principles and practices.GCP certifications (e.g., Cloud Digital Leader, Associate Cloud Engineer, DevOps Engineer).ITILv4 Foundation certification.Familiarity with GCP AI/ML services (AI Platform, AutoML ...