126 to 150 of 1,087 Site Reliability Engineering Jobs in the UK

Senior DevOps Engineer - AWS - Manchester

Hiring Organisation
Circle Group
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£70,000
also assist with CloudOps activities. Are you an experienced IT professional with a strong background in DevOps and Site Reliability Engineering (SRE)? Are you passionate about working with cutting-edge technologies, driving agile methodologies, and implementing CI/CD practices? Do you have knowledge of infrastructure … code? Experience required: - Solid experience in a similar role, working on DevOps or SRE initiatives within complex IT environments with Software Engineering - AWS environment - Proficiency in DevOps practices and related technologies, such as CI/CD pipelines & infrastructure as code tools such as Terraform, Ansible, Puppet or Bicep. - Strong ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
domains, and advise others on the technical and business issues facing them. You will will set the vision, strategy, and operating model for our SRE transformation - enabling our business-aligned support teams to deliver higher reliability, stronger resilience, and a measurably better end-user experience across the board. … responsibilities Defines the SRE vision, north-star outcomes, and multi-year roadmap for the Production Management team, aligned to both CIB and JPM Global Technology priorities. Establishes the SRE operating model across global regions (ways of working, intake, prioritization, engagement with engineering teams and production support). Partners with ...

Software Engineer, SRE

Location
United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 30/09/2026 You will have strong software engineering skills, approaching system reliability and observability … policy. Preferred Skills and Experience Knowledge and experience of modern software development techniques and lifecycles. Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 30/09/2026 You will have strong software engineering skills, approaching system reliability and observability … policy. Preferred Skills and Experience Knowledge and experience of modern software development techniques and lifecycles. Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. ...

Site Reliability Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
respect and fairness. If you're ready to push boundaries and challenge the status quo in security, we want to hear from you. Site Reliability Engineer Summary: Our growing technology company is seeking an experienced Site Reliability Engineer with deep Azure expertise to help maintain … availability, performance, and reliability of our critical SaaS applications. In this role, you will own and drive automation, monitoring, incident response, and infrastructure improvements across our multi-cloud, multi-region environment, working closely with senior engineering and cross-functional teams. What You'll Do: Own the availability ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Location
Greater London, England, United Kingdom
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing … production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
London, UK
Employment Type
Full-time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI.Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities ...

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

Head of Engineering

Hiring Organisation
WRK DIGITAL LTD
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Head of Engineering (Product Platforms) Manchester (Flexible Hybrid. Must be UK based) £140,000 + Bonus + Private Healthcare + Excellent Benefits WRK digital are delighted to be acting as the exclusive recruitment partner to a global FinTech organisation as they continue a significant investment in product engineering … transactions, client onboarding, risk management and business-critical financial operations. Following sustained growth and continued investment in technology, they are seeking a Head of Engineering to lead their product engineering function and help shape the next generation of their platform capability. This is an opportunity to join ...

Software Engineer/ SRE (Linux)

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
work that matters – to you, to your community, and to the world. Progress starts with you. Job DescriptionSite Reliability Engineering (SRE) is essential to Visa's Cloud platform strategy. In this role, you'll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. … full coverage. Hands-on expertise is required, especially with major DevTools like GitHub, Jenkins, Jira, and Artifactory. We seek a Software Engineer + SRE hybrid engineer. The ideal candidate deeply understands at least one major DevTool, quickly resolves tool-related issues in collaboration with developers, and applies systems thinking ...

Software Architect (Java or C#)

Location
United Kingdom
will be doing: Reliability Engineering Partner with Engineering teams to design resilient services, architectures, and deployment patterns. Define and promote SRE practices including SLIs, SLOs, error budgets, capacity planning, incident response, and post-incident learning. Identify systemic reliability risks and work with teams to address root … causes. Help reduce operational toil through automation, tooling, and better engineering practices. Architecture & Engineering Partnership Work actively with Engineering teams during design, development, and production-readiness reviews. Advise and challenge teams on service architecture, fault tolerance, scalability, observability, deployment safety, and operational readiness, helping them to make ...

Senior DevSecOps Engineer

Location
Greater London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Senior DevSecOps Engineer

Location
City Of London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Lead Site Reliability Engineer

Location
United Kingdom
support economic growth across the UK. The Department for Business, Innovation, Science and Trade (BIST), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability … UK. BIST's Digital, Data and Technology (DDaT) directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure ...

Junior SRE – Endpoint Focus

Location
Greater London, England, United Kingdom
technical components. Automation & Continuous Improvement: Proactively isolate recurring operational issues, eliminating manual workflow friction through shell scripting, automated provisioning design, and strategic process enhancements. SRE Transition: Partner closely with Senior SRE team members to progressively absorb production infrastructure tasks, system monitoring duties, and core site reliability principles. Qualifications … performance. Professional Trajectory: A clear, defined motivation to evolve technically and professionally into a Production Engineering or Site Reliability Engineering (SRE) role. Commute Compliance: Willingness and ability to work regularly from our client’s modern office facilities located in the Moorgate area of London (minimum ...

Security Operations Manager

Location
Greater London, England, United Kingdom
outsourced support, and set up how alerts are handled and escalated. You’ll also lead how we respond to incidents, working closely with our engineering and reliability colleagues, since a lot of this sits alongside how they already keep … services running. This is a deeply collaborative role. In particular you’ll work hand in hand with our site reliability engineering (SRE) team, since much of security monitoring and response builds on the same tooling and ways of working they already use to keep our services running. ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting … resume keywords AWS Cloud Platform Architecture Infrastructure as Code CI/CD Automation Platform-as-a-Product Principles Site Reliability Engineering (SRE) Practices ATS Optimization Keywords Hard Skills Cloud Infrastructure Distributed Systems Networking Security Performance Analysis Capacity Planning Scalability Engineering Operational Excellence Monitoring and Alerting Fault ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Site Reliability Engineer (DV Security Clearance)

Location
Manchester, England, United Kingdom
seeking an experienced and motivated Site Reliability Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production … environments. The successful candidate will work closely with engineering, platform, and operational support teams to strengthen monitoring and alerting capabilities, improve logging and traceability, troubleshoot incidents, support deployments, and automate operational processes wherever possible. The environment includes Kubernetes, Helm, the ELK stack, and a broad range of modern Site ...

Software Engineer, Model Deployment- ChatGPT Engineering

Location
Greater London, England, United Kingdom
Software Engineer, Model Deployment- ChatGPT Engineering Applied AI Engineering - London, UK About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond … operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking ...

SRE/Infra | Quant Trading

Location
Greater London, England, United Kingdom
Senior SRE/Platform Engineer | Elite Quant Hedge Fund | London Paragon Alpha is partnering with a leading ~$70BN AUM quantitative trading firm following an exceptional 2025 and continuing to invest aggressively across its global technology organisation. As the business scales, they're expanding a world-class Infrastructure/Site … just engineering metrics – they're a competitive advantage. This is a genuinely high-impact engineering role sitting at the intersection of SRE, Platform Engineering, and Software Engineering , where you'll build the tooling and infrastructure that enables researchers and traders to operate at scale. ...

Product Engineering Environment Lead

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Contract
Lead Location: UK (Hybrid with occasional travel) Contract: Interim 6 Months Rate: Inside IR35 Level: Senior Manager/Head of Function Reporting to: Product Engineering Director Essential Experience - Please Read Before Applying We are seeking a technically credible transformation leader who can operate at the intersection of engineering … role is likely to suit candidates from backgrounds such as: DevOps Leadership Platform Engineering Leadership Environment Management Site Reliability Engineering (SRE) Engineering Enablement Technology Operations Software Delivery Transformation Telecommunications Technology Large-scale Digital Engineering Organisations This role is NOT primarily looking for: A hands ...

Site Reliability Manager - Environment Strategy

Location
Greater London, England, United Kingdom
looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving … while ensuring strong technical standards, compliance and resilience across all production and non-production systems. Responsibilities Define and own the vision and roadmap for site reliability and environment strategy Lead, mentor and develop a team of DevOps and environment engineers Set and enforce standards for environment provisioning, lifecycle ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
Cost Explorer, Compute Optimizer, and EKS workload right‐sizing. Skills: 3+ years of hands‐on experience in DevOps, Site Reliability Engineering (SRE), infrastructure engineering, or closely related roles (focused on AWS/EKS environments). Proven track record building and maintaining production‐grade CI/… incidents/post‐mortems into actionable learning for personal and team growth. Proven ability to collaborate effectively with cross‐functional stakeholders, including engineering, SRE, security, and product teams. Maintain clear technical documentation, runbooks, and operational playbooks for production systems. We Offer: Experience a dynamic and team‐orientated work environment ...

Site Reliability Engineer

Location
Leeds, England, United Kingdom
Description At evoke, Site Reliability Engineering (SRE) is all about delivering exceptional customer experiences through reliable, scalable and high-performing technology. Operating at the heart of our betting and gaming platforms, our SRE team combines observability, automation and engineering excellence to ensure our systems perform when … systems can scale effectively to meet changing customer demand. Drive automation initiatives that improve operational efficiency, reduce manual effort and enhance service reliability. Promote SRE best practices throughout the organisation, influencing teams through data-driven recommendations and continuous improvement initiatives. Who we are looking for We are committed to responsible ...