126 to 150 of 1,094 Site Reliability Engineering Jobs in the UK

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
team you'll be working with: We are seeking an experienced Site Reliability Engineer (SRE)/Observability Technical Lead to join our team and drive the strategy and execution of observability and reliability projects across our clients. The ideal candidate will have deep expertise in Application Performance … will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and DevOps. What you'll be doing: Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business ...

Head of Engineering

Hiring Organisation
WRK DIGITAL LTD
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Head of Engineering (Product Platforms) Manchester (Flexible Hybrid. Must be UK based) £140,000 + Bonus + Private Healthcare + Excellent Benefits WRK digital are delighted to be acting as the exclusive recruitment partner to a global FinTech organisation as they continue a significant investment in product engineering … transactions, client onboarding, risk management and business-critical financial operations. Following sustained growth and continued investment in technology, they are seeking a Head of Engineering to lead their product engineering function and help shape the next generation of their platform capability. This is an opportunity to join ...

Software Architect (Java or C#)

Location
United Kingdom
will be doing: Reliability Engineering Partner with Engineering teams to design resilient services, architectures, and deployment patterns. Define and promote SRE practices including SLIs, SLOs, error budgets, capacity planning, incident response, and post-incident learning. Identify systemic reliability risks and work with teams to address root … causes. Help reduce operational toil through automation, tooling, and better engineering practices. Architecture & Engineering Partnership Work actively with Engineering teams during design, development, and production-readiness reviews. Advise and challenge teams on service architecture, fault tolerance, scalability, observability, deployment safety, and operational readiness, helping them to make ...

Manager, Site Reliability Engineering

Hiring Organisation
MasterCard
Location
Harrogate, United Kingdom
networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.Title and SummaryManager, Site Reliability EngineeringAbout MastercardMastercard is a global technology company in the payments industry, dedicated to enabling an inclusive digital economy that benefits everyone, everywhere. … household bills• Almost all state benefit paymentsJoining Vocalink means contributing to systems that millions rely on daily.The OpportunityWe are seeking a Manager, Site Reliability Engineer to support the Production Manager in overseeing a team of individual contributors who support the Card Transaction Services and Faster Payment Services. This ...

Senior DevSecOps Engineer

Location
Greater London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Senior DevSecOps Engineer

Location
City Of London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Senior Network Engineer- IP

Location
Birmingham, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Senior Network Engineer- IP

Location
Ipswich, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Senior Network Engineer- IP

Location
Greater London, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Lead Site Reliability Engineer

Location
United Kingdom
support economic growth across the UK. The Department for Business, Innovation, Science and Trade (BIST), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability … UK. BIST's Digital, Data and Technology (DDaT) directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure ...

Junior SRE – Endpoint Focus

Location
Greater London, England, United Kingdom
technical components. Automation & Continuous Improvement: Proactively isolate recurring operational issues, eliminating manual workflow friction through shell scripting, automated provisioning design, and strategic process enhancements. SRE Transition: Partner closely with Senior SRE team members to progressively absorb production infrastructure tasks, system monitoring duties, and core site reliability principles. Qualifications … performance. Professional Trajectory: A clear, defined motivation to evolve technically and professionally into a Production Engineering or Site Reliability Engineering (SRE) role. Commute Compliance: Willingness and ability to work regularly from our client’s modern office facilities located in the Moorgate area of London (minimum ...

Security Operations Manager

Location
Greater London, England, United Kingdom
outsourced support, and set up how alerts are handled and escalated. You’ll also lead how we respond to incidents, working closely with our engineering and reliability colleagues, since a lot of this sits alongside how they already keep … services running. This is a deeply collaborative role. In particular you’ll work hand in hand with our site reliability engineering (SRE) team, since much of security monitoring and response builds on the same tooling and ways of working they already use to keep our services running. ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting … resume keywords AWS Cloud Platform Architecture Infrastructure as Code CI/CD Automation Platform-as-a-Product Principles Site Reliability Engineering (SRE) Practices ATS Optimization Keywords Hard Skills Cloud Infrastructure Distributed Systems Networking Security Performance Analysis Capacity Planning Scalability Engineering Operational Excellence Monitoring and Alerting Fault ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Staff Software Engineer, AI Reliability Engineering London, UK About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed … from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between. Strong candidates may also Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems Have experience operating large-scale model serving or training infrastructure (>1000 GPUs ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Site Reliability Engineer (DV Security Clearance)

Location
Manchester, England, United Kingdom
seeking an experienced and motivated Site Reliability Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production … environments. The successful candidate will work closely with engineering, platform, and operational support teams to strengthen monitoring and alerting capabilities, improve logging and traceability, troubleshoot incidents, support deployments, and automate operational processes wherever possible. The environment includes Kubernetes, Helm, the ELK stack, and a broad range of modern Site ...

Software Engineer, Model Deployment- ChatGPT Engineering

Location
Greater London, England, United Kingdom
Software Engineer, Model Deployment- ChatGPT Engineering Applied AI Engineering - London, UK About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond … operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking ...

SRE/Infra | Quant Trading

Location
Greater London, England, United Kingdom
Senior SRE/Platform Engineer | Elite Quant Hedge Fund | London Paragon Alpha is partnering with a leading ~$70BN AUM quantitative trading firm following an exceptional 2025 and continuing to invest aggressively across its global technology organisation. As the business scales, they're expanding a world-class Infrastructure/Site … just engineering metrics – they're a competitive advantage. This is a genuinely high-impact engineering role sitting at the intersection of SRE, Platform Engineering, and Software Engineering , where you'll build the tooling and infrastructure that enables researchers and traders to operate at scale. ...

Product Engineering Environment Lead

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Contract
Lead Location: UK (Hybrid with occasional travel) Contract: Interim 6 Months Rate: Inside IR35 Level: Senior Manager/Head of Function Reporting to: Product Engineering Director Essential Experience - Please Read Before Applying We are seeking a technically credible transformation leader who can operate at the intersection of engineering … role is likely to suit candidates from backgrounds such as: DevOps Leadership Platform Engineering Leadership Environment Management Site Reliability Engineering (SRE) Engineering Enablement Technology Operations Software Delivery Transformation Telecommunications Technology Large-scale Digital Engineering Organisations This role is NOT primarily looking for: A hands ...

Site Reliability Manager - Environment Strategy

Location
Greater London, England, United Kingdom
looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving … while ensuring strong technical standards, compliance and resilience across all production and non-production systems. Responsibilities Define and own the vision and roadmap for site reliability and environment strategy Lead, mentor and develop a team of DevOps and environment engineers Set and enforce standards for environment provisioning, lifecycle ...

Site Reliability Engineer

Hiring Organisation
Trainline
Location
London, United Kingdom
millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices.The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. … willingness to challenge and be challenged — contributing to platform reliability while developing broader technical ownership with support from senior engineers.As an SRE at Trainline, you'll be working on...Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platformParticipating in production incident response, supporting investigation, mitigation ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
Cost Explorer, Compute Optimizer, and EKS workload right‐sizing. Skills: 3+ years of hands‐on experience in DevOps, Site Reliability Engineering (SRE), infrastructure engineering, or closely related roles (focused on AWS/EKS environments). Proven track record building and maintaining production‐grade CI/… incidents/post‐mortems into actionable learning for personal and team growth. Proven ability to collaborate effectively with cross‐functional stakeholders, including engineering, SRE, security, and product teams. Maintain clear technical documentation, runbooks, and operational playbooks for production systems. Chinese proficiency is preferred We Offer: Experience a dynamic ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, youhold a leadership role in your team, demonstrate … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team's applications ...

Site Reliability Engineer , Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. We are an operations team, but a key focus … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Software Engineer Lead - Site Reliability

Location
Telford, England, United Kingdom
proactive, self-starting engineer who enjoys getting things done and improving the reliability of live digital services, Standard Life could be the place for you. We’re looking for a Lead DevOps Engineer to join our Digital Engineering team. This role is focused on making immediate, practical improvements … GitHub/GitHub Actions, Azure DevOps, Terraform and automated testing. Improve deployment safety, release readiness and operational readiness for customer-facing digital services. Apply SRE principles pragmatically to improve availability, recoverability, monitoring and incident learning. Strengthen monitoring, logging, tracing, alerting and service-health dashboards across digitally connected workloads. Reduce single ...