226 to 250 of 278 Site Reliability Engineer Jobs

Platform Site Reliability Engineer

Location
Gloucester, England, United Kingdom
decisions clearly to non-technical stakeholders and customers Uphold a culture of: do, document, automate Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks. Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering … Requirements 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model in an SRE or equivalent role 3+ years experience in both running, deploying and optimising orchestration platforms with a strong emphasis on Kubernetes Expert-level Linux administration, especially Ubuntu distributions Proficiency ...

Site Reliability Engineer

Location
Cardiff, Wales, United Kingdom
Working pattern: Hybrid (1 day a week in Cardiff office) About the Role This company provides managed AI operations for technology businesses. The company operates, secures and governs the cloud, observability and AI runtime layer ...

SRE Engineer

Location
Greater London, England, United Kingdom
client is looking for a SRE Engineer combining software and IT engineering principles to build and maintain reliable, scalable, and high-performing systems. Job Responsibilities: Scope technical projects and break them down into user stories and tasks within an engineering team Directly contribute to the design and coding ...

Site Reliability Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
architecting everything around AI-native workflows. The Cloud Platform team is right at the center of that. We're looking for an SRE who doesn't just want to run reliable systems — but wants to use AI to make them more reliable. If you're already experimenting with … incident response, anomaly detection, or toil reduction, we want to talk. This isn't a "keep the lights on" SRE job. We're building the infrastructure that powers AI agents handling sensitive workflows for major financial and legal firms — and we want SREs who use AI as a force multiplier ...

Azure SRE Engineer - Systems Integrator

Location
York and North Yorkshire, England, United Kingdom
services, backed by UK data centres, a national 100Gb MPLS network, and 24/6 network operations centres. This opportunity is ideal for an SRE Engineer joining the Cloud Services Group to support the reliability, operation, security, and evolution of customer Azure platforms. The role involves designing … production environments – primarily the compute, networking, storage, database, costing, security and IAM, and management tools service groups.? Solid understanding of DevOps/SRE, continuous delivery and related principles with demonstrable experience using complex CI/CD implementations and IAC tools. e.g., Terraform, Bicep and ARM. Salary: £60,000 Basic Salary ...

Lead Site Reliability Engineer - Azure & CI/CD

Location
Telford, England, United Kingdom
Standard Life plc is seeking a Lead DevOps Engineer to strengthen the Digital Engineering team. The role focuses on improving reliability of live digital services, Azure API platform integration, cloud infra, CI/CD, observability and deployment safety. Hybrid roles span Birmingham, London, Telford, and Edinburgh with flexible ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments and full reproducibility, including ...

Staff Site Reliability Engineer

Location
Greater London, England, United Kingdom
that usable, and MystraAI is the agentic layer we are building on top of it. This is a Staff-level role that owns the reliability, performance, security and integrity of that infrastructure end-to-end — and sets the technical direction that other teams build on. You will lead … source level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data ...

Lead DevOps Engineer: Site Reliability (Hybrid)

Location
Telford, England, United Kingdom
Standard Life is seeking a hands-on Lead DevOps Engineer to strengthen the Customer Digital Platform. You’ll drive practical improvements, Azure API platform integration, cloud infra …/CD automation and observability, working with internal squads, partners and third parties. The role emphasizes deployment safety, release readiness and operational readiness, applying SRE principles to improve availability and incident learning in a hybrid work setup. #J-18808-Ljbffr ...

Waste Site SCADA Engineer – Reliability & OT Upgrades

Location
Phillack, England, United Kingdom
Thames Water is seeking a Waste Site SCADA Systems Engineer to maintain and enhance site SCADA across our extensive wastewater operations. You will support AVEVA System Platform/Wonderware, FactoryTalk View and Iconics, ensuring reliability and security across hundreds of sites. The role involves diagnosing faults ...

Platform Development Engineer: SRE, Kubernetes & Cloud

Location
Crewe, England, United Kingdom
Futura Design Ltd is seeking a Platform Development Engineer based in Crewe. This contract position is Inside IR35 with a proposed hourly pay rate of £33.64. The ideal candidate will possess 4+ years' experience working with Unix OS and container orchestration tools like Docker and Kubernetes, alongside strong Site Reliability Engineering experience. Applicants should have experience in operational monitoring, source code management, and CI/CD processes. Strong communication and collaboration skills are essential for working in cross-functional teams. #J-18808-Ljbffr ...

SRE Engineer: Build Reliable, Scalable Systems

Location
Greater London, England, United Kingdom
leading recruitment agency is seeking an SRE Engineer to combine software and IT engineering principles for building reliable systems. Responsibilities include scoping projects, designing software, and automating infrastructure management. Ideal candidates will have proficiency in programming languages like Python or Go, experience with Terraform, and familiarity with CI/ ...

Site Reliability Engineer Infrastructure View role →

Location
Greater London, England, United Kingdom
highly concurrent, agent-driven system. Drive cost and performance optimization across compute-heavy scanning workloads. What we're looking for 4+ years in SRE/infrastructure roles operating production Kubernetes at scale. Strong grasp of network and workload isolation (namespaces, gVisor/Firecracker-style sandboxing, or similar). Experience with ...

Software Engineering or SRE, PhD Intern, 2027

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
hackajob is partnering directly with Google to hire for this role. We offer a range of internships in either Software Engineering or Site-Reliability Engineering across EMEA. Durations and start dates will vary according to project and location. Our recruitment team will determine where you fit best based … continue to push technology forward. You will design, test, deploy and maintain software solutions as you grow and evolve during your internship. Site Reliability Intern: Our engineers create, fix, extend and scale the code to keep it working and to harden it against all the bad actors ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
hackajob is partnering directly with JPMorganChase to hire for this role. JOB DESCRIPTION There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skills to ...

Site Reliability Engineer Manchester

Hiring Organisation
Hackajob Ltd
Location
Gloucester, Gloucestershire, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
hackajob is partnering directly with BAE Systems Digital Intelligence to hire for this role. Location(s):UK, Europe & Africa : UK : Gloucester BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts. ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Swindon, Wiltshire, United Kingdom
Employment Type
Permanent
Salary
GBP 85,000 Annual
hackajob is partnering directly with Edenred PayTech to hire for this role. Take a step forward and let Edenred surprise you. Every day, we deliver innovative solutions to improve the life of millions of people ...

SRE Platform Engineer: Python Automation & Incident Response

Location
West of England, England, United Kingdom
Incite-Insight.co.uk is seeking an experienced Site Reliability/Platform Engineer to automate and modernise its infrastructure for large-scale environments. You'll build Python-based automation around incident management, runbooks, and routine tasks, and integrate monitoring … ITSM platforms via APIs, to reduce toil and speed up incident response. This role offers significant autonomy and an opportunity to shape the SRE capabilities within the organisation. #J-18808-Ljbffr ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
years, or PhD + 3 years in Computer Science, Software Engineering, or a related technical field. Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale. Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs … production automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands‐on experience building LLM pipelines, AI Agents ...

Observability SRE — Reliability & Telemetry Engineer, London

Location
Greater London, England, United Kingdom
HCLTech in London is seeking an Observability SRE to join the Group Platform Services & Engineering division. The role focuses on administering the production environment, building scalable monitoring, and embedding reliability in products and services. You will work with a global, agile team to enhance telemetry, observability and incident response ...

Senior Platform Engineer — SRE & Cloud

Location
Greater London, England, United Kingdom
Beamery in London seeks a hands-on Principal Platform Engineer to lead reliability and scalability across our cloud-based platform. You will shape architecture, drive incident response, and mentor engineers while partnering with product and leadership to deliver scalable services. You will own platform roadmap, advance Kubernetes, Terraform ...

Senior Platform Engineer – Remote UK (Cloud/SRE)

Location
Greater London, England, United Kingdom
Hudl in London, United Kingdom, is seeking a Senior Engineer to join our Platform Engineering team. You’ll work on site reliability, cloud infrastructure, observability and production operations to keep Hudl’s platform highly available, scalable and secure. You’ll lead with technical excellence, mentor engineers ...

Senior SRE & DevTools Engineer (CI/CD & Observability)

Location
Ham, England, United Kingdom
Visa is seeking a Software Engineer + SRE hybrid to join its UK Cloud platform team. You will safeguard reliability, automate resolution of recurring issues, and work with developers to optimize CI/CD pipelines. The role combines hands-on SRE with software engineering, requiring experience with GitHub ...

Storage Platform Engineer (SRE)

Location
Greater London, England, United Kingdom
Selection changes the language of the page/content London, England, United Kingdom Software and Services Apple Cloud infrastructure is BIG. The storage SRE teams of Apple Cloud are building and runningthe next generation distributed storage systems to support Apple’s most critical services. Operating atour scale, across multiple geographically … dispersed data centres, and servicing users with vast dataneed presents unique challenges. As a member of Storage SRE at Apple, you'll need to solve theseproblems using your deep understanding of infrastructure, storage, data analysis, programming,teamwork, and expertise in Linux system internals. Description We are looking for seasoned software ...