5 of 5 Site Reliability Engineer Jobs in Gloucester

HPC Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
experience operating large‐scale distributed systems and recent hands‐on expertise in high‐performance computing (HPC) and AI infrastructure. This is an operations‐first SRE role, working in a 24/7/365 on‐call environment, responsible for ensuring reliability, performance, and continuous improvement of mission‐critical infrastructure. … This role sits within a cross‐functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations. The ideal candidate has progressed through large‐scale, globally distributed or multi‐site infrastructure environments and has more recently specialised in GPU‐accelerated HPC systems. This ...

Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
leaders who can build strong teams, uphold high standards, and deliver reliably at pace. Job Summary We’re looking for an experienced Infrastructure Site Reliability Engineer to run and evolve our infrastructure stack. You’ll contribute across bare-metal, virtualization, and orchestration layers, keeping things stable … decisions clearly to non-technical stakeholders and customers Uphold a culture of: do, document, automate Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks. Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering ...

Site Reliability Engineer Manchester

Hiring Organisation
Hackajob Ltd
Location
Gloucester, Gloucestershire, South West, United Kingdom
Employment Type
Permanent, Work From Home
community engagement and outreach activities to help build tech and cyber skills in the region. What you could be doing for us: As an SRE, fundamentally you will be doing work that has historically been done by an operations team, but using software and systems engineering expertise to use automation … reduce human labour, limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time. Core role accountabilities include: Supporting and maintaining essential service that support core mission applications, proactively enhancing their availability, performance and stability; Finding innovative solutions to problems ...

Platform Site Reliability Engineer

Location
Gloucester, England, United Kingdom
decisions clearly to non-technical stakeholders and customers Uphold a culture of: do, document, automate Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks. Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering … Requirements 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model in an SRE or equivalent role 3+ years experience in both running, deploying and optimising orchestration platforms with a strong emphasis on Kubernetes Expert-level Linux administration, especially Ubuntu distributions Proficiency ...

Site Reliability Engineer Manchester

Hiring Organisation
Hackajob Ltd
Location
Gloucester, Gloucestershire, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
Location(s):UK, Europe & Africa : UK : Gloucester BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts. We work collaboratively across 10 countries to collect, connect and understand complex data, so ...