76 to 100 of 1,094 Site Reliability Engineering Jobs in the UK

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability … Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the large language model inference platform ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability … JPMorgan Chase in the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the LLM inference platform end to end. ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
systems run reliably in production at scale. In this role, you’ll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability … Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the large language model inference platform ...

Senior / Lead Site Reliability Engineer

Location
Watford, England, United Kingdom
customer-facing systems during both normal operation and peak lottery events. The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams. Objectives of the role Own reliability outcomes across services using SLOs, SLIs … performance optimisation: Latency reduction Throughput scaling Cost efficiency (AWS utilisation and associated log costs, observability license consumption) Backlog ownership & reporting Own and prioritise the SRE backlog, balancing: Reliability improvements Technical debt Automation opportunities to reduce/offload toil Produce structured reporting covering: SLO performance Incident trends and MTTR Platform ...

Service Manager – Site Reliability Engineering

Location
Belfast City District, Northern Ireland, United Kingdom
principles and IT Service Management processes Experience developing and maintaining operational documentation, runbooks, support procedures, recovery documentation, and knowledge articles Experience working within an SRE, DevOps, Cloud Operations, Platform Engineering, Enterprise Operations, or production support environment Experience supporting Identity and Access Management platforms, including IAM, ISAM, IBM Verify, SailPoint … comparable ITSM platform Knowledge of compliance, risk management, audit controls, operational resilience, business continuity, and disaster recovery practices Relevant professional certifications, such as ITIL, SRE, cloud, Kubernetes, security, ServiceNow, or comparable technology certifications Experience leading operational maturity assessments, service governance programs, or production readiness reviews Core Competencies Demonstrates expertise ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
period of growth and investing heavily in its engineering and platform capabilities. They're looking for an experienced Site Reliability Engineer (SRE) to join the team and play a key role in building highly reliable, scalable, and observable infrastructure. This is a hands-on role focused … experience Help improve platform resilience, scalability, and disaster recovery capabilities Contribute to capacity planning and performance optimisation as the platform scales Establish and champion SRE best practices across the wider engineering function What We're Looking For Proven commercial experience working as an SRE, DevOps Engineer, Platform Engineer ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms team, you hold a leadership role in your team … ability to expand and collaborate across different levels and stakeholder groups Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI-assisted operational recommendations ...

Director of Platform Operations

Location
United Kingdom
govern the SLO and error-budget framework that makes availability measurable, forecastable, and actionable rather than retrospective. Build the reliability discipline. Lead the SRE pillar to embed observability, capacity planning, performance engineering, resilience testing, and toil reduction as standing practices with clear owners and cadence. Make reliability … change failure rate, and time to restore - and use them to target investment where it demonstrably lifts delivery. About you Proven senior leadership of SRE, platform, or infrastructure functions at scale, with accountability for availability against a defined SLA in a customer-facing SaaS environment. Deep, practical command of reliability ...

Site Reliability / Software Engineer - SC Cleared

Hiring Organisation
Searchability NS&D
Location
Gloucestershire, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £65,000 per annum
Site Reliability/Software Engineer (SC Cleared) Location: Gloucestershire SC Clearance required to start with the opportunity to be sponsored through DV Clearance after joining Salary: Up to £65,000 + Clearance Bonus To appy, email: Overview An exciting opportunity has arisen for a technically versatile engineer … Agile teams Desirable Skills Exposure to infrastructure automation tools Experience with microservices architectures Familiarity with MongoDB, Elasticsearch or similar technologies Knowledge of DevOps and SRE best practices Understanding of operational support within secure environments Experience improving system observability and performance Additional Information Active SC Clearance is required for this role ...

Site Reliability Engineer (SRE) - Glasgow, UK

Location
Glasgow, Scotland, United Kingdom
# Site Reliability Engineer (SRE) - Glasgow, UKGlasgowApply for this job* Permanent* Experienced Professionals* Software Engineering* ID 553300-en\\_GB**About the Job you are considering:**We are seeking an experienced **Site Reliability Engineer** SRE AWS DevOps Engineer with strong expertise in AWS Cloud DevOps practices … automation and operational support This role is primarily focused on maintaining supporting and enhancing critical production systems while driving reliability scalability and operational excellence across cloud environments**Hybrid working:**The places that you work from day to day will vary according to your role, your needs, and those ...

Lead Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, United Kingdom
Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a Technical Lead SRE, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. You will collaborate … person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability ...

Senior Lead Site Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part … provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world. Provide technical guidance and serve as a function-wide SRE subject matter expert, driving reliability decisions across multiple products and influencing the adoption of leading-edge observability technologies. #J-18808-Ljbffr ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 05/08/2026 You will have strong software engineering skills, approaching system reliability and observability … policy. Preferred Skills and Experience Knowledge and experience of modern software development techniques and lifecycles. Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
JP Morgan Chase
Location
Glasgow, United Kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Chief Technology Office, you will solve complex and broad business problems, demonstrate … engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilitiesDemonstrates and champions site reliability culture and practices and exerts technical influence throughout your teamLeads initiatives to improve the reliability and stability of your team ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
Site Reliability/Platform Engineer We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments. This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating … that infrastructure more reliable. You'll be joining a growing organisation where you'll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment. Salary: TBC Location/hybrid working: TBC #J-18808-Ljbffr ...

Site Reliability Engineer

Location
Leeds, England, United Kingdom
Site Reliability EngineerAdvertising locationLeedsHours37.5I'm interestedShareJob descriptionAt evoke, Site Reliability Engineering (SRE) is all about delivering exceptional customer experiences through reliable, scalable and high-performing technology. Operating at the heart of our betting and gaming platforms, our SRE team combines observability, automation and engineering … systems can scale effectively to meet changing customer demand.* Drive automation initiatives that improve operational efficiency, reduce manual effort and enhance service reliability.* Promote SRE best practices throughout the organisation, influencing teams through data-driven recommendations and continuous improvement initiatives.Who we are looking forWe are committed to responsible gambling ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Chief Technology Office, youwill solve complex and broad business problems, demonstrate … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team's applications ...

Senior Lead Site Reliability / DevOps Engineer

Location
Auchentibber, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part … Drive significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices ...

Site Reliability Engineer

Location
Newcastle upon Tyne, England, United Kingdom
procedures for incident response and operational tasks. Collaborate with cross-functional teams to review and provide feedback on technical designs, ensuring alignment with SRE principles. Participate in on-call rotations and handle critical incidents with confidence and expertise. Continuously improve documentation for systems and services, contributing to a knowledge-sharing … incident management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized SRE observability experience with New Relic or DataDog. Familiarity with OpenTelemetry, AIOps, MLOps, or SecOps. Location: Newcastle, UK - In-Office (at least 4 days per week ...

Site Reliability Engineer

Hiring Organisation
E-Solutions IT Services UK Ltd
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£280.00 - £300.00 per day
onsite Mode of employment: Inside IR 35 Contract About the Role: We are looking for a passionate and experienced Site Reliability Engineer (SRE) to join our Cloud Platform team. The ideal candidate will have hands-on experience managing large-scale Kubernetes clusters on public cloud environments … strong understanding of modern SRE and DevOps practices. You will be responsible for ensuring high availability, reliability, scalability, and performance of our cloud-native infrastructure and CI/CD systems. Key Responsibilities: • Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Title: Site Reliability Engineer (SRE) III – Platform Engineering & Systems Reliability The Role: CME Group is seeking a Site Reliability Engineer (SRE) III to engineer reliability for our Google Cloud (GCP) infrastructure, Middleware Platform Engineering team, and core technology foundations powering our Clearing … improvement suggestions to the Product backlog. Collaboration & Leadership: Lead technical discussions for assigned scope, present solution options, collaborate across functional teams, and mentor junior SRE colleagues. What We're Looking For Engineering & Scripting Discipline: Programming and scripting skills in high-level languages such as Python, Go, Java, or Bash ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce downtime. What … Experience monitoring high-availability technical systems. Understanding of incident management, root cause analysis, reliability engineering or Site Reliability Engineering (SRE) principles. Experience with HTTP/REST APIs, Git, containerisation, Kubernetes or Infrastructure as Code. Continuous improvement, ownership or leadership experience. Why Join OQC You will ...

Senior Software Engineer

Location
Cheltenham, England, United Kingdom
work-life balance and offer a range of working patterns, including full-time, part-time, and compressed hours. Hybrid working, which combines working on-site and from home, may be more limited due to the nature of the work. However, some homeworking may be available depending on business needs. … existing systems, establish and promote best practices, and deliver high-quality software solutions. Drawing on your expertise in a range of software engineering methodologies, you’ll introduce fresh ideas and innovative approaches that make a real impact at the core of our mission: keeping the UK safe, both ...

Senior Software Engineer Ref. 3839

Location
Manchester, England, United Kingdom
work-life balance and offer a range of working patterns, including full-time, part-time, and compressed hours. Hybrid working, which combines working on-site and from home, may be more limited due to the nature of the work. However, some homeworking may be available depending on business needs. … existing systems, establish and promote best practices, and deliver high-quality software solutions. Drawing on your expertise in a range of software engineering methodologies, you’ll introduce fresh ideas and innovative approaches that make a real impact at the core of our mission: keeping the UK safe, both ...

Site Reliability Engineer III JBLE1 NI

Hiring Organisation
CME Technology Support Services Ltd
Location
Belfast, UK
Title: Site Reliability Engineer (SRE) III - Platform Engineering & Systems Reliability The Role: CME Group is seeking a Site Reliability Engineer (SRE) III to engineer reliability for our Google Cloud (GCP) infrastructure, Middleware Platform Engineering team, and core technology foundations powering our Clearing … improvement suggestions to the Product backlog. Collaboration & Leadership: Lead technical discussions for assigned scope, present solution options, collaborate across functional teams, and mentor junior SRE colleagues. What We're Looking For Engineering & Scripting Discipline: Programming and scripting skills in high-level languages such as Python, Go, Java, or Bash ...