101 to 125 of 275 Permanent Site Reliability Engineer Jobs

Senior Site Reliability Engineer – Cloud & Platform Automation

Location
Belfast City District, Northern Ireland, United Kingdom
Group is seeking a Site Reliability Engineer (SRE) III to strengthen our Google Cloud platform and middleware stack. You will help design resilient, low-latency systems across CME’s core derivatives applications and mentor junior engineers. In addition to hands-on engineering, you will drive cloud transformation … participate in disaster recovery planning, and lead reliability improvements across cross-functional teams in a hybrid UK-based role. #J-18808-Ljbffr ...

Site Reliability Engineer

Hiring Organisation
RSA TECH GROUP
Location
Liberty, Texas, United States
Employment Type
Any
Salary
USD Annual
Site Reliability Engineer Dallas TX (Day 1 onsite) Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc Good to have skills: Cloud Concepts/AWS, Terraform, MongoDB, Spring Boot Framework, Kafka, RESTful API, Camunda Should ...

Site Reliability Engineer (London) - Banking & Finance

Location
Greater London, England, United Kingdom
collaboration and technical excellence, the organisation continues to push the boundaries of low-latency infrastructure and reliable system design. The team is hiring a Site Reliability Engineer (London) to build, monitor, and optimise mission-critical trading systems. The role will focus on automation, system scalability, and incident … improving the infrastructure. Drive automation and operational excellence by leveraging your Linux expertise, Kubernetes, and Python scripting skills. Monitor and ensure high availability and reliability of trading applications while being on top of system alerts and incidents. Key Requirements: 1-5 years working experience The right candidate will come ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
hackajob is partnering directly with BNY Mellon to hire for this role. Were seeking a future team member for the role of Vice President - Site Reliability Engineer to join our team. This role is located in London. Role Summary BNY is seeking a Vice President - Site Reliability Engineer to design, build, deploy, and scale resilient, automated, and centrally managed engineering solutions for Production click apply for full job details ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more. As a Site Reliability Engineer (SRE), you will work at the intersection of production operations and software development as you improve, manage, and monitor production-critical infrastructure and data pipelines. ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City, London, United Kingdom
Employment Type
Permanent
Salary
GBP 85,000 Annual
Site Reliability Engineer Up to £85,000 + Benefits Central London Hybrid (2/3 days a week in the office) Build, Scale & Improve the Reliability of a Fast-Growing SaaS Platform We're partnering with a fast-growing SaaS company that's going through ...

Senior / Lead Site Reliability Engineer

Location
Watford, England, United Kingdom
customer-facing systems during both normal operation and peak lottery events. The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams. Objectives of the role Own reliability outcomes across services using SLOs, SLIs, and error … performance optimisation: Latency reduction Throughput scaling Cost efficiency (AWS utilisation and associated log costs, observability license consumption) Backlog ownership & reporting Own and prioritise the SRE backlog, balancing: Reliability improvements Technical debt Automation opportunities to reduce/offload toil Produce structured reporting covering: SLO performance Incident trends and MTTR Platform ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‐healing workflows, and resilience engineering. Use enterprise‐authorized AI capabilities within the work environment to accelerate major‐incident triage ...

Site Reliability Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 Annual
Site Reliability Engineer Location: Manchester or Gloucester (Hybrid) Salary: £40,000£65,000 Sector: Consulting National Security Who you'll be working with You'll be joining a growing National Security consulting team, working with a range of clients to deliver technology solutions across some ...

Site Reliability Engineer

Hiring Organisation
Anson Mccade
Location
Gloucester, Gloucestershire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 Annual
Site Reliability Engineer Location: Manchester or Gloucester (Hybrid) Salary: £40,000£65,000 Sector: Consulting National Security Who you'll be working with You'll be joining a growing National Security consulting team, working with a range of clients to deliver technology solutions across some ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Lead Site Reliability Engineer

Location
Westminster, West End, United Kingdom
trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
Southampton, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£60000 - £70000/annum
Senior Site Reliability Engineer Southampton HQ - 2 Times a week in Office Cloud, SaaS, AWS, The company deliver cutting-edge enterprise software solutions across both cloud and on-premises environments, empowering organisations to enhance customer experiences, maintain regulatory compliance, and proactively fight fraud. The company are trusted … Datadog, PagerDuty, or Rundeck Experience using configuration management platforms like Ansible, Puppet, or Chef Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar credentials Do You Have What It Takes? 3-6 years of hands-on experience ...

eDV-Cleared Site Reliability Engineer – Platform Reliability

Location
Cheltenham, England, United Kingdom
Forward Role Recruitment is seeking a Site Reliability Engineer (eDV) for national security environments. You will own production reliability, instrument services, and drive improvements in deployment, monitoring and performance. You'll work across cloud, platform and DevOps domains with AWS, Kubernetes, Terraform, Linux, CI/… scripting in Python or Bash. Active eDV clearance is required; role is on-site at national security hubs. #J-18808-Ljbffr ...

HPC Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
experience operating large‐scale distributed systems and recent hands‐on expertise in high‐performance computing (HPC) and AI infrastructure. This is an operations‐first SRE role, working in a 24/7/365 on‐call environment, responsible for ensuring reliability, performance, and continuous improvement of mission‐critical infrastructure. … This role sits within a cross‐functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations. The ideal candidate has progressed through large‐scale, globally distributed or multi‐site infrastructure environments and has more recently specialised in GPU‐accelerated HPC systems. This ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
/ML Technologies and Professional services in the UK and EU market. Job Description Overview We are seeking a highly experienced Senior Network SRE with deep expertise across multi-vendor network infrastructure, automation, and reliability engineering. The ideal candidate will possess strong technical leadership, hands‐on engineering capabilities … resilient, scalable, and observable network environments. Key Responsibilities Design, implement, and maintain highly available network solutions across routing, switching, firewalling, and wireless technologies. Apply SRE principles to improve network reliability, scalability, and performance. Develop and maintain automation workflows using Ansible, Salt, and related frameworks to reduce operational toil. Build ...

Lead Site Reliability Engineer

Location
United Kingdom
support economic growth across the UK. The Department for Business, Innovation, Science and Trade (BIST), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability … UK. BIST's Digital, Data and Technology (DDaT) directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure ...

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners Admin
Location
United Kingdom
Employment Type
Permanent
looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate, distributed tracing and application/infrastructure monitoring Experience troubleshooting … mentoring other engineers The opportunity You'll join a sizeable engineering capability working across complex, large-scale environments, taking a leading role in SRE and observability engineering. There is flexibility around some of the wider cloud/platform technology stack for candidates with genuinely strong Dynatrace and SRE expertise. ...

Lead Site Reliability Engineer (Dynatrace)

Location
London, United Kingdom
looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate, distributed tracing and application/infrastructure monitoring Experience troubleshooting … mentoring other engineers The opportunity You'll join a sizeable engineering capability working across complex, large-scale environments, taking a leading role in SRE and observability engineering. There is flexibility around some of the wider cloud/platform technology stack for candidates with genuinely strong Dynatrace and SRE expertise. ...

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners
Location
South West England, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £100,000 per annum
looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate, distributed tracing and application/infrastructure monitoring Experience troubleshooting … mentoring other engineers The opportunity You'll join a sizeable engineering capability working across complex, large-scale environments, taking a leading role in SRE and observability engineering. There is flexibility around some of the wider cloud/platform technology stack for candidates with genuinely strong Dynatrace and SRE expertise. ...

Hybrid AI-Driven SRE & Reliability Engineer

Location
United Kingdom
bet365 Group is seeking a Site Reliability Engineer to enhance system reliability, observability, and performance. You will treat reliability as a software problem, protecting uptime and driving improvements across critical systems. Responsibilities include building tools, dashboards, and automation, contributing to live incident resolution and post ...

Senior Site Reliability Engineer – AI-Driven Production Resilience (Hybrid)

Location
Greater London, England, United Kingdom
London seeks a Site Reliability Engineer/Senior Engineer to join Corporate Bank Production. The role focuses on embedding SRE principles, improving production stability, and partnering with global teams to manage risks and changes across platforms. You will define and monitor SLOs, enhance observability, and explore ...

Site Reliability Engineer – Platform & Automation

Location
Cambridge, England, United Kingdom
Bango plc is seeking a Site Reliability Engineer to own reliability, performance and continuous improvement of the Bango Platform end-to-end. You’ll merge platform engineering, automation and delivery ownership with proactive incident response and customer impact management, shaping the automation and security posture across ...

Site Reliability Engineer (Edv) - National Security

Location
Cheltenham, England, United Kingdom
Site Reliability Engineer (eDV) - National Security Location: Key National Security hubs Salary: £65,000-£90,000+ depending on experience Clearance: Active eDV required Level: Mid-level through to Senior … spending more time firefighting than improving reliability, this one should resonate. I'm supporting a National Security engineering team looking for an SRE who wants to work closer to the platform, improve how services behave in production and take real ownership of resilience rather than simply reacting to incidents. ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security. The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers … platform in close partnership with our lead/backend/staff engineers. Who you are (must-haves) 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm ...