51 to 75 of 664 Site Reliability Engineering Jobs in the UK

Site Reliability Engineer (SRE) - Cloud & Automation

Hiring Organisation
Spencer Rose Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 60,000 - 70,000 Annual
Site Reliability Engineer (SRE) - Cloud & Automation London, Docklands (hybrid) £60,000 - £70,000 per annum + annual discretionary bonus On behalf of a leading financial services organisation, I'm looking for a highly capable Site Reliability Engineer (SRE) to drive the adoption of SRE methodologies across … days per week in their Canary Wharf office, therefore you must be within a reasonable commute of London. Responsibilities: Lead the implementation of SRE practices across the organisation, working closely with infrastructure teams to optimise deployment processes and embed automation and operational excellence. Enhance observability and reliability , defining ...

Senior Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part … provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world. Provide technical guidance and serve as a function-wide SRE subject matter expert, driving reliability decisions across multiple products and influencing the adoption of leading-edge observability technologies. #J-18808-Ljbffr ...

SRE Technical Lead

Hiring Organisation
83zero Limited
Location
Wokingham, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
SRE Technical Lead Location: Hybrid (UK - office, client site and home-based working) Salary: Up to £100,000 + 5% bonus We're looking for an experienced SRE Technical Lead to take ownership of the reliability, availability and operational excellence of critical platforms within complex, multi-vendor environments. … clearance. Be a sole UK national. Unfortunately, candidates who do not meet both of these essential requirements cannot be considered. The Role As the SRE Technical Lead, you will: Define and drive the SRE strategy, standards, SLAs, SLOs and error budgets. Embed reliability engineering principles into platform ...

Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
containerization technologies (Docker)and orchestration platforms (Kubernetes, e.g., Amazon EKS, GoogleKubernetes Engine) to support cloud-native application deployments. Site Reliability Engineering (SRE): Embrace a "you build it, you run it" mindset. Take ownership ofthe reliability, performance, and availability of the cloud platforms andservices you build. Implement … monitoring, logging, and alerting tools (e.g.,Prometheus, Grafana, Splunk, ELK stack, CloudWatch, Stackdriver). A strong commitment to Site Reliability Engineering (SRE)principles and practices, including operational ownership. Preferred Qualifications & Skills Public cloud provider certifications (e.g., AWS CertifiedSolutions Architect, AWS Certified DevOps Engineer, Google CloudProfessional Cloud Architect ...

Senior Associate Engineer (Digital Product Team)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
broad, T-shaped engineer, you will also bring a particular interest in the operational side of engineering, Site Reliability Engineering (SRE), observability, and the delivery pipelines that get our work safely to production. This is a blended role: you will spend plenty of your time building … while actively reducing toil, technical debt, and manual effort through automation. Take ownership of personal learning paths, continually acquiring new skills, particularly across operations, SRE, and observability, to enhance team solutions and individual performance. Skills, Knowledge, and Experience You are a capable, hands-on software developer who enjoys designing ...

Senior Associate Engineer (Digital Product Team)

Hiring Organisation
Canada Life
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
broad, T-shaped engineer, you will also bring a particular interest in the operational side of engineering, Site Reliability Engineering (SRE), observability, and the delivery pipelines that get our work safely to production. This is a blended role: you will spend plenty of your time building … while actively reducing toil, technical debt, and manual effort through automation. Take ownership of personal learning paths, continually acquiring new skills, particularly across operations, SRE, and observability, to enhance team solutions and individual performance. Skills, Knowledge, and Experience You are a capable, hands-on software developer who enjoys designing ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Chief Technology Office, youwill solve complex and broad business problems, demonstrate … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team’s applications ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Newcastle upon Tyne, England, United Kingdom
procedures for incident response and operational tasks. Collaborate with cross-functional teams to review and provide feedback on technical designs, ensuring alignment with SRE principles. Participate in on-call rotations and handle critical incidents with confidence and expertise. Continuously improve documentation for systems and services, contributing to a knowledge-sharing … incident management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For: Specialized SRE observability experience with New Relic or DataDog. Familiarity with OpenTelemetry, AIOps, MLOps, or SecOps. Logistics: Location: Newcastle, UK - In-Office (at least 4 days ...

DevOps Engineer- London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Develop and maintain automation solutions using scripting and DevOps tooling Support observability initiatives through monitoring, logging, and alerting platforms Collaborate with Technology, Infrastructure, Data Engineering, and Production Support teams to resolve issues and improve service performance About You ------------- Experience supporting production batch processing environments within complex enterprise platforms Strong … Shell scripting and Ansible Strong incident management, root cause analysis and operational support experience Knowledge of DevOps and Site Reliability Engineering (SRE) principles and best practices Experience supporting Market Risk, FRTB, VaR, SVaR, P&L or other regulatory reporting environments is advantageous Minimum of 5 years' experience ...

Senior Lead Site Reliability / DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top‐notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability/DevOps Engineer at JPMorgan Chase within the Commercial & Investment Bank … Drive significant business impact through your capabilities and contributions, and apply deep technical expertise and problem‐solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices ...

Senior DevOps Engineer

Hiring Organisation
Anson McCade
Location
Manchester Area, United Kingdom
while continuing to develop your expertise. Why Join This Team? Work on large-scale cloud and platform engineering projects Exposure to modern DevSecOps, SRE and AI-assisted engineering practices Access to industry-leading training and certifications Join a thriving engineering community of over 1,000 specialists Opportunity … solutions Implementing Infrastructure as Code using Terraform Driving DevSecOps best practices across engineering teams Supporting production environments and resolving complex platform issues Applying SRE principles to improve resilience, reliability and performance Implementing observability solutions including monitoring, metrics, logging and alerting Supporting incident management and post-incident improvement activities ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
support this growth by designing, provisioning then supporting the platforms to enable this. We are seeking an experienced Site Reliability Engineer (SRE) to join our Cloud & Infrastructure team. The successful candidate will be responsible for designing, operating, automating, and continuously improving enterprise-scale Azure platforms, ensuring high availability … resiliency, security, and performance. The role combines software engineering, cloud architecture, infrastructure automation, and operational excellence to improve service reliability and reduce operational overhead through automation and engineering best practices. The ideal candidate will have strong experience with Microsoft Azure, DevOps, Terraform, API, disaster recovery planning ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Senior Site Reliability EngineerApplylocations: London (82)time type: Full timeposted on: Posted Todayjob requisition id: JR101516**Senior Site Reliability Engineer (SRE) - GCP/Kubernetes****About the Role**We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small … agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership.The successful candidate will split their time between providing expert operational support for our critical systems and leading exciting new infrastructure ...

Lead Cloud Infrastructure & Site Reliability Engineer - Contract

Hiring Organisation
Caspian One Ltd
Location
Sheffield, Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 525 - 625 Daily
partnering with a leading global organisation undergoing significant investment in its cloud and data platforms. As part of a high-performing engineering team, you'll play a key role in operating, improving, and automating a large scale Azure-based platform that supports critical cybersecurity and analytics capabilities. This … Maintaining engineering standards, documentation, and change controls Required Experience We're particularly interested in candidates with: Strong Site Reliability Engineering (SRE) or Cloud Infrastructure Engineering experience Deep Azure platform knowledge Proven Terraform and Infrastructure-as-Code expertise Experience with Azure DevOps and CI/ ...

Software Engineer III, Site Reliability Engineering, Traffic Network Load Balancing

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Science or Engineering. 2 years of experience designing, analyzing, and troubleshooting large-scale distributed systems. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault‐tolerant systems. SRE ensures that the company Cloud's services … internally critical and our externally‐visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever‐watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Ipswich, England, United Kingdom
This is a blended Network Operations and Site Reliability Engineering role within Professional Services, combining hands‐on network engineering with SRE principles to ensure the reliability of BT's fixed network infrastructure. You will implement flawless network change, resolve network and platform issues, and drive … automation to improve reliability and efficiency — developing your skills across network operations and SRE disciplines to deliver brilliant customer experience. You will build effective working relationships internally and externally and contribute as a technically talented network and reliability expert, challenging those around you to perform at their best. ...

Site Reliability Engineer- London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
extend and will be a hybrid role that will be based in London. Our client is seeking an experienced Site Reliability Engineer (SRE) with a strong focus on Observability and Monitoring Platforms. The successful candidate will play a key role in enhancing the organisation's monitoring, alerting … streamline operational processes and improve reliability. Collaborate with engineering, infrastructure, and support teams to improve system resilience and operational performance. Define and implement SRE best practices, including monitoring standards, alert management, incident response, and operational readiness. Perform troubleshooting and root cause analysis of platform and application issues. Support capacity ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing … production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems ...

Software & Data Engineers

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
native, low‐latency system that operates at global scale and underpins critical investment products used worldwide. You will work on challenging problems across software engineering where performance, data quality, and reliability are non‐negotiable, leveraging modern cloud (AWS) and AI‐assisted development tooling to accelerate delivery without compromising … encouraged to apply. Areas We Value Experience In Software Engineering Data Engineering Cloud & Platform Engineering Site Reliability Engineering (SRE) AI & Machine Learning DevOps & Automation Architecture & Distributed Systems Analytics & Data Platforms Benefits LSEG offers a range of tailored benefits and support, including healthcare, retirement planning ...

Platform Engineer (Security)

Hiring Organisation
Jobleads-UK
Location
United Kingdom
manages all applications and next steps. Our partner is looking for a Platform Engineer (Security) based in United Kingdom. Join a cloud-native engineering environment where security is embedded into every stage of platform development and operations. In this role, you will strengthen the security posture of modern infrastructure … integrate security into platform operations and development workflows. Requirements 3–6 years of experience in Platform Engineering, Site Reliability Engineering (SRE), DevOps, Cloud Engineering, Security Engineering, or a related field. Hands-on experience securing production cloud environments on AWS, Microsoft Azure, Google Cloud Platform ...

SRE Technical Lead

Hiring Organisation
Adecco
Location
Reading, Berkshire, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 90,000 Annual
SRE Technical Lead Reading/Hybrid (UK-based - mix of home, office, and client site) Must be eligible for SC Clearance We are seeking an experienced SRE Technical Lead to act as the technical authority for Site Reliability Engineering across complex, large-scale platforms. This … , availability, and operational excellence across multi-team and multi-vendor environments. You will combine hands-on engineering expertise with strategic leadership, ensuring SRE practices are Embedded across the full service life cycle-from design through to production operations. As the SRE Technical Lead, you will: Define and implement ...

Engineering Manager, DevOps

Hiring Organisation
Jobleads-UK
Location
United Kingdom
About the Engineering Organization The Engineering Team at Loop is a balance of agility, consistency, and performance. These are the pillars that allow the team to constantly and consistently deliver value that matters to customers. That customer intimacy is what allows our engineering teams … joining high-severity incident calls when needed. Your experience: 2+ years of DevOps leadership managing DevOps or Site Reliability Engineering (SRE) teams. 7+ years in hands-on platform or infrastructure roles. You have a strong self-service track record, having delivered internal platforms or portals that empowered ...

SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed … Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self ...

Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office)

Hiring Organisation
Hamilton Barnes
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP 450 - 475 Daily
enabling proactive monitoring, faster incident resolution, and improved platform reliability through modern observability practices - with deep expertise in Grafana, OpenTelemetry, distributed tracing, SRE, event-driven architecture, and Azure Integration Services. Key Responsibilities Define and implement enterprise observability strategies, standards, and governance frameworks Design and manage observability solutions covering metrics … service health across Azure services Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience Drive root cause analysis, incident investigations, and continuous service improvement initiatives Champion operational excellence through proactive monitoring, automation ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...