651 to 675 of 1,104 Site Reliability Engineering Jobs in the UK

Senior Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
2025. Their platform is redefining the sector, and with revenues nearly 10x since the start of last year, they're continuing to expand their Engineering team to match the ambition of their product and customers. The product is real-time, data-rich and AI-native, creating complex engineering … infrastructure as code Strong understanding of networking, IAM, managed services and secure cloud architecture Experience owning observability, logging, APM or distributed tracing tooling An SRE mindset across SLOs, reliability, incident response and toil reduction Clear communication skills and the ability to influence how engineers build and operate services This ...

SRE Engineer - Global Trading Platform | Cloud + Automation

Location
Greater London, England, United Kingdom
Mandarin Recruitment is looking for a Site Reliability Engineer to join a fast-growing Singapore-based financial services company. The ideal candidate will have a bachelor's degree in computer science and at least three years of experience in DevOps, ideally focusing on infrastructure support. Successful applicants will ...

Senior Platform Engineer – Remote UK (Cloud/SRE)

Location
Greater London, England, United Kingdom
Hudl in London, United Kingdom, is seeking a Senior Engineer to join our Platform Engineering team. You’ll work on site reliability, cloud infrastructure, observability and production operations to keep Hudl’s platform highly available, scalable and secure. You’ll lead with technical excellence, mentor engineers ...

Staff Cloud SRE for AI/ML Platform & GPU Compute

Location
Greater London, England, United Kingdom
Wayve is seeking a founding Staff Cloud Site Reliability Engineer to shape the reliability of large-scale AI systems and GPU compute infrastructure. … will build and scale the reliability foundations of our AI cloud platform, including model development and GPU compute environments. This is a founding SRE role situated at the intersection of AI research, cloud infrastructure, and operations. You will define frameworks, automate standards, and ensure scalable, secure deployments with ...

Cloud SRE: Build Reliable, Scalable Platforms

Location
Manchester, England, United Kingdom
Lloyds Banking Group is seeking a Site Reliability Engineer to bolster cloud-hosted applications and services. You will join a collaborative team focused on reliability, scalability and performance of platforms, while supporting DevOps initiatives and automation. You'll work with GCP, Terraform, GitHub and CI/ ...

Network & Infrastructure Tooling / Automation Specialist/ Architect - freelance - hybrid, London, UK

Location
Greater London, England, United Kingdom
driven workflows. Observability, Operations & Resilience Develop automation for fault detection, root cause analysis, and remediation. Integrate telemetry, logs, and metrics into observability platforms. Support SRE-style practices including error budgets, reliability metrics, and continuous feedback loops. Improve MTTR and operational consistency through standardised tooling and workflows. Governance, Compliance & Secure … MPLS, Traffic engineering, QoS, ExpressRoute, Direct Connect Observability & Operations : Observability platforms, Telemetry, Logs, Metrics, Fault detection automation, Root cause analysis, Remediation automation, SRE practices, Reliability metrics Platform, Cloud & Service Integration: IPAM integration, CMDB integration, Platform engineering, DevOps, Lifecycle management, Cloud networking, AWS, Azure, Data centre networking Security ...

Principal DevOps Engineer

Location
Greater London, England, United Kingdom
define standards and reference architectures, lead complex initiatives across Windows/Linux platforms (with deep expertise in Windows clustering), and own end-to-end reliability—from Infrastructure as Code to release engineering and observability. What you will be doing: Act as technical lead for DevOps/Platform/… platform roadmaps What you will bring to the role: Bachelors degree in Computer Science or related field 7+ years (or equivalent) in DevOps/SRE/Infrastructure Engineering , including leadership in complex environments Expert-level experience designing and operating Windows Server HA and clustering (Failover Clustering and related components ...

VP DevOps/SRE Lead — Cloud Infra, Kubernetes & CI/CD

Location
Glasgow, Scotland, United Kingdom
J.P. Morgan seeks a hands-on Lead DevOps/SRE Engineer - Vice President to own automation, reliability and production operations of AI/ML platforms. You will build CI/CD, observability, and incident-management practices that keep services stable across international markets. As part of the IPB Tech … AIML team, you will lead reliability engineering in a regulated environment, mentor engineers, and collaborate with cross-functional teams to raise standards and security." #J-18808-Ljbffr ...

Staff Site Reliability Engineer

Hiring Organisation
Genomics
Location
Oxford, United Kingdom
that usable, and MystraAI is the agentic layer we are building on top of it. This is a Staff-level role that owns the reliability, performance, security and integrity of that infrastructure end-to-end — and sets the technical direction that other teams build on.You will lead the analytical … engine at source level rather than as a black box — and ideally have contributed code upstream.Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines.Data-as-a-Service productisation. You think in terms of data ...

Software Engineer — Observability Instrumentation

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
ideas take time to evolve. Together we're building a world-class platform to amplify our teams' most powerful ideas. As part of our engineering team, you'll shape the platforms and tools that drive high-impact research - designing systems that scale, accelerate discovery and support innovation across … engineering teams to debug and operate servicesDevOps tooling experience using tools such as Terraform, ArgoCD, Helm or JenkinsInterest in AI engineering and SRE practices to improve incident response and RCADesirable but not essential experience includes: Prometheus/PromQL, VictoriaMetrics, OpenSearch, Grafana or similar observability backendsAuto-instrumentation, distributed tracing ...

Production Engineering Lead

Hiring Organisation
Ncounter Limited
Location
SW1, Latchmere, Greater London, United Kingdom
Employment Type
Permanent
Salary
£200000 - £250000/annum plus Bonus & Package
Production Engineering Lead £200,000-£220,000 London, New York, Montreal or Singapore | Permanent Ncounter is looking for a Production Engineering Lead to take technical and leadership ownership of a global … engineering team responsible for the reliability of large-scale, business-critical production infrastructure. This is an opportunity for an experienced Production Engineer, SRE, Platform Engineer or Infrastructure Engineer who has remained technically hands-on while progressing into leadership. Financial services experience is not essential. We are particularly interested ...

Platform Reliability Engineer - Automation & Observability

Location
Manchester, England, United Kingdom
Onyx-Conseil is seeking an experienced Site Reliability Engineer to join a high-performing team supporting data product and platform groups. The role focuses on reliability, scalability, observability, deployment, and operational support across complex production environments. You will … collaborate with engineering, platform, and operations teams to enhance monitoring, logging, incident response, and automation, leveraging Kubernetes, Helm, the ELK stack, and modern SRE practices to ensure #J-18808-Ljbffr ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Cheltenham, England, United Kingdom
Envitia in the UK is hiring a DevOps/Platform Engineer to design, build and operate secure cloud capabilities within classified government environments. You will work across AWS and cloud-native technologies, delivering reliable, scalable ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Manchester, England, United Kingdom
Envitia in the UK is hiring a DevOps/Platform Engineer to design, build and operate secure cloud capabilities within classified government environments. You will work across AWS and cloud-native technologies, delivering reliable, scalable ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Greater London, England, United Kingdom
Envitia in the UK is hiring a DevOps/Platform Engineer to design, build and operate secure cloud capabilities within classified government environments. You will work across AWS and cloud-native technologies, delivering reliable, scalable ...

Client Implementation Engineer

Hiring Organisation
Bright Purple Resourcing
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£45,000
available either at their HQ in Glasgow, or there are flexible/remote options available for those further afield. The role sits between platform engineering, networking and client delivery, helping deploy and support a real-time analytics platform used within financial markets. This is high end, market leading stuff. … product as well as work directly with clients. They're looking for someone with experience in areas like Linux systems, DevOps, integration engineering, SRE or technical infrastructure. Financial markets experience is useful but definitely not essential. The role: Design, deliver and support analytics solutions for electronic trading clients, from ...

Test Environment Manager

Location
Greater London, England, United Kingdom
This role ensures that test environments are reliable, scalable, automated, observable, and aligned with modern SDLC and DevOps practices. The TEM drives technical excellence, SRE-inspired culture, and continuous improvement across development, QA, and operations teams. Key Operational Responsibilities: Environment Automation & Lifecycle Management Design and implement Infrastructure as Code … based on reliability KPIs. Promote a culture of shared ownership, blameless problem-solving, and strong cross-team collaboration among development, QA, DevOps, and SRE teams. Capacity Planning & Scalability Forecast environment capacity requirements based on usage trends, test cycles, and upcoming projects. Ensure infrastructure elasticity and scalability to meet future ...

Cloud Software Engineer

Location
Mansfield, England, United Kingdom
within platform services and APIs Assist in building and maintaining monitoring, alerting and dashboarding solutions using Grafana and OpenTelemetry Work with the Cloud SRE to understand platform reliability metrics and ensure application code supports robust, performant and maintainable services Develop a working understanding of cloud infrastructure across GCP, Azure … collaborating with senior engineers and the Cloud SRE on infrastructure requirements Develop and deploy containerised application workloads using Docker and Kubernetes, working with the Cloud SRE on cluster-level operations as needed Work within CI/CD pipelines using GitHub Actions, delivering application code through automated workflows from commit through ...

Observability SRE AVP: Cloud Observability & Migrations

Location
Greater London, England, United Kingdom
Citi is seeking an experienced Site Reliability Engineer to lead end-to-end observability and resiliency initiatives in a large-scale environment. You will migrate monitoring tooling … Google Cloud Observability and Grafana, implement OpenTelemetry instrumentation, and author reusable deployment solutions for OpenShift/Kubernetes and VM environments. The role requires deep SRE knowledge, strong collaboration with application teams, and a focus on reliability, performance, and automation within #J-18808-Ljbffr ...

DevOps Engineer, Blockchain Infra (Fully Remote)

Hiring Organisation
Binance
Location
London, United Kingdom
Dubai/Ireland, Dublin/United Kingdom, London/Europe/Southern Europe/Eastern EuropeEngineering – DevOps/SRE/Full-time: Remote/RemoteBinance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million … improve deployment processes and platformdeveloper experience.Participate in on-call rotations and incident response.Requirements 3+ years of experience as a DevOps Engineer, Platform Engineer, SRE, or Infrastructure Engineer.Strong experience with at least one major cloud platform: AWS/GCP/AzureHands-on experience managing Kubernetes production environments.Experience with Karpenter for Kubernetes ...

Principal Engineer I — Prepurchase Platform

Location
City Of London, England, United Kingdom
Summary: JOB DESCRIPTION –Principal Engineer I, Prepurchase Platform Location:London, UK Division:TicketmasterUK Limited LineManager:VP Engineering, Accounts, Identity and Pre-Purchase Contract Terms:Full-time permanent, 40h/per week THE TEAM The Prepurchase Platform group owns the backend services that power every fan's path … strategies. Collaborate with the Enterprise Architecture team on architectural decisions and evolve engineering standards across the Prepurchase domain. Partner with product, security, and SRE teams to align technical decisions with business priorities. Drive observability improvements - ensuring services are instrumented for monitoring, alerting, and rapid incident response. Identifyandeliminatesingle points ...

Senior DevOps Engineer

Location
Harwell, England, United Kingdom
create GitHub Actions/Argo CD pipelines for zero-touch deployments. Security hardening – lead security-posture reviews, implement GuardDuty, CloudWatch and IAM best practices. SRE & monitoring – uphold SLAs through observability stacks, proactive alerting and performance tuning of distributed systems. Collaboration & enablement – automate repetitive tasks, mentor developers and champion DevSecOps best … comfortable working alone, supporting team members as well as interacting with clients if required Qualifications, experience and skills Proven experience in DevOps/SRE, with recent focus on AWS cloud engineering. Expert in AWS core services (EC2, VPC, IAM, S3, ALB/ELB, CloudFront, ECR/ECS, Control Tower ...

Site Reliability Engineer Infrastructure View role →

Location
Greater London, England, United Kingdom
highly concurrent, agent-driven system. Drive cost and performance optimization across compute-heavy scanning workloads. What we're looking for 4+ years in SRE/infrastructure roles operating production Kubernetes at scale. Strong grasp of network and workload isolation (namespaces, gVisor/Firecracker-style sandboxing, or similar). Experience with ...

Principal Platform Engineer

Location
Caerphilly, Wales, United Kingdom
virtualisation infrastructure using VMware and/or Hyper-V Manage backup and disaster recovery using Veeam Use Infrastructure as Code to automate environments Security & Reliability Apply strict security standards required for public sector clients Lead incident response and write post-incident reports Improve system reliability and performance Leadership … security, scalability, and cost Mentor other engineers and lead technical decisions on platform projects Experience Required Senior-level experience as a platform, infrastructure, or site reliability engineer Strong experience with private cloud or on-premise infrastructure Hands-on experience with VMware and/or Hyper-V, and Veeam ...

SRE Lead: Scalable, Reliable Betting Platform (Hybrid)

Location
Leeds, England, United Kingdom
evoke is seeking a Site Reliability Engineer to join our betting and gaming platforms, focusing on reliability, scalability and performance. You will work with observability, automation and engineering to deliver robust customer experiences across services. You will lead incident response, perform postmortems and drive resilience improvements. ...