226 to 250 of 518 Site Reliability Engineering Jobs in London

Senior Cloud Engineer

Location
Greater London, England, United Kingdom
operational excellence, security, reliability, and lifecycle management of cloud applications across multiple regions. The role sits at the intersection of cloud architecture, SRE, security, and developer experience, helping to raise the overall standard of cloud usage across PL Re. While experience in platform provisioning is required, this role does … hours on-call rotas as required. Support the full lifecycle of application workloads, including maintenance, patching, backups, upgrades, and controlled decommissioning. Reliability, SRE & Observability Drive improvements in reliability, resilience, performance, and efficiency through strong SRE practices. Own and enhance observability and monitoring, ensuring meaningful alerting, clear operational dashboards ...

Kubernetes SRE: DevOps, Observability & Reliability

Location
Greater London, England, United Kingdom
Cisco's Webex Engineering Group in London is seeking a Site Reliability Engineer to own the design, deployment, and operation of Kubernetes-based microservices. You will build reusable configurations and advance a robust, scalable platform. In this role, you will implement GitOps workflows with Argo CD, manage … canary releases and secret management with Vault, and monitor reliability with Prometheus and Grafana, collaborating across teams in a hybrid London office. #J-18808-Ljbffr ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
sits within a specialist engineering function responsible for designing and operating distributed platforms at scale. You will work closely with software engineers, architects, SRE teams and business stakeholders to modernise infrastructure, improve platform reliability and deliver robust cloud-native solutions. The ideal candidate will have significant experience with … databases. Experience with Kafka or other distributed messaging platforms. Understanding of CAP Theorem, consistency models and distributed database design principles. Platform Engineering and SRE experience. Experience supporting mission-critical financial services or transaction-processing systems. Exposure to multi-region or active-active architectures. Experience within Financial Services, FinTech, Payments ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
credentials Do You Have What It Takes? 3-6 years of hands-on experience in a similar role, with a strong emphasis on systems engineering, automation, and service reliability Proficient in at least one programming language such as Python, Go, Java, or C#, along with scripting skills … PowerShell Solid grasp of cloud platforms like AWS, including an understanding of how core services like EC2, ECS, Lambda, and DynamoDB operate under reliability constraints Practical experience using infrastructure-as-code tools like CloudFormation or Terraform In-depth knowledge of CI/CD principles and hands-on experience with ...

Principal Platform Security Engineer

Location
Greater London, England, United Kingdom
Type:**Permanent**Build a brilliant future with Hiscox****Position**: Principal Platform Security Engineer, London MarketAs a senior and influential member of the London Platform Engineering Chapter, the Principal Platform Security Engineer will set direction and lead by example in maturing our platform security practices. You will guide multiple Innovation … squads and Engineering Chapters, driving cloud-first adoption and championing secure-by-design initiatives.You will be an integral member of a Chapter spanning Platform, DevOps, and Site Reliability Engineers, and a core contributor to a Platform Engineering squad focused on continuous improvement across our cloud ...

Principal Platform Security Engineer

Hiring Organisation
Hiscox AG
Location
London, UK
Employment Type
Full-time
Type: PermanentBuild a brilliant future with HiscoxPosition: Principal Platform Security Engineer, London MarketAs a senior and influential member of the London Platform Engineering Chapter, the Principal Platform Security Engineer will set direction and lead by example in maturing our platform security practices. You will guide multiple Innovation squads … Engineering Chapters, driving cloud-first adoption and championing secure-by-design initiatives. You will be an integral member of a Chapter spanning Platform, DevOps, and Site Reliability Engineers, and a core contributor to a Platform Engineering squad focused on continuous improvement across our cloud ...

Front-Office SRE Lead: Observability & AI-Driven Reliability

Location
Greater London, England, United Kingdom
JPMorgan Chase & Co. in London is seeking a Lead Site Reliability Engineer to shape next‐gen SRE patterns, observability, and reliability across globally distributed trading systems. You will partner with front‐office traders, contribute production code (Java/Python/Kotlin), drive incident response, and lead … assisted reliability initiatives while collaborating with infrastructure, cloud, and security teams. #J-18808-Ljbffr ...

Director of DevOps & SRE

Location
Greater London, England, United Kingdom
opportunities, create and manage relationships, and turn relationship insights into action with increased productivity and transparency. We are looking for a Director of DevOps & SRE to lead the team responsible for the reliability, security, and scalability of theInvestorFlowplatform. This is a strategic leadership role. You will set technical direction … DevOps/SRE roadmap, and lead a team of technical leads, senior and mid-level engineers across both disciplines,while partnering closely with Engineering and Product on technology and product roadmap planning, resourcing, and delivery sequencing. You will stay technical enough to lead architectural review, challenge design decisions ...

SC Cleared DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £570 per day + Inside IR-35
must have been used within the last 12 months, with a minimum of 3 years remaining before expiry. This is a hands-on platform engineering role focused on creating secure, resilient and highly automated AWS environments that enable teams to develop, test and deploy software efficiently while maintaining strong … platform engineering best practices. What We're Looking For * Strong AWS cloud engineering and platform experience. * Background in DevOps, Platform Engineering, SRE or Cloud Infrastructure roles. * Experience with Infrastructure as Code, ideally Terraform, CloudFormation or AWS CDK. * Strong CI/CD pipeline experience. * Experience with containerisation technologies ...

Senior Nework Programmer - Algo trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Network SRE – 250,000-300,000 total compensation – 4 days in officeQuant Capital is urgently looking Network SRE for our high profile client. Our client is a leading quantitative trading company and liquidity provider. Their focus on technology has allowed them to deeply penetrate the market and gain market share. … Shared Engineering team that focuses on designing, developing, and maintaining infrastructure and tools. The team requires a Network Site Reliability Engineer (SRE) with strong network fundamentals, problem-solving skills, and a keen interest in diverse tools and techniques. The role involves collaborative work across various teams, exploring ...

MLOps Engineer

Location
Greater London, England, United Kingdom
2025. Learn more at www.coreweave.com. We're proud to be a Living Wage accredited Employer. What You'll Do CoreWeave’s Physical AI Platform Engineering team builds and scales the data and workflow backbone powering advanced engineering simulation and AI workflows. Our ambition is to become the super … engineers on production‐grade ML practices. Who You Are 5–6+ years of professional experience in MLOps, ML platform engineering, ML infrastructure, or SRE/DevOps for production machine learning systems. Proven experience building, operating, and automating production ML pipelines covering experiment tracking, model registries, artifact versioning, dataset management ...

Observability AVP: SRE & Cloud Telemetry Lead

Location
Greater London, England, United Kingdom
Citigroup Inc. is seeking a Site Reliability Engineer (SRE) to lead hands-on deployment of observability principles in a large-scale environment. The role focuses on reliability, performance, and resiliency across applications and services. You will migrate monitoring tools to Google Cloud Observability and Grafana using OpenTelemetry … author reusable deployment solutions, and onboard engineering teams to new telemetry standards. This is a senior technical role within Production Management. #J-18808-Ljbffr ...

SRE Engineer – FinTech Reliability, Observability & Cloud

Location
Greater London, England, United Kingdom
Hamilton Barnes Associates Limited is seeking a Site Reliability Engineer to work at the intersection of software engineering and infrastructure. You'll develop internal platforms, tooling, and automation across Linux, distributed systems, and cloud-native technologies to improve reliability and operational efficiency in a global production ...

Senior DevOps Analyst

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
Azure/GCP.Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or ARM templates. Build and manage containerized applications using Docker and Kubernetes. Ensure system reliability, scalability, security, and performance. Implement monitoring, logging, and alerting solutions. Collaborate with development, QA, and security teams. Troubleshoot production issues and lead incident response. … stack, Splunk. Version Control: Git (GitHub/GitLab/Bitbucket).Security: IAM, secrets management, vulnerability scanning. Experience & Qualifications8+ years of experience in DevOps/Site Reliability/Infrastructure Engineering. Strong understanding of DevOps and CI/CD best practices. Experience supporting high-availability, production systems. Experience in Agile ...

Architect & Delivery Lead (68018)

Location
Greater London, England, United Kingdom
Mandatory Skills Broad and deep technical knowledge across infrastructure, cloud, networking, security, IAM, SRE, and operations. Key Responsibilities Define and own the end‐to‐end technical architecture and target operating model across all six service towers Lead the three‐phase transformation programme (Stabilise → Optimise → Transform), managing timelines, milestones, risks … dependencies Make architectural decisions across IAM, Cloud, SRE, Network, Data, and Security — ensuring coherence, reusability, and alignment with business objectives Establish and chair the Architecture & Engineering Governance board, providing technical assurance across all workstreams Own the programme roadmap, resource plan, and financial model — tracking cost savings, team reduction trajectory ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
service operations including incident, problem, change, event, monitoring, resilience, continuity, capacity, and on‐call models. Strong understanding of ITIL practices blended with modern DevOps, SRE, Agile and platform‐engineering approaches. Broad technical awareness across cloud platforms, application architectures, data platforms, networks, observability tooling, security‐by‐design, and automation. Ability … engineering, service readiness, and live‐service best practices. Key experiences Running and growing operational or engineering teams in a Managed Service, SRE, DevOps, or live‐service environment—with responsibility for hiring, coaching, development and performance. Leading high‐pressure operational functions including incident management, problem resolution, major incident ...

Principal Recruitment Consultant

Hiring Organisation
Harrison Clarke
Location
London Area, United Kingdom
class engineering teams. Our mission is simple: align exceptional engineers with groundbreaking businesses that are shaping the future. We operate across Cloud (DevOps, SRE, Platform Engineering, DevSecOps, Performance Engineering) and Data & AI (Machine Learning, MLOps, Data Engineering, AI Research, and related domains). Through our portfolio … recruitment consultant with experience placing technical talent into startups or technology companies Strong domain knowledge or demonstrable interest in Cloud, Data, AI, DevOps/SRE, Platform Engineering, Machine Learning or related fields Excellent stakeholder management skills, comfortable engaging at founder and senior engineering leadership level Entrepreneurial attitude with ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Take on ambiguous reliability, scalability, and efficiency challenges and drive solutions across SRE and development teams Build and run large-scale, massively distributed, fault-tolerant systems supporting the Genesis platform Optimize existing systems and build infrastructure Eliminate toil through automation to improve uptime and rate of change Cultivate … culture of reliability throughout the organization Guide technical decisions balancing system health with product priorities Ensure long-term health, maintainability, and reliability of services Perform capacity planning and performance analysis Proactively prevent incidents Work across teams to build robust, reusable solutions Requirements Strong software engineering skills ...

DevOps Solution Architect

Location
London, United Kingdom
security teams to ensure alignment and successful delivery. Provide mentorship and technical leadership to Platform Engineers and wider engineering teams. Promote DevOps and SRE best practices across the organisation. Enable teams by defining clear platform capabilities, usage patterns, and self-service models. Security, Risk & Compliance Ensure secure-by-design … Skills Experience defining platform strategies or roadmaps at an organisational level. Knowledge of observability and monitoring solutions (Prometheus, Grafana, ELK, Splunk). Experience implementing SRE practices (SLIs, SLOs, error budgets). Familiarity with compliance frameworks and continuous compliance tooling. Experience with large-scale or enterprise platform transformation initiatives. Exposure ...

Senior SRE Technical Lead — Reliability & Observability

Location
Greater London, England, United Kingdom
leading global financial markets infrastructure provider is seeking a Technical Lead SRE in Greater London. In this role, you will enhance the reliability engineering capabilities, collaborating with various teams to establish observability standards and ensure operational excellence. The ideal candidate will have over 10 years of experience … SRE or related fields, strong AWS and Kubernetes skills, and a proven track record in building resilient platforms. Join us to make a significant impact in financial markets infrastructure. #J-18808-Ljbffr ...

Senior SRE, Observability & Cloud Reliability

Location
Greater London, England, United Kingdom
company is seeking a Principal Site Reliability Engineer, Infrastructure Observability to guide a team of SREs focused on observability, reliability, and scalable cloud/on‐prem solutions. … role requires hands‐on expertise and collaboration with diverse partners to drive measurable improvements. The ideal candidate has extensive cloud experience, DevOps/SRE leadership, and strong automation skills, with a track record in designing resilient systems and implementing effective monitoring and #J-18808-Ljbffr ...

Data Platform SRE — Cloud-Scale Reliability

Location
Greater London, England, United Kingdom
Apple Inc. in London is seeking a Data Platform Infra Site Reliability Engineer to own the reliability, performance, and scale of Apple's data platform services. You will work on large-scale distributed systems, debugging replication and consensus issues, and writing production code in Go or Python ...

SRE for AI-Driven Financial Infrastructure

Location
Greater London, England, United Kingdom
company Group is seeking a Site Reliability Engineer (SRE) to improve, manage, and monitor production-critical infrastructure and data pipelines. You will work on production systems, sometimes embedded with software teams, to increase reliability and reduce operational risk. We value a growth mindset, strong problem-solving ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
interconnect enables a step change in system level performance. The Role We're hiring a senior infrastructure engineer to own the internal platform our engineering teams build, test, and collaborate on, spanning compute, storage, networking, build and CI/CD systems, and the security that protects … manual steps, and reliability gaps that slow the wider engineering team down. Skills & Experience 5+ years in infrastructure, platform, or DevOps/SRE engineering, with a track record of owning systems end to end. Strong infrastructure‐as‐code and automation skills: Terraform, Ansible, and scripting in Python ...

Senior SRE: GCP & Kubernetes, Automation Lead

Location
Greater London, England, United Kingdom
seeking a Senior Site Reliability Engineer to join our agile engineering team in London. You will drive reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership. You will design, implement, and operate cloud-native infrastructure on GCP using Kubernetes ...