76 to 100 of 572 Site Reliability Engineering Jobs in London

Principal Platform Engineer (12 Month FTC)

Location
Greater London, England, United Kingdom
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
Site Reliability Engineer (SRE) - Assistant Vice President is a technical professional responsible for the hands‐on execution, technical implementation, and deployment of SRE and observability principles in a complex, critical, and large-scale multi-disciplinary environment. In this role, you will apply a deep understanding of multiple technology … authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical onboarding support to application teams. We are seeking a passionate and experienced SRE to join our Production Management team. In this role, you will be instrumental in executing our strategy for end‐to‐end observability and resiliency, collaborating ...

AI Technical Platform Leader

Location
Greater London, England, United Kingdom
business stakeholders to mold and implement strategy. The role applies broad technical knowledge with depth in generative AI, agentic systems, enterprise platforms and engineering governance to ensure that AI platforms at WTW are optimally configured to achieve company vision and imperatives. It has end-to-end ownership of implementation … operations in a large, global enterprise: Software or platform engineering for global, enterprise-scaled solutions Management of Site reliability engineering (SRE) programs for mission-critical systems Creation and management of DevOps practices for automated, consistent, and secure solution deployment in regulated environments. Literacy in global compliance ...

TechOps & Support Engineer, Amazon MGM Studios | Technology Operations & Support

Location
Greater London, England, United Kingdom
software and devices across Amazon MGM Studios global productions. TechOps teams support the Studios production personnel (cast and crew) and Studios business teams (development, engineering, programming and marketing). We work in a team environment and regularly interact with production personnel and studio executives at all levels.Regular activities include … Bachelor's degree in Systems Engineering, Computer Science, or related field or relevant work experience - Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration - Experience working with Linux - Experience in systems engineering - Experience ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
employ over 130 colleagues across Jersey, Geneva, London, Singapore, New York and Shanghai. Purpose and Overview of Role This is a hands‐on senior engineering role within the Platform Engineering team, which forms part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure … platforms and developer tooling that enable our engineering and quantitative research teams to deliver software reliably, securely and at scale. The role will contribute to the design, build, automation and operation of a hybrid production platform across AWS and on‐premises environments, with a particular focus on the HashiCorp ...

Platform Engineer

Location
Greater London, England, United Kingdom
Platform Engineer Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Segun Ikuesan Description This is a hands‐on engineering role within the Platform Engineering team, which forms part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure, platforms and developer … tooling that enable our engineering and quantitative research teams to deliver software reliably, securely and at scale. The role will contribute to the design, build, automation and operation of a hybrid production platform across AWS and on‐premises environments, with a particular focus on the HashiCorp platform, including Nomad ...

Site Reliability Engineer - Frontend

Hiring Organisation
Capital On Tap
Location
London, UK
Employment Type
Full-time
just getting started! ðLondon, Old Street | ð 2 Days in OfficeSRE at Capital On Tap ðAt Capital On Tap, we run a hybrid embedded SRE model - We aim to work closely with the teams to provide them the best support. As a Site Reliability Engineer (SRE) you will … apply. Interview process ðFirst stage: 30 minute intro and values call with Talent PartnerSecond stage: 60 minute CV overview and technical chat with the SRE team lead and the SRE & Platform Engineering ManagerThird stage: 75 minute technical exercise & questions with the SRE lead Final stage: 30 minute chat with ...

Devops SRE

Location
Greater London, England, United Kingdom
Cloud Engineering team is seeking a seasoned and passionate Senior Cloud Engineer with deep hands‐on development and cloud engineering expertise. In this role, you will serve as a key technical contributor within a cloud‐focused engineering team, working on one of the Group’s flagship initiatives … best practices and business goals. Required Skills & Experience Core Cloud & DevOps Competencies Extensive experience in DevOps or Site Reliability Engineering (SRE) roles across consumer or SaaS environments. Strong expertise in deploying and managing production‐grade Kubernetes clusters and containerised services. Hands‐on experience with Kubernetes ...

Venue & Studio Deployments System Engineer, Event Productions

Hiring Organisation
Amazon
Location
London, UK
Employment Type
Full-time
Bachelor's degree in Systems Engineering, Computer Science, or related field or relevant work experience- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

BXTI, Site Reliability Engineer - Data, Cloud & Developer Experience

Location
Greater London, England, United Kingdom
systems and services to meet the needs of the business. This is achieved through collaboration with the development and engineering teams to leverage SRE practices and principles. You’ll have the opportunity to identify and solve new problems as they arise, deploy and maintain observability systems and pipelines, mature … operational efficiency, and ensure the high quality outputs in all that we do. Key Responsibilities: Provide technical leadership in the understanding and adoption of SRE methodologies across the firm Incorporating observability standards into code and deployment pipelines. Evolving the SRE standards that are adopted across all teams Partnering with colleagues ...

Interim Principal Platform Engineer

Location
Greater London, England, United Kingdom
computer vision to provide the most comprehensive view of surgery. This role is focused on developing and maintaining critical software engineering tools and architecture to drive engineering efficiency, reliability, and continuous improvement. The responsibilities will be diverse, ranging from maintaining AWS resources through Infrastructure as Code … OpenTofu) through to enhancing observability tools (we embrace OpenTelemetry) and supporting AI agentic workflows. You will be working with different software engineering groups to understand their problems and coming up with a pragmatic, maintainable solution to solve their problems. Primary Responsibilities Enhance and maintain our Infrastructure as Code repository ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
+ 3 years in Computer Science, Software Engineering, or a related technical field. Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale. Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation. … Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands‐on experience building LLM pipelines, AI Agents, Model Context ...

Systems Engineer, Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. As a Systems Engineer at Amazon you will utilize … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

AWS Cloud Architect, Technology Consulting- London, Leeds, Manchester or Newcastle

Hiring Organisation
Momentum Worldwide
Location
London, UK
Employment Type
Full-time
security and automation technologies. You will help design scalable, secure and resilient cloud solutions for our clients, enabling digital transformation through modern architecture, platform engineering and DevOps practices. You should bring practical experience of cloud technologies and infrastructure modernisation, alongside an understanding of how AI-enabled tooling and modern … capabilities- Strong understanding of cloud networking, identity, security and resilience patterns- Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) teams- Cloud architecture certifications in additional public cloud platforms (Azure or GCP)- Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

AWS Cloud Architect, Technology Consulting

Location
Greater London, England, United Kingdom
security and automation technologies. You will help design scalable, secure and resilient cloud solutions for our clients, enabling digital transformation through modern architecture, platform engineering and DevOps practices. You should bring practical experience of cloud technologies and infrastructure modernisation, alongside an understanding of how AI-enabled tooling and modern … capabilities Strong understanding of cloud networking, identity, security and resilience patterns Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) teams Cloud architecture certifications in additional public cloud platforms (Azure or GCP) Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

Lead Platform Engineer

Location
Greater London, England, United Kingdom
operation of enterprise platforms supporting major client transformation programmes. Acting as both a technical leader and trusted client advisor, the role combines hands‐on engineering, people leadership and customer engagement. The successful candidate will work directly with client stakeholders, architects and programme teams to ensure platform solutions are secure … self‐service engineering models. Experience leading large‐scale cloud migration and application modernisation programmes. Knowledge of Site Reliability Engineering (SRE) principles and operational excellence practices. Experience with enterprise networking, identity management and Zero Trust architectures. Exposure to data and AI platforms. Relevant industry certifications, such ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
welcome; join us and let's build what's next - together! Role OverviewWe are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused scenario … using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your SkillsHands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure. Proven track record of deploying, scaling, and operationalising machine learning models ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
welcome; join us and let’s build what’s next - together! Role Overview We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused … tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure. Proven track record of deploying, scaling, and operationalising machine learning models ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
other to deliver impact. Make Wayve the experience that defines your career! The role This is a rare opportunity to be a founding Staff SRE shaping the reliability of large-scale AI systems and GPU compute infrastructure from the ground up. As a Staff Cloud Site Reliability … Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won’t inherit a mature SRE function, you’ll help create it. You will define the frameworks, automation, and operational standards that ensure ...

Systems Engineer, Database Services (AWS)

Location
Greater London, England, United Kingdom
supported in the workplace and at home, there’s nothing we can’t achieve. Basic Qualifications Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration Experience working with Linux Experience in any of the following: Python, Java, Perl ...

Principal AWS Cloud Architect, Technology Consulting- London, Leeds, Manchester or Newcastle

Hiring Organisation
Momentum Worldwide
Location
London, UK
Employment Type
Full-time
cloud solutions that accelerate digital transformation and operational excellence. You will help clients realise the full value of cloud adoption through modern architecture, platform engineering, DevOps practices and infrastructure automation, delivering sustainable outcomes in complex and regulated environments. This is a permanent within our Technology Solutions function. At Credera … models- Strong understanding of cloud networking, identity, security and resilience patterns- Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) transformations- Cloud architecture certifications in additional public cloud platforms (Azure or GCP)- Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Site Reliability Engineer Reports to: Labs Team Lead Looper Insights Remote-first, with regular visits to our Byfleet and Hounslow data centres The company Looper Insights builds analytics products that help the world’s leading media and entertainment companies understand how their content is performing across digital platforms … collaborate closely to produce a valuable service for an industry about which we are all passionate. The role We’re looking for a Site Reliability Engineer to keep the global LooperBox fleet running, the physical backbone behind every piece of data Looper Insights produces. LooperBoxes sit in front ...

Senior Site Reliability Engineer (LON)

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (London) We’re working in collaboration to source a Senior Site Reliability Engineer for a large UK client. The role is mostly working remotely, with only 1 day per week being required to work in the London office. In this key role … frequently to other teams, customers and stakeholders The skills you’ll need At least 10 years of hands-on experience, including as a Senior SRE with a proactive approach to spotting problems, areas for improvement, and performance bottlenecks. Experience working with cloud-native microservices, including containerisation, management of Kubernetes workloads ...

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
data for swift root cause identification. Drive post-incident reviews and implement long-term solutions to enhance system resilience.* Collaborate & Influence: Partner with Development, SRE, and Infrastructure leaders to embed observability into the entire technology lifecycle. Influence and drive the adoption of observability best practices across the global organization. Champion … this.**Job Requirements:**Essential Qualifications* Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects.* Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. ...

Military Data Centre Engineering Operations (DCEO), Amazon Web Services (AWS)

Location
Greater London, England, United Kingdom
Military Data Centre Engineering Operations (DCEO), Amazon Web Services (AWS) Job ID: 10509900 | Amazon Data Services UK Limited This role focuses on those who have military experience interested in working in the private sector. Amazon Web Services (AWS) is seeking a Critical Facilities Technician to join our Data Center … Engineering Operations (DCEO) team. In this role you will support the operation, monitoring and maintenance of the electrical and mechanical infrastructure that powers AWS cloud services. AWS data centers operate 24/7 and rely on highly reliable power and cooling systems. As a technician in this environment ...