26 to 50 of 262 Site Reliability Engineering Jobs in London

Site Reliability Engineer - SRE Fleet

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and efficiency of the infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US, Canada, and the UK, we combine deep infrastructure expertise with a strong focus on automation … reliability, and operational excellence. We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide. The team operates with a high degree of autonomy, giving engineers the opportunity to drive both critical initiatives ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing … production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems ...

Software & Data Engineers

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
native, low‐latency system that operates at global scale and underpins critical investment products used worldwide. You will work on challenging problems across software engineering where performance, data quality, and reliability are non‐negotiable, leveraging modern cloud (AWS) and AI‐assisted development tooling to accelerate delivery without compromising … encouraged to apply. Areas We Value Experience In Software Engineering Data Engineering Cloud & Platform Engineering Site Reliability Engineering (SRE) AI & Machine Learning DevOps & Automation Architecture & Distributed Systems Analytics & Data Platforms Benefits LSEG offers a range of tailored benefits and support, including healthcare, retirement planning ...

Software & Data Engineers

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
underpins critical investment products used worldwide.We are recruiting for individual contributors across a varied technical tech. You'll work on challenging problems across software engineering where performance, data quality, and reliability are non-negotiable, while leveraging modern cloud (AWS) and AI-assisted development tooling to accelerate delivery without … more of the following areas:* Software Engineering* Data Engineering* Cloud & Platform Engineering* Site Reliability Engineering (SRE)* AI & Machine Learning* DevOps & Automation* Architecture & Distributed Systems* Analytics & Data PlatformsBring your curiosity, expertise, and ambition—and help build what's next for global investment and market intelligence. ...

Lead Product Manager AIOPs

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
Greater London, United Kingdom
Employment Type
Full Time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed … Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Senior Site Reliability Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Role Overview Sr. Manager, Site Reliability Engineering (London) is an experienced leader responsible for overseeing a globally distributed team of SRE technologists with diverse skills in software development, systems, network, application, and/or database management. This role ensures seamless, continuous coverage of Cboe's real‐time … features; monitor development activities, change‐management tickets, evaluate impact; approve and execute daily change tickets; organize testing prior to deployment; work with software engineering to resolve systemic issues; ensure compliance obligations are met. Incident Response & Escalation Management: Serve as senior escalation point for production incidents across European ...

Site Reliability Engineer

Hiring Organisation
Randstad Digital
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
Site Reliability Engineer (SRE) - 100% Remote Location: Fully Remote Duration: Permanent Are you passionate about building unbreakable systems and automating away the noise? We are looking for a dedicated Site Reliability Engineer (SRE) to join our remote team. Your primary mission will be to design, implement … complex challenges in the Azure ecosystem and sharing your knowledge with others, we want you on our team! What You Will Do As an SRE, you will be accountable for the delivery and support of production and non-production systems within the Azure ecosystem. Your day-to-day responsibilities will ...

TAM Manager, Google Cloud Consulting

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Master’s degree in a Management, Technical, or Engineering field. Experience managing teams focused on cloud operations, Site Reliability Engineering (SRE), or technical account management. Experience collaborating with cross-functional teams (e.g., sales, support, product) to drive customer outcomes. Understanding of cloud computing concepts (e.g., networking … Identity and Access Management (IAM), compute, storage) and operational frameworks (e.g., Information Technology Infrastructure Library (ITIL), Developer Operations (DevOps), SRE). Ability to attract and develop talent, with a track record of coaching team members to improve performance. Excellent communication skills and the ability to manage multiple priorities ...

Site Reliability Engineer (SRE)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
become available for a Site Reliability Engineer to join our team to help us transform our existing operational workloads to an SRE approach. Key Responsibilities Integrating tightly with our Product Engineering teams Following SRE practices and maintaining high standards of compliance Implementing a new standard of observability … part in the daily stand-ups and keeping sprints on track Keeping up-to-date documentation in the JIRA & Confluence tools Taking part in SRE Incident Management processes Acting as a key Incident Commander within the Incident Management process Taking part in SRE On Call Ensuring a focus on cost ...

Site Reliability Engineer - NS London

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission. The role holder: As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

Site Reliability Engineer – NS London

Hiring Organisation
BAE Systems
Location
Greater London, United Kingdom
Employment Type
Full Time
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission. The role holder: As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‐healing workflows, and resilience engineering. Use enterprise‐authorized AI capabilities within the work environment to accelerate major‐incident triage ...

Site Reliability Engineer

Hiring Organisation
Randstad Digital
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 60,000 Annual
Site Reliability Engineer (SRE) - 100% Remote Location: Fully Remote Duration: Permanent Are you passionate about building unbreakable systems and automating away the noise? We are looking for a dedicated Site Reliability Engineer (SRE) to join our remote team click apply for full job details ...

Senior AWS Platform Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
standards, and take architectural ownership of the platform What We're Looking For Strong background in Platform Engineering, DevOps, Cloud Engineering, or SRE Deep hands on AWS experience, including networking, IAM, and security Commercial experience running Kubernetes in production, ideally EKS Expert level Terraform and Infrastructure as Code … Platform Engineer London (Hybrid, 2 days per week) Up to £90,000 + Benefits AWS | EKS | Kubernetes | Terraform | RDS | Linux | GitHub Actions | Python | Go | SRE | CI/CD | Infrastructure as Code | Platform Engineering ...

Project Manager (DV Security Clearance)

Hiring Organisation
CGI
Location
Greater London, United Kingdom
Employment Type
Full Time
team, you will help create the conditions for successful delivery while supporting innovation, accountability, and operational excellence. Key responsibilities: ~Lead & Coordinate delivery across multiple SRE and product teams ~Develop & Maintain project plans, roadmaps, milestones, and reporting artefacts ~Facilitate & Drive agile ceremonies including sprint planning, reviews, retrospectives, and backlog refinement ~Manage … communication, presentation, and stakeholder management skills ~Strong organisational skills with the ability to drive accountability and delivery across multiple workstreams Desirable experience: ~Experience supporting SRE, DevOps, Cloud, Platform Engineering, or Infrastructure teams ~Knowledge of IT Service Management and operational delivery practices ~Experience managing project budgets, forecasting, and resource planning ...

IT Recruiter

Hiring Organisation
Cpl Life Sciences
Location
City of London, London, United Kingdom
business leaders to develop sourcing strategies aligned with workforce plans and hiring priorities. Build pipelines of diverse technology talent across multiple domains including: Cloud Engineering Platform Engineering & Operations Software Engineering Data Engineering & Analytics Cybersecurity AI/Machine Learning/Generative AI Enterprise Architecture Infrastructure Engineering DevOps & Site Reliability Engineering Technology Risk & Governance Technology Modernization & Digital Transformation Develop proactive sourcing strategies utilizing LinkedIn Recruiter, talent mapping, networking, referrals, and market intelligence. Evaluate candidates for both technical capability and organizational fit through structured screening and assessment processes. Consult with hiring managers regarding market ...

Principal Platform Engineer (SRE/Cloud)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
honesty ensuring our workforce is able to bring their full selves to work. ABOUT THE ROLE Principal Platform Engineers at Beamery solve the toughest reliability, scalability and infrastructure problems with the highest impact. Together they collaborate to set the standards for how Engineering will build, run and operate … engineering organisation WHO ARE WE LOOKING FOR? We are seeking a hands-on technical leader with deep Site Reliability Engineering (SRE) and Cloud expertise who can set direction across the engineering organisation. Key skills/experience: A proven track record of designing and delivering scalable ...

DevSecOps Engineer

Hiring Organisation
167 Solutions Ltd
Location
North West London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
+ Benefits Type: Permanent About the Company 167 Solutions is partnering with an innovative technology organisation that is scaling its cloud engineering and platform capabilities. We are seeking a hands-on DevSecOps Engineer who can embed security directly into the software development lifecycle while remaining actively involved in engineering … will be responsible for automating security controls, improving cloud security posture, developing CI/CD pipelines, and implementing security tooling within a fast-paced engineering environment. The successful candidate will have strong software engineering capabilities alongside cloud and security expertise. Key Responsibilities Design, build and maintain secure ...

DevSecOps Engineer

Hiring Organisation
167 Solutions Ltd
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£40,000 - £70,000 per annum
+ Benefits Type: Permanent About the Company 167 Solutions is partnering with an innovative technology organisation that is scaling its cloud engineering and platform capabilities. We are seeking a hands-on DevSecOps Engineer who can embed security directly into the software development life cycle while remaining actively involved … engineering, automation, cloud infrastructure, and platform delivery. This is not a traditional security administration or governance role. We are looking for an engineer who writes code, builds automation, develops cloud-native solutions, and integrates security into modern software delivery practices. The Opportunity As a DevSecOps Engineer, you will work ...

IAM Secrets Management Engineering - SRE Platform Engineer - VP - London

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
Secrets Management Engineering - SRE Platform Engineer - VP - London Job Description The Role We are seeking a skilled Lead Site Reliability Platform Engineer (SRE) to join our team. The ideal candidate will be responsible for ensuring the reliability, performance, and scalability of mission‐critical, high‐availability, high … throughput systems and infrastructure. This role involves leading collaboration with cross‐functional teams, and implementing best practices in SRE, DevOps, and cyber security to enhance our operational efficiency and security posture. System Reliability, Performance, and Security: Design, implement, and maintain highly available, scalable, and secure systems on AWS cloud ...

Principal Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
help us transform the future of work once and for all. ABOUT THE ROLE Principal Platform Engineers at Beamery solve the toughest reliability, scalability and infrastructure problems with the highest impact. Together they collaborate to set the standards for how Engineering will build, run and operate services … engineering organisation. WHO ARE WE LOOKING FOR? We are seeking a hands-on technical leader with deep Site Reliability Engineering (SRE) and Cloud expertise who can set direction across the engineering organisation. Key skills/experience: A proven track record of designing and delivering scalable ...