26 to 50 of 357 Site Reliability Engineering Jobs in London

Senior Software Engineer - Permanent - London/Hybrid - £70,000 - 85,000

Hiring Organisation
Robson Bale Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 85,000 Annual
efficiency and reduce MTTR. Use Datadog, Kibana and Heap to investigate issues, understand customer impact and improve service health. Partner with Software Engineering, SRE and Platform teams to improve observability, reliability and operational readiness. Contribute to service onboarding, post-incident reviews and continuous improvement initiatives. What … Looking For Experience in Application Operations, Production Engineering, SRE or a similar operational engineering environment. Experience supporting cloud-hosted applications, ideally PHP services running on Kubernetes (EKS) with MySQL databases. Strong troubleshooting and diagnostic skills, with the ability to understand systems end-to-end. Hands-on experience with ...

Software Engineer III, Site Reliability Engineering, Traffic Network Load Balancing

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Science or Engineering. 2 years of experience designing, analyzing, and troubleshooting large-scale distributed systems. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault‐tolerant systems. SRE ensures that the company Cloud's services … internally critical and our externally‐visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever‐watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure ...

SRE Managing Consultant

Hiring Organisation
Akkodis
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £100000/annum
SRE Managing Consultant Cloud Operating Model & Reliability Transformation Security Clearance: SC eligible (UK residency required) Shape the Future of Cloud Reliability Are you passionate about building resilient, scalable cloud platforms that truly support the business? Do you thrive at the intersection of engineering excellence, operating models … senior stakeholder advisory? We're looking for a Managing Consultant in Site Reliability Engineering (SRE) to help organisations shift from reactive operations to measurable, product-aligned reliability - embedding SRE as a core engineering discipline across cloud and hybrid environments. You'll work with senior leaders ...

Site Reliability Engineer - SRE Fleet

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and efficiency of the infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US, Canada, and the UK, we combine deep infrastructure expertise with a strong focus on automation … reliability, and operational excellence. We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide. The team operates with a high degree of autonomy, giving engineers the opportunity to drive both critical initiatives ...

Site Reliability Engineer- London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
extend and will be a hybrid role that will be based in London. Our client is seeking an experienced Site Reliability Engineer (SRE) with a strong focus on Observability and Monitoring Platforms. The successful candidate will play a key role in enhancing the organisation's monitoring, alerting … streamline operational processes and improve reliability. Collaborate with engineering, infrastructure, and support teams to improve system resilience and operational performance. Define and implement SRE best practices, including monitoring standards, alert management, incident response, and operational readiness. Perform troubleshooting and root cause analysis of platform and application issues. Support capacity ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing … production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems ...

Software & Data Engineers

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
native, low‐latency system that operates at global scale and underpins critical investment products used worldwide. You will work on challenging problems across software engineering where performance, data quality, and reliability are non‐negotiable, leveraging modern cloud (AWS) and AI‐assisted development tooling to accelerate delivery without compromising … encouraged to apply. Areas We Value Experience In Software Engineering Data Engineering Cloud & Platform Engineering Site Reliability Engineering (SRE) AI & Machine Learning DevOps & Automation Architecture & Distributed Systems Analytics & Data Platforms Benefits LSEG offers a range of tailored benefits and support, including healthcare, retirement planning ...

Software & Data Engineers

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
underpins critical investment products used worldwide.We are recruiting for individual contributors across a varied technical tech. You'll work on challenging problems across software engineering where performance, data quality, and reliability are non-negotiable, while leveraging modern cloud (AWS) and AI-assisted development tooling to accelerate delivery without … more of the following areas:* Software Engineering* Data Engineering* Cloud & Platform Engineering* Site Reliability Engineering (SRE)* AI & Machine Learning* DevOps & Automation* Architecture & Distributed Systems* Analytics & Data PlatformsBring your curiosity, expertise, and ambition—and help build what's next for global investment and market intelligence. ...

Lead Product Manager AIOPs

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
Greater London, United Kingdom
Employment Type
Full Time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed … Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self ...

Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office)

Hiring Organisation
Hamilton Barnes
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP 450 - 475 Daily
enabling proactive monitoring, faster incident resolution, and improved platform reliability through modern observability practices - with deep expertise in Grafana, OpenTelemetry, distributed tracing, SRE, event-driven architecture, and Azure Integration Services. Key Responsibilities Define and implement enterprise observability strategies, standards, and governance frameworks Design and manage observability solutions covering metrics … service health across Azure services Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience Drive root cause analysis, incident investigations, and continuous service improvement initiatives Champion operational excellence through proactive monitoring, automation ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Site Reliability Engineer

Hiring Organisation
Randstad Digital
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
Site Reliability Engineer (SRE) - 100% Remote Location: Fully Remote Duration: Permanent Are you passionate about building unbreakable systems and automating away the noise? We are looking for a dedicated Site Reliability Engineer (SRE) to join our remote team. Your primary mission will be to design, implement … complex challenges in the Azure ecosystem and sharing your knowledge with others, we want you on our team! What You Will Do As an SRE, you will be accountable for the delivery and support of production and non-production systems within the Azure ecosystem. Your day-to-day responsibilities will ...

Site Reliability Engineer (SRE)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
become available for a Site Reliability Engineer to join our team to help us transform our existing operational workloads to an SRE approach. Key Responsibilities Integrating tightly with our Product Engineering teams Following SRE practices and maintaining high standards of compliance Implementing a new standard of observability … part in the daily stand-ups and keeping sprints on track Keeping up-to-date documentation in the JIRA & Confluence tools Taking part in SRE Incident Management processes Acting as a key Incident Commander within the Incident Management process Taking part in SRE On Call Ensuring a focus on cost ...

Site Reliability Engineer - NS London

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission. The role holder: As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

Site Reliability Engineer – NS London

Hiring Organisation
BAE Systems
Location
Greater London, United Kingdom
Employment Type
Full Time
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission. The role holder: As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‐healing workflows, and resilience engineering. Use enterprise‐authorized AI capabilities within the work environment to accelerate major‐incident triage ...

Site Reliability Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About the Role Nottingham Trent House (95002), United Kingdom, Nottingham, Nottinghamshire Site Reliability Engineering Manager This role requires a proven leader to develop technical staff, drive service excellence, and implement significant reliability improvements within complex, large‐scale, highly regulated systems. What You’ll Do Lead … cross‐functional group of software engineers focused on managing and optimizing applications to maintain and improve reliability for our customers. Coach and nurture engineers to attain their technical, business, and personal goals. Collaborate with Senior Software Engineering managers to deliver improvements aligned with the technical roadmap and customer ...

Senior AWS Platform Engineer

Hiring Organisation
ReVybe IT Recruitment Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
£85000 - £90000/annum
standards, and take architectural ownership of the platform What We're Looking For Strong background in Platform Engineering, DevOps, Cloud Engineering, or SRE Deep hands on AWS experience, including networking, IAM, and security Commercial experience running Kubernetes in production, ideally EKS Expert level Terraform and Infrastructure as Code … London (Hybrid, 2 days per week) Up to £90,000 + Bonus + Benefits AWS | EKS | Kubernetes | Terraform | RDS | Linux | GitHub Actions | Python | Go | SRE | CI/CD | Infrastructure as Code | Platform Engineering ...

Cloud Operations Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Engineer****Overview** At Mimecast, we operate at scale—protecting billions of emails daily across a global hybrid-cloud infrastructure. If you’re passionate about reliability, automation, and solving complex challenges in mission-critical environments, this is the opportunity for you.**Why Join Us** “This is a hands-on role … performing out-of-hours maintenance as necessary.**What You’ll Bring:*** **Technical expertise**: Experience in roles such as Site Reliability Engineer (SRE), Platform Engineer, Cloud Operations, or Sys Admin.* **Kubernetes proficiency**: Strong hands-on experience with Kubernetes (k8s) and containerization in production environments.* **Linux expertise**: Proficiency in maintaining ...

Project Manager (DV Security Clearance)

Hiring Organisation
CGI
Location
Greater London, United Kingdom
Employment Type
Full Time
team, you will help create the conditions for successful delivery while supporting innovation, accountability, and operational excellence. Key responsibilities: ~Lead & Coordinate delivery across multiple SRE and product teams ~Develop & Maintain project plans, roadmaps, milestones, and reporting artefacts ~Facilitate & Drive agile ceremonies including sprint planning, reviews, retrospectives, and backlog refinement ~Manage … communication, presentation, and stakeholder management skills ~Strong organisational skills with the ability to drive accountability and delivery across multiple workstreams Desirable experience: ~Experience supporting SRE, DevOps, Cloud, Platform Engineering, or Infrastructure teams ~Knowledge of IT Service Management and operational delivery practices ~Experience managing project budgets, forecasting, and resource planning ...

Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£49,734 - £57,176 per annum, Pro-rata, Inc benefits
economy! The Department for Business, Innovation, Science and Trade ("BIST") and Inspire People are partnering together to bring you an exciting opportunity for a Site Reliability Engineer to join a team that ensures BIST's digital services work as users expect, working with development teams giving them … dependent on location and technical skills as assessed at interview. Flexible, hybrid working from London, Cardiff, Darlington. Birmingham, Salford, Edinburgh or Belfast. As a Site Reliability Engineer, you will pro-actively engage development teams and use initiative to develop the tools for their job, including application performance monitoring ...

Principal Platform Engineer (SRE/Cloud)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
honesty ensuring our workforce is able to bring their full selves to work. ABOUT THE ROLE Principal Platform Engineers at Beamery solve the toughest reliability, scalability and infrastructure problems with the highest impact. Together they collaborate to set the standards for how Engineering will build, run and operate … engineering organisation WHO ARE WE LOOKING FOR? We are seeking a hands-on technical leader with deep Site Reliability Engineering (SRE) and Cloud expertise who can set direction across the engineering organisation. Key skills/experience: A proven track record of designing and delivering scalable ...