26 to 50 of 538 Site Reliability Engineering Jobs in London

Site Reliability Engineer - Fintech / Linux

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Site Reliability Engineer – Fintech/Linux Site Reliability Engineer – Fintech/Linux85,000 Plus BonusQuant Capital is urgently looking for a Site Reliability Engineer to join our high profile client. Our client is a major global financial exchange, driven by technology. They … trading environment for their clients. They have grown massively and recently were voted in the top 50 fintech firms globally. Day to Day the Site Reliability Engineer will: Analyzing and optimizing trading platform Monitoring development activities, change management tickets Monitoring U.S. production, disaster recovery, and certification systems ...

Site Reliability Engineer

Hiring Organisation
SR2 | Socially Responsible Recruitment | Certified B Corporation™
Location
City of London, London, United Kingdom
Site Reliability Engineer (SRE) DevSecOps | Cloud Engineering | Observability | Production Environments | London SR2 is supporting a major 3-year programme and looking for an experienced Site Reliability Engineer (SRE) to join the Production Engineering team. This function underpins the reliability, security, and performance … likely) IR35: Inside Location: London twice a week (hybrid model) Clearance: SC level may be required depending on deployment If you’re an experienced SRE who thrives on building reliable, secure, and cost-efficient production systems, apply now for immediate consideration. ...

Software Engineer III, Site Reliability Engineering, GCE AI

Location
Greater London, England, United Kingdom
Science or Engineering. 2 years of experience designing, analyzing, and troubleshooting large-scale distributed systems. About The Job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both … internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure ...

Site Reliability Engineer, Studios

Location
Greater London, England, United Kingdom
rotations, to support live operations and critical systems. Occasional travel may be required depending on project and client needs. IMG is looking for a Site Reliability Engineer to help design, build, operate, and continuously improve resilient, secure, and highly available platforms that underpin our digital, cloud, and broadcast … adjacent services. This role is suited to someone who combines strong infrastructure and software engineering capability with an operational mindset, and who can help embed reliability engineering practices across systems that support live, business-critical environments. The successful candidate will play a key role in improving service ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
Meta is seeking a Production Engineering Manager to lead a team responsible for the reliability, scalability, and operational excellence of Meta's production infrastructure and services. In this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning … performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure Meta's services operate at global scale with high availability and efficiency.Production Engineering Manager Responsibilities:Manage a team of production ...

Site Relaibility Engineer

Hiring Organisation
Bristow Holland
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£55,000 - £60,000 per annum
exciting global technology organisation is looking for a Site Reliability Engineer (SRE) to join its growing engineering team. This is a fully remote position, offering the opportunity to work on large-scale, business-critical platforms used by customers around the world. The role would suit an experienced … Site Reliability, DevOps, Platform or Cloud Engineer with strong hands-on experience across Kubernetes and Microsoft Azure who enjoys solving complex production problems, improving reliability and automating manual processes. You will work closely with Development and DevOps teams, helping to design, build, operate and scale highly available ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
domains, and advise others on the technical and business issues facing them. You will will set the vision, strategy, and operating model for our SRE transformation - enabling our business-aligned support teams to deliver higher reliability, stronger resilience, and a measurably better end-user experience across the board. … responsibilities Defines the SRE vision, north-star outcomes, and multi-year roadmap for the Production Management team, aligned to both CIB and JPM Global Technology priorities. Establishes the SRE operating model across global regions (ways of working, intake, prioritization, engagement with engineering teams and production support). Partners with ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Location
Greater London, England, United Kingdom
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing … production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
London, UK
Employment Type
Full-time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI.Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities ...

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
team you'll be working with: We are seeking an experienced Site Reliability Engineer (SRE)/Observability Technical Lead to join our team and drive the strategy and execution of observability and reliability projects across our clients. The ideal candidate will have deep expertise in Application Performance … will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and DevOps. What you'll be doing: Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business ...

Senior DevSecOps Engineer

Location
Greater London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Senior DevSecOps Engineer

Location
City Of London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

DevSecOps Engineer

Location
City of Westminster, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting … resume keywords AWS Cloud Platform Architecture Infrastructure as Code CI/CD Automation Platform-as-a-Product Principles Site Reliability Engineering (SRE) Practices ATS Optimization Keywords Hard Skills Cloud Infrastructure Distributed Systems Networking Security Performance Analysis Capacity Planning Scalability Engineering Operational Excellence Monitoring and Alerting Fault ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Software Engineer, Model Deployment- ChatGPT Engineering

Location
Greater London, England, United Kingdom
Software Engineer, Model Deployment- ChatGPT Engineering Applied AI Engineering - London, UK About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond … operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking ...

SRE/Infra | Quant Trading

Location
Greater London, England, United Kingdom
Senior SRE/Platform Engineer | Elite Quant Hedge Fund | London Paragon Alpha is partnering with a leading ~$70BN AUM quantitative trading firm following an exceptional 2025 and continuing to invest aggressively across its global technology organisation. As the business scales, they're expanding a world-class Infrastructure/Site … just engineering metrics – they're a competitive advantage. This is a genuinely high-impact engineering role sitting at the intersection of SRE, Platform Engineering, and Software Engineering , where you'll build the tooling and infrastructure that enables researchers and traders to operate at scale. ...

Site Reliability Manager - Environment Strategy

Location
Greater London, England, United Kingdom
looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving … while ensuring strong technical standards, compliance and resilience across all production and non-production systems. Responsibilities Define and own the vision and roadmap for site reliability and environment strategy Lead, mentor and develop a team of DevOps and environment engineers Set and enforce standards for environment provisioning, lifecycle ...

Platform Engineer

Location
Greater London, England, United Kingdom
workflows, eliminate toil, and promote consistent best practices, so software engineers can focus on high-value customer delivery. Support reliability and sacabibility through SRE principles and proactive operations. In partnership with product and engineering teams, you'll participate in incident response, help define and monitor SLI/… teams building a customer-focused SaaS product, weighing delivering value to clients against technical excellence. 4+ years of experience in Platform Engineering, SRE, or DevOps roles. Good working knowledge of cloud platforms (AWS and/or GCP), managed services and infrastructure automation tools (Terraform or similar). Proficiency ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
Cost Explorer, Compute Optimizer, and EKS workload right‐sizing. Skills: 3+ years of hands‐on experience in DevOps, Site Reliability Engineering (SRE), infrastructure engineering, or closely related roles (focused on AWS/EKS environments). Proven track record building and maintaining production‐grade CI/… incidents/post‐mortems into actionable learning for personal and team growth. Proven ability to collaborate effectively with cross‐functional stakeholders, including engineering, SRE, security, and product teams. Maintain clear technical documentation, runbooks, and operational playbooks for production systems. We Offer: Experience a dynamic and team‐orientated work environment ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
TITLE: Site Reliability Engineer SALARY: Up to £80,780 (Outside London), £93,390 (London) LOCATION: Remote (UK Based) HOURS: Full Time (35 Hours) WORKING PATTERN: Our work style is hybrid, which involves spending at least two days per week, or 40% of our time … with workplace adjustments including hybrid working expectations in line with our Flexibility Works policy. What you\'ll be doing We\'re looking for a Site Reliability Engineer to help build, scale and operate the platforms that power Curve\'s products and services. Working closely with engineering teams ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
london, south east england, united kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, youhold a leadership role in your team, demonstrate … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team's applications ...

Site Reliability Engineer , Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. We are an operations team, but a key focus … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Software Engineer Lead - Site Reliability

Location
City of Westminster, England, United Kingdom
strategic partners and third parties., As a Lead DevOps Engineer, you will be a hands-on contributor and technical lead focused on improving the engineering foundations of the Customer Digital Platform. You will work across internal squads and with strategic partners, third parties, service, architecture and security teams … GitHub/GitHub Actions, Azure DevOps, Terraform and automated testing. Improve deployment safety, release readiness and operational readiness for customer-facing digital services. Apply SRE principles pragmatically to improve availability, recoverability, monitoring and incident learning. Strengthen monitoring, logging, tracing, alerting and service-health dashboards across digitally connected workloads. Reduce single ...