76 to 100 of 796 Site Reliability Engineering Jobs in England

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT
Location
London, United Kingdom
Salary
£ 80 K
team you'll be working with: We are seeking an experienced Site Reliability Engineer (SRE)/Observability Technical Lead to join our team and drive the strategy and execution of observability and reliability projects across our clients. The ideal candidate will have deep expertise in Application Performance … will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and DevOps. What you'll be doing: Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
team you'll be working with: We are seeking an experienced Site Reliability Engineer (SRE)/Observability Technical Lead to join our team and drive the strategy and execution of observability and reliability projects across our clients. The ideal candidate will have deep expertise in Application Performance … will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and DevOps. What you'll be doing: Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business ...

Salesforce Engineering Lead

Location
Greater London, England, United Kingdom
Ready to take your career global? Make your mark at one of the biggest names in payments. We’re looking for a Salesforce Engineering Lead to join and lead the strategic direction, delivery, and operational excellence of our enterprise Salesforce ecosystem across Enterprise (LOB). This is a senior … leadership role responsible for building and leading a high-performing engineering organisation that delivers secure, scalable, and innovative solutions supporting the end-to-end customer lifecycle. As a member of the technology leadership team, you will be accountable for engineering strategy, organisational capability, product delivery, workforce planning, budget ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting … resume keywords AWS Cloud Platform Architecture Infrastructure as Code CI/CD Automation Platform-as-a-Product Principles Site Reliability Engineering (SRE) Practices ATS Optimization Keywords Hard Skills Cloud Infrastructure Distributed Systems Networking Security Performance Analysis Capacity Planning Scalability Engineering Operational Excellence Monitoring and Alerting Fault ...

SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Location
Greater London, England, United Kingdom
balance reliability with feature velocity Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement Collaborate with Platform Engineering and Cloud teams to embed … Strong analytical and problem-solving mindset with attention to detail Ability to manage competing priorities across multiple workstreams simultaneously QUALIFICATIONS & EXPERIENCE 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity Proven track record of improving availability, reducing MTTR, and implementing self ...

Site Reliability Engineer - Java - Fintech

Hiring Organisation
Rothstein Recruitment Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 85,000 Annual
Site Reliability Engineer - Java - Fintech Excellent opportunity opens to join a High-Growth Fintech as their new Site Reliability Engineer. With a focus on its flagship platform, you will take ownership of the platform's B2B, B2C, and SaaS production environments. You will also lead … areas for improvement in stability and efficiency using tools like Datadog, Rootly, and CloudWatch/AppDynamics. Qualifications & Experience Bachelor's degree in computer science, Engineering, or a related field. Minimum of 5+ years of experience in sustaining engineering, DevOps, or software engineering with a focus on incident ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer (SRE) - GCP/Kubernetes We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Senior Site Reliability Engineer

Hiring Organisation
Brevan Howard
Location
London, UK
Employment Type
Full-time
Senior Site Reliability Engineer (SRE) - GCP/KubernetesAbout the RoleWe are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Senior Site Reliability Engineer

Hiring Organisation
Brevan Howard
Location
London, United Kingdom
Salary
£ 80 K
Senior Site Reliability Engineer (SRE) - GCP/KubernetesAbout the RoleWe are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance ...

Oracle Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, United Kingdom
Salary
£ 60 K
DescriptionPurpose of the roleTo apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them. AccountabilitiesAvailability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.Resolution, analysis and response … incident management, and production troubleshooting.Knowledge of Oracle RAC, Data Guard, GoldenGate, ASM, and disaster recovery architectures.Exposure to cloud platforms (OCI, AWS, Azure) and modern SRE practices such as reliability engineering, capacity planning, and service resilience.You may be assessed on the key critical skills relevant for success in role ...

Oracle Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, UK
Employment Type
Full-time
DescriptionPurpose of the roleTo apply software engineering techniques, automation, and best practices in incident response, to ensure the reliability, availability, and scalability of the systems, platforms, and technology through them. AccountabilitiesAvailability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning. Resolution, analysis … production troubleshooting. Knowledge of Oracle RAC, Data Guard, GoldenGate, ASM, and disaster recovery architectures. Exposure to cloud platforms (OCI, AWS, Azure) and modern SRE practices such as reliability engineering, capacity planning, and service resilience. You may be assessed on the key critical skills relevant for success in role ...

SRE Architect (68019)

Hiring Organisation
Hitachi
Location
London, United Kingdom
Salary
£ 80 K
balance reliability with feature velocity• Conduct chaos engineering exercises and game days to validate resiliency and uncover hidden failure modes• Mentor 2 SRE Engineers, establish engineering standards, and build a culture of reliability and continuous improvement• Collaborate with Platform Engineering and Cloud teams to embed … consensus• Strong analytical and problem-solving mindset with attention to detail• Ability to manage competing priorities across multiple workstreams simultaneouslyQUALIFICATIONS & EXPERIENCE• 7+ years in SRE, DevOps, or production engineering with 3+ years in a senior or lead capacity• Proven track record of improving availability, reducing MTTR, and implementing self ...

Site Reliability Engineer

Hiring Organisation
Trainline
Location
London, United Kingdom
Salary
£ 70 K
millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong DevOps and SRE practices.The Reliability & Operations Engineering team (ReliabilityOps) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. … willingness to challenge and be challenged — contributing to platform reliability while developing broader technical ownership with support from senior engineers.As an SRE at Trainline, you'll be working on...Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platformParticipating in production incident response, supporting investigation, mitigation ...

Site Reliability Engineer , Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, United Kingdom
Salary
£ 80 K
high availability environment, building and operating critical Cryptography, Access and Identity services for our customers.This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. We are an operations team, but a key focus … supported in the workplace and at home, there’s nothing we can’t achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Security Engineer (Site Reliability Engineering)

Hiring Organisation
Sanderson Government and Defence
Location
London, United Kingdom
Employment Type
Permanent
Salary
£545 - £590 per day + Inside IR35
Length: 6-18 Months Rate: £545-£590 per day (Inside IR35) Positions Available: 2 About the Role We are seeking two experienced Security Engineers (SRE) to join a specialist consultancy delivering cyber security services across a portfolio of government projects and digital transformation programmes. This role is ideal for security … best practices. Drive continual improvements across security engineering and platform security functions. Essential Skills & Experience Strong experience as a Security Engineer, Security-focused SRE, Platform Security Engineer, or similar role. Advanced knowledge of Enterprise Security Architecture principles. Hands-on experience withSplunk, including: Security monitoring Dashboard development Alerting and reporting ...

Lead Site Reliability Engineer

Hiring Organisation
JP Morgan Chase
Location
London, United Kingdom
Salary
£ 80 K
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … directly to the codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation.Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major ...

Strategic DevSecOps Consultant

Hiring Organisation
CloudBees
Location
London, United Kingdom
Salary
£ 80 K
transform how software is built, secured, and delivered. As a member of the Services team, you will work at the intersection of DevSecOps, Platform Engineering, AI, and software delivery innovation, helping customers accelerate outcomes and unlock new levels of engineering productivity.You will collaborate with some of the world … customer obsession.What You'll BringRequired:5+ years of experience in consulting, solutions architecture, platform engineering, DevOps, Site Reliability Engineering (SRE), or related customer-facing technical roles.Proven experience advising customers on DevSecOps, software delivery modernization, platform engineering, or cloud transformation initiatives.Strong understanding of modern software delivery ...

Site Reliability Engineer - NS London

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission. The role holder: As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‐healing workflows, and resilience engineering. Use enterprise‐authorized AI capabilities within the work environment to accelerate major‐incident triage ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Lead Site Reliability Engineer

Location
Westminster, West End, United Kingdom
trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Engineering Lead

Hiring Organisation
Moorepay
Location
Manchester, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Engineering Lead is a hands-on technical leader responsible for guiding a cross-functional engineering squad in the delivery of high-quality, secure, reliable, and scalable software. Combining the responsibilities of a senior engineer with those of a team enabler, the role promotes strong engineering practices, effective … Owner. Cross-Functional Collaboration: Partner with the Solutions Architect to ensure architectural consistency and alignment with strategic design patterns. Work closely with the Lead SRE to embed reliability, observability, and CI/CD into day-to-day engineering workflows. Collaborate with Product Owners to refine requirements, manage scope ...

Infrastructure Engineer-Hyper-V

Hiring Organisation
NEEV LIMITED
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
From £400 to £450 per day 400 - 450 GBP/day InsideIR35
Required Technical Skills Site Reliability Engineering Strong understanding of Site Reliability Engineering principles and operational excellence. Experience with infrastructure reliability, service availability, resiliency, and performance optimization. Storage Space Direct a nd failover clustering technical expertise . (Storage Spaces Direct enables you to build … within Hyper-V environments including virtual switches, VLANs, NIC Teaming, QoS, and network performance tuning. System Center Virtual Machine Manager ( SCVMM ). Automation & Platform Engineering Strong PowerShell scripting and automation experience. Experience automating infrastructure deployment, operational tasks, health checks, and reporting. Familiarity with Infrastructure as Code concepts and configuration ...