3 of 3 Site Reliability Engineer Jobs in Berkshire

Site Reliability Engineer

Location
Slough, England, United Kingdom
Site Reliability Engineer (SRE) DevSecOps | Cloud Engineering | Observability | Production Environments | London SR2 is supporting a major 3-year programme and looking for an experienced Site Reliability Engineer (SRE) to join the Production Engineering team. This function underpins the reliability, security, and performance … likely) IR35: Inside Location: London twice a week (hybrid model) Clearance: SC level may be required depending on deployment If you’re an experienced SRE who thrives on building reliable, secure, and cost-efficient production systems. #J-18808-Ljbffr ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security. The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers … platform in close partnership with our lead/backend/staff engineers. Who you are (must-haves) 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm ...

Platform Engineering Manager (SRE)

Location
Bracknell, England, United Kingdom
cloud products fast, reliable, and trusted by some of the world's largest SAP-run businesses. We are hiring a Platform Engineering Manager (SRE) to own reliability and platform engineering for Klario and our Private Cloud estate. This is a hands-on leadership role and a genuine build. … management, observability, and developer tooling. Drive standardisation and reduce engineering friction through automation and self-service. Site Reliability Engineering Introduce and embed SRE practices across Engineering. Improve reliability, resilience, recoverability, and operational readiness. Stand up monitoring, alerting, logging, and service health capabilities, including public-facing uptime ...