Principal Site Reliability Engineer, Infrastructure Observability
- Location
- Greater London, England, United Kingdom
with implementation and operation of the chaos model at scale Strategic and program-level implementation experience Demonstrable experience implementing new technology, tools, and platforms System administration and scripting experience Demonstrable experience leveraging automation to proactively prevent or quickly remediate incidents Fluent in multiple programming languages (e.g., Python, Java … database development (SQL Server, PostgreSQL, MySQL, etc) Proficiency with defining, right-sizing, tracking, and reporting on Service Level Objectives (SLOs), Service Level Indicators (SLIs), system availability, and the progress and outcomes related to reliability Experience with implementing and managing Error Budgets Proficiency with understanding and explaining incident situations ...