Principal Platform Engineer - London
- Location
- Greater London, England, United Kingdom
take increasing responsibility for leading incidents end‐to‐end. Improve operational reliability: Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements. Own parts of the production environment: Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing … Kubernetes and containerised workloads. Infrastructure as Code experience (Terraform or similar). Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc). Scripting or automation experience (Python, Bash, or similar). Nice to have: Experience leading incidents or mentoring others during on‐call. Experience in regulated or security‐sensitive ...