Site Reliability Engineer

Site Reliability Engineer Responsibilities :

  • Proactively identify reliability risks and independently drive initiatives to address them before they become incidents
  • Participate in and continuously improve our on-call rotation, including incident response, triage, and leading blameless post-incident reviews
  • Define and implement monitoring, logging, and distributed tracing strategies; build and maintain dashboards; set meaningful alerts; and drive SLO/SLI/SLA and error budget adoption across services
  • Scope technical projects and break them down into user stories and tasks, driving them to completion with minimal oversight
  • Make sound technical decisions, leveraging input from teammates and contributing to technical conversations across engineering teams
  • Automate the provisioning and management of infrastructure using Infrastructure as Code (IaC) tools such as Terraform

You may be a good fit if :

  • You have at least 4 years of experience working in a professional environment as a Software Engineer (with some SRE or operations responsibilities)
  • You have contributed to the design and build of cloud-native applications written in Python, Java or Go
  • You have strong hands-on experience with observability
  • You understand the difference between monitoring and observability, and can articulate how metrics, logs, and traces work together
  • You have worked with Infrastructure As Code tooling, for example Terraform
  • You have participated in on-call rotations and are comfortable leading incident response under pressure, communicating clearly with stakeholders throughout
  • You are comfortable taking ownership of initiatives or projects independently, from scoping through to delivery
  • You build effective working relationships, give and receive constructive feedback openly, and are trusted by colleagues at all levels

Technologies we use include :

  • Python, Java, and Go are our primary server languages
  • Our browser applications are based on Angular and React –
  • Code lives in GitHub and flows to production through a CI/CD pipeline built on GitHub Actions, with some workloads on Jenkins
  • Infrastructure runs on AWS (EC2) with workloads on Kubernetes-managed Docker containers
  • Datadog is our primary observability platform — experience with Datadog APM, dashboards, monitors, and RUM is a plus - Infrastructure is managed as code using Terraform

Job Details

Company
Infinity Quest
Location
City of London, London, United Kingdom
Posted