Senior Lead Site Reliability Engineer
- Location
- Glasgow, Scotland, United Kingdom
level indicators, service level objectives, and error budgets Designs, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem / cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch Drives the assessment, refactoring, and incremental migration … track record in system health monitoring, capacity management, and blameless postmortems for high-availability services Deep understanding of distributed system design principles, networking (TCP / IP, DNS, load balancing), and Linux internals Contributions to open-source observability or telemetry projects Experience working with agent ...