Lead Site Reliability Engineer
- Location
- Southampton, England, United Kingdom
Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor. Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc. Developing bicep modules for monitoring infrastructure and deploy it. Optimize system performance, cost, and security through regular reviews and tuning. … etc.) Experience with infrastructure/configuration as code and version control (ARM, BICEP, Git) Strong Experience managing monitoring, alerting and dashboarding platforms (Azure Monitor, Prometheus, Grafana, Elasticsearch) Demonstrable experience of supporting live cloud services and platforms Expert in developing queries for dashboards and alerting for microservices. Collaborate with DevOps ...