Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office)

Observability SME | 1 year | London, UK (Hybrid - 3 days/week in office)

Role Overview

We're recruiting for an experienced Observability SME to define, implement, and govern enterprise-wide observability capabilities for cloud-native and distributed applications running on Microsoft Azure, for a leading organisation. The role requires establishing observability standards, ensuring end-to-end visibility across business-critical platforms, and enabling proactive monitoring, faster incident resolution, and improved platform reliability through modern observability practices - with deep expertise in Grafana, OpenTelemetry, distributed tracing, SRE, event-driven architecture, and Azure Integration Services.

Key Responsibilities

  • Define and implement enterprise observability strategies, standards, and governance frameworks
  • Design and manage observability solutions covering metrics, logs, traces, and application telemetry
  • Establish monitoring, alerting, and diagnostics best practices across cloud-native platforms and microservices
  • Design and implement distributed tracing solutions using OpenTelemetry and modern observability tools
  • Develop and maintain Grafana dashboards for engineering, operations, business, and leadership stakeholders
  • Monitor platform performance, availability, reliability, and service health across Azure services
  • Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs
  • Collaborate with development, platform engineering, and SRE teams to improve system observability and resilience
  • Drive root cause analysis, incident investigations, and continuous service improvement initiatives
  • Champion operational excellence through proactive monitoring, automation, and reliability engineering practices

What You Will Ideally Bring

  • Strong experience in enterprise observability, monitoring, and operational support
  • Expertise in Grafana dashboard development and observability platform management
  • Hands-on experience with OpenTelemetry, distributed tracing, and telemetry frameworks
  • Strong understanding of Site Reliability Engineering (SRE) principles and practices
  • Experience monitoring cloud-native applications, microservices, and distributed systems
  • Knowledge of Azure monitoring services, diagnostics, and observability ecosystems (Azure Monitor, Log Analytics, Event Hub, Service Bus, Functions)
  • Experience implementing alerting strategies, incident management, and performance monitoring solutions
  • Strong understanding of Event-Driven Architecture and asynchronous application behaviour
  • Experience defining and measuring performance, reliability, and availability metrics
  • Excellent problem-solving, analytical, stakeholder management, and communication skills
  • Cosmos DB and/or PostgreSQL experience (desirable)

Contract Details

  • Duration: Initial 12 months
  • Location: 3 days onsite in London
  • Rate: £450-475/day (inside)

Job Details

Company
Hamilton Barnes
Location
London, United Kingdom
Hybrid / Remote Options
Employment Type
Contract
Salary
GBP 450 - 475 Daily
Posted