Embed observability-led operations through monitoring, logging, tracing, alerting, dashboards and actionable runbooks across critical services Improve incident prevention and recovery through better alerting, runbook automation, expedited L2/L3 escalation and reduced MTTR and disruptive minutes Govern operational resilience activity including IBS/TBSL governance, TRCB engagement, disaster recovery … service resilience and reliability outcomes including MTTR reduction and service stability improvement Apply strong SRE practices including SLIs, SLOs, reliability targets, actionable alerting and runbook design Use observability tooling and practices including monitoring, logging, tracing, dashboards and customer journey visibility Show engineering experience that embeds build-time resilience, testing ...