champions blameless postmortem cultureCollaborates with team members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgetsDesigns, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB … respectRequired qualifications, capabilities, and skillsFormal training or certification on software engineering concepts and advanced applied experience delivering system design, application development, testing, and operational stabilityAdvanced knowledge of reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with considerable in-depth knowledge ...