Site Reliability Engineer, K8s (Remote International)

DescriptionWebMD and its affiliates is an Equal Opportunity/Affirmative Action employer and does not discriminate on the basis of race, ancestry, color, religion, sex, gender, age, marital status, sexual orientation, gender identity, national origin, medical condition, disability, veterans status, or any other basis protected by law. Build the platform behind PulsePointPulsePoint operates large-scale data and advertising platforms that power business-critical services used every day across the company.Our Platform Engineering team owns the foundation that enables engineering teams to move quickly and safely. We build and operate the infrastructure that supports Kubernetes workloads, data platforms, developer tooling, observability and production operations at scale.Unlike many cloud-only environments, we own the full lifecycle of the infrastructure - from bare-metal hardware and networking to Kubernetes, observability and developer experience.The environment supports large-scale Kubernetes workloads, multi-petabyte data systems and business-critical services used across multiple engineering organizations.We're looking for an experienced engineer to help shape its next stage of growth.This is a hands-on role focused on architecture, reliability, automation and operational excellence. You will work on complex distributed systems, drive platform improvements and help shape the technical direction of infrastructure across the company.What you'll work onYou'll help design, build and operate the Kubernetes platform used across PulsePoint.Examples of challenges you may work on include:Platform architecture and Kubernetes lifecycle managementReliability, observability and incident responseInfrastructure automation and GitOps workflowsNetworking, service connectivity and platform securityDeveloper experience and self-service platform capabilitiesLarge-scale distributed systems running on bare-metal infrastructureTechnologyYou'll work in an environment that includes:Kubernetes and platform servicesMulti-petabyte data infrastructureBare-metal and cloud environmentsGitOps and infrastructure automationModern observability and reliability engineering practicesTechnologies commonly used across the environment include Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis and Ceph.Experience with every technology is not required.Who we’re looking forSuccess in this role is not measured by the number of tickets closed or clusters operated. Success means building platform capabilities that make engineering teams more reliable, productive and autonomous.We're especially interested in engineers who:Have operated production infrastructure at meaningful scaleUnderstand how distributed systems fail and recoverPrefer automation over repetitive operational workEnjoy simplifying systems rather than adding complexityTake ownership beyond the boundaries of a single componentWilling and able to work 9am-6pm ET U.S. hours. You can work fully remotelyWhy this roleThis is not a ticket-driven operational role.You'll help define platform architecture, influence engineering standards and work on infrastructure that supports multiple engineering organizations.The engineer joining this role is expected to become a key technical contributor shaping the future of the platform.We try to keep the process focused and practical.Introductory conversation (~60m)Learn about your background and discuss the role.Technical discussion (~60m)Deep dive into systems engineering, Kubernetes and operational experience.Architecture discussion (~60m)Explore platform design, distributed systems and technical decision making.Leadership conversation (~30m)Meet engineering leadership and discuss team, strategy and long-term direction.

Job Details

Company
PulsePoint
Location
United Kingdom, UK
Hybrid / Remote Options
Posted