Site Reliability Engineer, Big Data (Remote, International)
- Hiring Organisation
- PulsePoint
- Location
- United Kingdom, UK
- Employment Type
- Full-time
hybrid infrastructure: bare-metal on-prem, cloud, and the integration between them. You own the lifecycle from architecture through deployment to capacity planning and incident response. What you'll work onKafka architecture, topic design, governance, partition strategy, throughput and latency optimization. Ceph operations, pool design, placement optimization, capacity planning. … Operational automation, reduce manual work, faster incident response, preventive systems. SQL Server backup and recovery pipelines, basic cluster support. Data team tooling with self-service capabilities and observability. TechnologyApache Kafka for messaging layerHadoop and Ceph as distributed storage layerSQL Server backup and recoveryTerraform, Ansible, Puppet, ArgoCD for operational ...