1,726 to 1,750 of 2,348 Remote/Hybrid Observability Jobs

Remote Head of Engineering, POS Application Platform

Hiring Organisation
grabjobs
Location
Skelmorlie, North Ayrshire, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Stornoway, Eilean Siar, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Dungannon, Co. Tyrone, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Southend-on-Sea, Essex, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Ballyclare, Co. Antrim, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Craigavon, Co. Armagh, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
grabjobs
Location
Bolton le Sands, Lancashire, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Pontypridd, Rhondda Cynon Taf, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

Site Reliability Engineer NEW Posted today Hemel Hempstead Haven Haven

Location
Hemel Hempstead, England, United Kingdom
Tech Leads to design, implement and support the systems that guests, owners and colleagues rely on every day. From CI/CD pipelines and observability through to database reliability, incident management and disaster recovery, this role touches every layer of our stack. This is also a great time to join. … developers and engineers to troubleshoot build and deployment issues and unblock delivery Contribute to and maintain internally developed engineering tools Own monitoring, tracing and observability so we are first to know when something is (or is about to be) an issue, and can diagnose it quickly Drive database reliability across ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
across our systems. Improve developer experience through automation, self-serve tooling and infrastructure-as-code practices that increase engineering velocity. Own and evolve our observability, incident response and reliability practices to minimise downtime and improve system performance. Partner closely with product and engineering teams to support new services, migrations … production environments. Strong experience operating and scaling infrastructure components such as Kafka, Redis and managed relational databases like RDS. Strong understanding of networking, security, observability and distributed systems fundamentals. Experience improving CI/CD pipelines and deployment automation in fast-moving engineering teams. Comfortable debugging production issues and participating ...

Senior Software Engineer, ML Infrastructure

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Principal AI Platform Engineer

Location
City of Edinburgh, Scotland, United Kingdom
security, performance and availability.* Using technologies such as containers, Kubernetes, vLLM and AI gateway platforms, you will deploy and operate scalable inference services, improve observability and performance, and investigate complex technical issues across the platform. You will also help establish engineering standards for operating AI services within a secure enterprise … such as LiteLLM, Bifrost or similar* GPU workloads, including performance, utilisation and resource management* Programming and scripting, such as Python, Bash or Go* Monitoring, observability and SRE practices* Secure, resilient and scalable service design* Authentication, access control, rate limiting and service integration* Technical documentation and operational guidance **Security Clearance ...

Senior Software Engineer: Agentic Development Enablement

Location
Greater London, England, United Kingdom
Claude Code and GitHub Copilot Design and implement practical guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Help define how controls should work consistently across local development environments and CI/CD pipelines Explore changes to the development environment … background as a software engineer Broad technical understanding across several of the following: developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end‐user application delivery Ability to work in ambiguous ...

Senior Platform Engineer

Location
Warminster, England, United Kingdom
Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This is not a pure cloud or greenfield platform role. … failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. Understanding of Distributed Systems in production. ...

Senior Software Engineer - Backend

Location
Manchester, England, United Kingdom
caching and data-access strategies Building reliable transactional workflows Developing applications and services within AWS Deploying and operating containerised applications Improving monitoring, alerting and observability Exploring AI-assisted engineering and modern development tooling At Senior level, you’ll be expected to understand the wider system rather than only the individual … traffic or data volumes Performance optimisation, caching and latency reduction Designing for resilience and failure Containerisation and orchestration technologies such as Docker and Kubernetes Observability, monitoring and operating production systems Automated testing and modern engineering practices We don’t expect candidates to have worked with every technology in our stack. ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...

Senior DevOps / Platform Engineer - Autonomous Vulnerability Research (Harness Engineering)

Location
Greater Manchester, England, United Kingdom
reach sanctioned targets. Build CI/CD pipelines with integrated supply-chain security: SBOMs, image signing, artefact provenance, and automated policy gates. Deliver comprehensive observability - logs, metrics, distributed traces, and cost telemetry - across long-running, non-deterministic agent workloads. Build evidence-capture pipelines: immutable audit trails, artefact retention, and reproducible … egress control, network policy, and segmentation in cloud-native environments. Strong CI/CD engineering skills and experience embedding security controls into delivery pipelines. Observability expertise across logging, tracing, and metrics, including designing for auditability and evidence retention. A security-first mindset with the judgement to balance researcher velocity against ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
England, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
Lincoln, Lincolnshire, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
Barrowford, Lancashire, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
Hingham, Norfolk, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
Thornbury, Gloucestershire, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...

Remote Data Engineering Manager

Hiring Organisation
grabjobs
Location
Prestatyn, Denbighshire, UK
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Nice-to-have: Experience implementing data ...