1,876 to 1,900 of 2,348 Remote/Hybrid Observability Jobs

Software Development Engineer in Test

Location
Greater London, England, United Kingdom
quality gates, and make production failures easier to identify and diagnose. You will work hands-on across C#, TypeScript, APIs, data, pipelines, and observability rather than focusing solely on UI testing. Your work will help Athos reduce manual checking, catch failures earlier, and create stronger evidence behind every release. … quality gates into Azure DevOps CI/CD pipelines. Investigate failures across application logs, APIs, databases, telemetry, test environments, and production systems. Improve observability and production quality through monitoring, alerting, and better visibility into feed and syndication failures. Contribute to test strategy, including determining what should be tested, at which ...

Software Engineering Manager

Location
Bristol, England, United Kingdom
risks early and transparently. Establish engineering guardrails across scope, quality, and non-functional requirements, enabling teams to design optimal solutions within them. Champion observability and operational excellence, ensuring system health, SLOs, and alerting are visible and actively managed. Partner with Tech Leads and Architects on system design and evolution, bringing … architecture, including API design and integration, performance optimisation, security, and microservice or event‐driven patterns. Experience with CI/CD, modern development workflows, and observability practices. Proven track record of leading high‐performing teams in fast‐paced, complex or regulated environments. Passion for mentoring and developing engineers through coaching, feedback ...

Platform Engineer

Location
City of Edinburgh, Scotland, United Kingdom
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort.We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration* Cloud platforms, virtualisation and containers* Infrastructure as code and configuration management* Programming and scripting, such as Python, Bash or Go* Monitoring, observability and SRE practices* Infrastructure, networking and performance troubleshooting* Secure, resilient and scalable system design* Technical documentation and operational guidance* Technical leadership and mentoringBeneficial skills include:* Strong ...

Engineering Manager (Remote - UK)

Hiring Organisation
Reonomy
Location
Manchester, UK
Employment Type
Full-time
building and evolving AWS-native and cloud-based platforms, alongside legacy systemsSet and uphold strong engineering standards across code quality, testing, CI/CD, observability and documentationStay close to technical decisions through design reviews, architecture discussions and hands-on coachingBalance new feature delivery with technical debt, reliability, security and long … distributed systemsProficiency in at least one modern programming languageStrong grasp of system design and software engineering fundamentalsExperience with Infrastructure as Code, CI/CD, observability and secure production systemsAble to communicate technical ideas clearly to both technical and non-technical stakeholdersUnlock your Altus Experience! If you're looking to advance ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
engineers to turn agreed product behaviour into implementation-ready technical solution designs covering assumptions, alternatives, service interactions, data and schema changes, failure modes, security, observability, rollout, rollback, and tests. Use agentic coding tools such as Codex, Claude Code, Cursor, or equivalent to generate, refactor, test, review, and document code within … with product and engineering colleagues through concise written analysis and appropriate architecture, sequence, data-flow, and entity-relationship diagrams. Sound judgement around automated testing, observability, production reliability, security, and data integrity. Strong working fluency with agentic coding tools and spec-driven development workflows, including specification writing, code generation, test generation ...

Senior Engineering Manager - 9-10 month FTC

Location
Manchester, England, United Kingdom
practices, including test‐driven development and automated testing, helping teams build quality into the development process. Support strong operational ownership through CI/CD, observability, production support and the principle that teams own the systems they build and run. Help teams identify and address technical debt sustainably while maintaining appropriate … outcomes and translating these into clear engineering priorities. Strong understanding of modern software engineering practices, including automated testing and TDD principles, CI/CD, observability and production ownership. Experience facilitating technical decisions, bringing the right people and evidence together and constructively challenging thinking when needed. Comfortable balancing feature delivery ...

Senior Engineering Manager - 9-10 month FTC

Location
Greater London, England, United Kingdom
practices, including test‐driven development and automated testing, helping teams build quality into the development process. Support strong operational ownership through CI/CD, observability, production support and the principle that teams own the systems they build and run. Help teams identify and address technical debt sustainably while maintaining appropriate … outcomes and translating these into clear engineering priorities. Strong understanding of modern software engineering practices, including automated testing and TDD principles, CI/CD, observability and production ownership. Experience facilitating technical decisions, bringing the right people and evidence together and constructively challenging thinking when needed. Comfortable balancing feature delivery ...

Forward Deployed Engineer

Hiring Organisation
Janus Henderson
Location
London, UK
Employment Type
Full-time
architecture and design decisions for your solutions, keeping them secure, scalable, and aligned to the enterprise core stack. Take solutions through evaluation, observability, and our AI governance checkpoints, and keep them healthy afterwards. Build investment capability on NexusBuild advanced investment capability on Nexus and make it available across the firm. … portfolio construction. Snowflake, Microsoft Fabric/OneLake, or comparable enterprise data platforms. Azure AI Foundry, model gateways, inference routing, or AI evaluation and observability in production. TypeScript or a second production language, and front-end experience for user-facing AI applications. Supervisory responsibilitiesNo. This is an individual-contributor role. Forward ...

Senior SRE: AWS, Dynatrace & Observability Lead (Remote)

Location
Birmingham, England, United Kingdom
Partners is seeking a Senior Site Reliability Engineer with strong AWS expertise and Dynatrace implementation experience to own observability and reliability across large-scale enterprise platforms in the UK. You will design and deploy Dynatrace across complex cloud environments, build dashboards, set monitoring standards, and drive improvements with IaC (Terraform ...

SRE: Cloud Reliability, DevSecOps & Observability — Hybrid London

Location
Slough, England, United Kingdom
will apply software engineering principles to automate, scale, and secure cloud-native environments. Responsibilities include building and maintaining production and demo environments, implementing observability with Prometheus, Grafana, and Loki, and guiding project teams in DevSecOps practices. Hybrid London model, SC level clearance may be required. #J-18808-Ljbffr ...

Observability Manager

Location
Greater London, England, United Kingdom
What you'll bring to the team Observability Manager Location: London/Hybrid Hours: 37.5 hours per week Contract: Permanent - Salaried At Merlin Entertainments , our purpose is simple but powerful: to bring joy, create connections and make memories . Merlin is embarking on an exciting Digital and Data Transformation focused … continue our ambitious global transformation journey, technology plays a critical role in enabling sustainable growth and unforgettable guest experiences across our iconic destinations. The Observability Manager is responsible for enabling deep, end-to-end visibility of IT services to support effective Incident, Major Incident, Problem, Change, and Service Improvement practices. ...

Staff Software Engineer, Observability & Profiling

Location
Greater London, England, United Kingdom
engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at Anthropic depends on—from metrics … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Senior AI / Agentic Engineer London

Location
Greater London, England, United Kingdom
Agentic ecosystem, responsible for the high-level design choices that define how agents run at PhysicsX. You will cover topics such as: Agent Observability: Own the implementation to enforce deep tracing, granular cost tracking, and observability across the lifecycle. Agent Deployment: Deliver an intuitive deployment lifecycle which simplifies questions around … behalf of users in a regulated enterprise environment. The Tech Stack Core Platform: Python (Primary), Go or TypeScript (Secondary), Kubernetes, Docker, Terraform. Observability & Evals: OTel, LangSmith, Arize, Braintrust. Who You Are An Architect at Heart: You have strong, reasoned opinions on Durable Execution vs. Standard Async, Vector Search vs. Keyword ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health ...

Senior SRE Engineer

Hiring Organisation
SF Partners
Location
Birmingham, West Midlands (County), United Kingdom
Employment Type
Permanent
Salary
£100000 - £110000/annum great training and progression opp
largest and most complex enterprise platforms. Working within a highly skilled, multi-disciplinary engineering team, this Senior SRE Engineer will take technical ownership of observability and reliability across large-scale AWS environments, helping engineering teams improve platform performance, resilience and availability. A key focus of the role will … enterprise environments - End-to-end Dynatrace implementation experience - Experience designing and deploying Dynatrace across complex cloud environments - Strong understanding of APM, infrastructure monitoring and observability - Dynatrace configuration, dashboards, alerting and performance monitoring - Experience integrating Dynatrace into AWS and wider engineering/tooling ecosystems - Strong understanding of SLIs, SLOs, availability, reliability ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets … operational tooling to reduce manual processes Requirements Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices Deep understanding ...

Senior DevOps Engineer

Location
East Midlands, England, United Kingdom
/CD pipelines and deployment orchestration. Support Kubernetes and OpenShift platform troubleshooting and optimisation. Deliver secure and compliant infrastructure solutions. Implement monitoring, logging, and observability tooling across environments. Collaborate with engineering, architecture, and delivery teams to improve deployment efficiency and platform reliability. Champion automation-first approaches to infrastructure and application … designing and maintaining enterprise-scale CI/CD pipelines . Strong understanding of cloud security and secure delivery practices. Experience implementing monitoring, logging, and observability solutions. Ability to define technical standards, governance, and reusable deployment frameworks. Experience working within large-scale enterprise transformation programmes. Desirable Skills Experience within highly regulated ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Hiring Organisation
Learning Technologies Group
Location
United Kingdom, UK
Employment Type
Full-time
This is a hands-on infrastructure role. You'll work across the full stack — from Terraform modules and Puppet manifests to Python automation and observability pipelines. The platform is not containerised — there is no Kubernetes here — so we're looking for someone who understands Linux systems deeply and can reason … service discovery and configuration management (etcd)Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache)Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotationsEvaluating and implementing distributed storage solutions as the platform evolvesImproving deployment workflows and release ...

Senior Observability Solution Architect – Pre-Sales

Location
Greater London, England, United Kingdom
leading observability platform in Greater London is seeking an experienced Solution Engineer to join their team. This role involves collaborating with account executives on technical sales cycles, delivering impactful presentations, and overseeing technical aspects of the process. The ideal candidate will have a minimum of 5 years in a customer ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Location
United Kingdom
This is a hands-on infrastructure role. You'll work across the full stack - from Terraform modules and Puppet manifests to Python automation and observability pipelines. The platform is not containerised - there is no Kubernetes here - so we're looking for someone who understands Linux systems deeply and can reason … service discovery and configuration management (etcd) Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache) Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotations Evaluating and implementing distributed storage solutions as the platform evolves Improving deployment workflows ...

Senior Backend Engineer (Python | AI | 3D Environments | £130,000)

Hiring Organisation
Paradigm Talent
Location
Cardiff, United Kingdom
implement distributed systems for async processing, ML workflows, and asset pipelines. Own authentication, billing, and subscription systems — ensuring reliability and seamless user experience. Drive observability, performance tuning, and deployment automation across the stack. Collaborate closely with product, frontend, and ML teams to deliver features that delight and scale. You should … Azure). Familiarity with auth, billing, or subscription systems . Background in 3D graphics, creative tooling, or ML pipelines . Knowledge of observability tools like Grafana, Prometheus, or OpenTelemetry. This is a rare opportunity to join an early-stage team backed by leading deep-tech investors, building the foundation ...

Staff Analytics Platform Engineer

Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Site Reliability Engineer

Location
Fenny Stratford, England, United Kingdom
hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Reliability testing and capacity planning Toil reduction We offer a hybrid working arrangement with one day per week in our Milton Keynes office. Key Responsibilities: Support … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands‐on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Site Reliability Engineer

Hiring Organisation
Connells Limited
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
hands-on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post-incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer a hybrid working arrangement with one day per week … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands-on experience with Azure Monitoring (Application Insights, Alerts, Action ...