401 to 425 of 502 Remote/Hybrid Observability Jobs

Data Services Sales Leader – Remote-First, EMEA/APJ

Hiring Organisation
Jobleads-UK
Location
Windsor, England, United Kingdom
Manager for our Data Services portfolio to lead a high-performing team of sales specialists across EMEA and APJ, driving the ARR target for Observability and Cyber Resilience solutions in large enterprise accounts. You will develop and execute a comprehensive sales strategy, oversee pipeline reviews and revenue forecasting, coach ...

Site Reliability Engineering Manager

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
Reliability Engineers. Shape and deliver our Site Reliability Engineering roadmap alongside the Head of Platform. Champion modern engineering practices including SLIs, SLOs, error budgets, observability and automation. Improve the reliability, scalability and performance of our cloud platforms and digital services. Partner with Engineering, Security, Data and Product teams to embed … operational excellence from design through to production. Drive the adoption of our observability platform, helping teams gain deeper insight into the health and performance of their services. Lead incident learning, continuous improvement and automation initiatives that reduce operational toil. Provide technical leadership across AWS, Kubernetes, Infrastructure as Code, CI/ ...

Azure Platform Engineering Consultant

Hiring Organisation
Morgan McKinley
Location
Newbury, Berkshire, England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
platform templates and landing zone patterns. CI/CD & Automation: Build and refine automated deployment pipelines, environment management, and release practices. Platform Quality: Embed observability (monitoring, logging, alerting), resilience, security, and FinOps principles directly into platform assets. Co-Delivery & Knowledge Transfer: Work closely alongside client engineering teams to pair, document … Core compute, networking, storage, identity, security, and platform services. Infrastructure as Code: Strong proficiency with Terraform AND Terragrunt using modular, reusable implementation patterns. DevOps & Observability: Strong experience with CI/CD tools (Azure DevOps/GitHub Actions) and monitoring stacks (Prometheus, Grafana, Azure Monitor, etc.). FinOps: Practical knowledge ...

Permanent position: Senior Python Backend Engineer (AI Systems)

Hiring Organisation
Nicoll Curtin Technology
Location
United Kingdom
Employment Type
Permanent
Salary
GBP Annual
responsible for building and operating the infrastructure and workflows that allow AI features to execute end-to-end, with a strong focus on reliability, observability, latency, and user experience. Areas of responsibility include: Building production Back End systems that serve AI-powered product features Designing retrieval, inference, and orchestration workflows … workflows Tool calling and workflow orchestration LLM-based applications used by real users Vector databases and semantic search AI evaluation and monitoring frameworks Production observability and reliability engineering Distributed systems and high-throughput services End-to-end ownership of Back End systems in production Technical Environment Python Node.js SQL/ ...

Forward Deployed AI Engineer

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
enabled systems. You’ll bring deep expertise across modern full-stack technologies (.NET, Azure, SQL, React/Angular), along with experience in distributed systems, observability, and AI tooling such as LLMs, retrieval pipelines, and agentic workflows. Acting as a bridge between business and technology, you’ll work across product, data … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Head of Platform Engineering, you'll define and deliver robust platform services, establish the patterns for self-service infrastructure, CI/CD, and observability, and drive the integration of AI/ML capabilities across the organisation. You'll translate platform strategy into a coherent technical roadmap, setting standards that teams … these are sustained as the platform evolves. Defining and maintaining platform standards, patterns, and reusable components that drive consistency across teams. Leading improvements to observability, monitoring, and alerting capabilities across systems and services. Mentoring and developing engineers across the organisation, raising the collective standard of platform engineering practice. Identifying ...

AI Platform Architect

Hiring Organisation
83zero Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £100000/annum
AI Platform Architect Location: London (Hybrid - 1-2 days per week) Salary: £80,000 - £100,000 + Bonus & Excellent Benefits We're partnering with a global technology consultancy delivering one of the UK's largest ...

SRE Managing Consultant - Cloud Operating Model

Hiring Organisation
Capgemini
Location
Manchester, United Kingdom
Employment Type
Full Time
Budgets : Establish service measures and targets (SLIs/SLOs) and introduce Error Budgets to enable data-driven trade-offs between reliability and delivery velocity. Observability & Operational Insight: Shape observability approaches (metrics/logs/traces) and operational monitoring models that make reliability risks visible and actionable, improving operational decision-making. … large‐scale delivery contexts; associate‐level certifications are desirable but not mandatory. Design, establish, and evolve SRE‐led centres of excellence (e.g. Reliability, Observability, or Operational Excellence), setting enterprise‐level standards for SLIs/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms. Exposure to modern observability ...

Observability Engineer

Hiring Organisation
Hays Technology
Location
Telford, Shropshire, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £590/day Per Day
large-scale, complex programmes across the public sector. Working within a collaborative and forward-thinking environment, you will play a key role in enhancing observability capabilities and driving proactive service management through modern monitoring and automation practices. Your new role As an Observability Engineer, you will be responsible for designing … monitoring solutions across a range of technologies and platforms. You will work closely with architects, engineers, and project teams to deliver end-to-end observability solutions that provide valuable performance insights, improve service stability, and support proactive incident management. Key responsibilities include: - Configuring and optimising Dynatrace monitoring solutions - Translating monitoring ...

Observability Engineer

Hiring Organisation
Hays Specialist Recruitment
Location
Telford, Shropshire, UK
Employment Type
Full-time
large-scale, complex programmes across the public sector. Working within a collaborative and forward-thinking environment, you will play a key role in enhancing observability capabilities and driving proactive service management through modern monitoring and automation practices. \n Your new role \n As an Observability Engineer, you will be responsible … monitoring solutions across a range of technologies and platforms. You will work closely with architects, engineers, and project teams to deliver end-to-end observability solutions that provide valuable performance insights, improve service stability, and support proactive incident management. \n Key responsibilities include: \n - Configuring and optimising Dynatrace monitoring solutions ...

Site Reliability Engineer (SRE)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
approach. Key Responsibilities Integrating tightly with our Product Engineering teams Following SRE practices and maintaining high standards of compliance Implementing a new standard of observability utilising SLI/SLO/Error Budgets Continually evolving our observability platforms for greater coverage Using a code-first approach to build and changes … ongoing communication with stakeholders Skills Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer Observability product experience (eg Datadog) Managing services using SLI/SLO & Error Budgets Experience with AWS or other cloud providers Experience in HA environments Automation skills ...

Senior Data Practice Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
pick the right approach for the problem Data testing: know what to catch at build time (schema contracts, assertions, transformation logic) versus defer to observability, and can make that call for a team Data observability: treat data reliability like site reliability, with measurable indicators, alerting, incident response, and root cause ...

Senior Backend Engineer - Databases Pyroscope | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
United Kingdom (Remote) Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations … cost that makes sense. Turn Pyroscope into a platform capability inside Grafana: bi-directional trace-to-profile correlation, integration with Kubernetes Monitoring and App Observability, and profiles surfaced where engineers already start their investigations. Prepare Pyroscope for an agent-driven world: APIs, CLI, and docs designed so AI agents ...

Software Engineer - Synthetic Monitoring | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … Monitoring as code and with confidence Break down complex, ambiguous problems into incremental deliverables and iterate quickly based on feedback Ensure quality through testing, observability of your own systems, and documentation : our checks are something customers alert on, so reliability is a feature Be a part of the team ...

Global DevOps Lead

Hiring Organisation
Stott & May Professional Search Limited
Location
United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
with engineering, cloud, and operations teams to deliver a modern, automated, and scalable platform. You'll drive DevOps strategy across infrastructure, CI/CD, observability, SRE, and cloud optimisation while influencing senior stakeholders across the business. Key Responsibilities - Define and implement a global DevOps operating model, including governance, standards … initiatives. - Partner with engineering and cloud teams to establish clear ownership across DevOps and infrastructure. - Lead the implementation and optimisation of enterprise monitoring and observability using Datadog. - Build scalable deployment pipelines that improve release quality and speed. - Establish and monitor DORA metrics, driving improvements in deployment frequency, lead time, change ...

Site Reliability Engineer

Hiring Organisation
Connells Limited
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
hands-on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post-incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer a hybrid working arrangement with one day per week … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands-on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Azure DevOps Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
modules aligned to business and security requirements Implement and enforce Azure security controls and policies Lead platform security initiatives across Azure environments Contribute to observability, alerting, and Site Reliability Engineering practices Collaborate with engineering teams to deliver resilient and scalable solutions Security & Platform Focus Areas Implement perimeter security using Azure … Terraform Kubernetes certification and experience with AKS Deep understanding of DevOps and platform engineering principles Strong knowledge of cloud security best practices Experience with observability, monitoring, and SRE concepts Why Apply Fully remote within the UK Work with a mission‐driven, highly respected health tech organisation Modern cloud environment with ...

Data Platform Lead

Hiring Organisation
Jobleads-UK
Location
Cardiff, Wales, United Kingdom
foundations that enable better decision-making across the business. From building resilient data pipelines and reusable data models to improving platform reliability, governance and observability, you'll create the capabilities that allow teams to move faster with confidence. You'll combine hands-on technical leadership with people management, building … future AI capabilities. Partner with Analytics, Product, Engineering and Architecture teams to ensure the platform supports current and future business needs. Improve platform reliability, observability, data quality and operational resilience. Ensure datasets are well-modeled, documented, discoverable and governed to support self-service analytics and AI readiness. Translate business ...

Senior Platform Engineer - Developer Experience

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
paved roads that other engineers use. You will work across the software development lifecycle, from creating a new service through to testing, deployment, observability and operating it in production. You will join an established Platform team and work alongside our existing Developer Experience Engineer. You will speak directly with engineers … reliability and usability of our CI/CD systems. Developing reusable platform capabilities that product engineers can consume through self-service. Helping engineers use observability effectively, with good defaults for logs, metrics, traces and service‐level indicators. Working directly with engineers to understand friction, test ideas and support adoption. Using ...

Backend Engineer - Platform - Stacks | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand ...

Senior Software Engineer - Customer Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
designing and maintaining RESTful APIs, web hooks, and service-to-service integrations Comfortable operating in AWS, GCP, or Azure environments and working with modern observability tooling, logging, and monitoring platforms A track record of delivering performant, reliable and scalable applications Excellent collaboration and communication skills in cross-functional teams, including … build systems, but why architectural decisions matter You can balance scalability, reliability, maintainability, and speed of execution You think critically about security, observability, and operational excellence from day one Love the idea of blending software development, distributed systems and data-intensive applications Strong familiarity with authentication and identity technologies such ...

SRE Technical Lead

Hiring Organisation
Capgemini
Location
Surrey, United Kingdom
Employment Type
Full Time
point for major incidents and high risk releases, protecting service stability and ensuring blameless post incident reviews lead to measurable improvement. • Define and govern observability and capacity practices so reliability risks are visible, actionable, and proactively managed. • Ensure SRE practices align with service governance, security, and compliance requirements, and contribute … including: • Strong expertise in Kubernetes and OpenShift. • Experience with multi cloud and hybrid architectures, including service mesh (e.g. Istio). • Hands on experience with observability platforms such as Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. • Strong Infrastructure as Code and GitOps experience (Helm, Kustomize, ArgoCD, Tekton). • Experience with CI/ ...

Senior Software Engineer - Customer Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
designing and maintaining RESTful APIs, web hooks, and service-to-service integrations Comfortable operating in AWS, GCP, or Azure environments and working with modern observability tooling, logging, and monitoring platforms A track record of delivering performant, reliable and scalable applications Excellent collaboration and communication skills in cross-functional teams, including … build systems, but why architectural decisions matter You can balance scalability, reliability, maintainability, and speed of execution You think critically about security, observability, and operational excellence from day one Love the idea of blending software development, distributed systems and data-intensive applications Strong familiarity with authentication and identity technologies such ...

Hybrid SRE Engineer — Observability & Cloud (London)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
help transform current workloads toward an SRE model while working in a hybrid setup, visiting the London office twice weekly. The role focuses on observability, high availability and incident management, with collaboration across Product Engineering and Infrastructure teams. Strong AWS, Terraform, Python and Kubernetes skills are valued. #J-18808-Ljbffr ...

SRE Consultant

Hiring Organisation
Akkodis
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £100000/annum
hold currently). The Role As a Site Reliability Engineer (SRE) you will lead site reliability engineering initiatives with a strong emphasis on observability, ensuring high performance and reliability of applications & infrastructure. Provide strategic insights to shape the overall SRE strategy while collaborating on the design and implementation of scalable … solutions. Establish effective monitoring, alerting and incident response strategies to maintain system availability and promote continuous improvement by collaborating with team members to deliver observability best practices and SRE methodologies. The Responsibilities Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to measure and maintain system ...