3,226 to 3,250 of 3,972 Observability Jobs

Lead AI Architect: Agentic Systems & Observability

Location
Greater London, England, United Kingdom
clients transform their digital platforms and products. As an AI Engineer, you'll own the development of the core agentic frameworks, evaluation pipelines, and observability tooling that let it operate with safety, trust, and intelligence at scale. You won't just be integrating AI into a product ...

Network Splunk Engineer - Telemetry & Observability (Hybrid)

Location
Wrexham, Wales, United Kingdom
Splunk Developer for a 12-month contract in Chester with hybrid work 3 days per week. The role focuses on telemetry ingestion, monitoring, and observability across an enterprise network environment. You will own development of Splunk dashboards, alerts and analytics to support Network Operations. Ideal candidates will bring extensive hands ...

M365 Copilot Incident & Observability Lead

Location
Birmingham, England, United Kingdom
ensure audit-ready artifacts within enterprise-scale environments. You will handle complex escalations across Copilot, Entra ID, and SharePoint/OneDrive permissions, contributing to observability and monitoring across Microsoft 365 services. Candidates should have 5–8+ years of enterprise M365 experience, and familiarity with PowerShell and Graph API will ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City, London, United Kingdom
Employment Type
Permanent
Salary
GBP 780 - 830 Daily
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone ...

Network Engineer – Observability, Automation & Reliability

Location
Greater London, England, United Kingdom
drive reliability, performance, resilience and security of our network and security infra from our London HQ. You will own end-to-end operations, enhance observability, and champion automation across Cisco and Arista environments. The role requires hands-on experience with BGP/OSPF, EVPN/VXLAN, MPLS, and network monitoring ...

Solutions Engineer - Data & Observability for SaaS Security

Location
Greater London, England, United Kingdom
Splunk, a Cisco company, is seeking a Solutions Engineer to support the UK Commercial Sales team. You will be the technical authority on Splunk solutions, explaining how they integrate within client workflows and tooling and ...

Observability Solutions Consultant - Client-Facing

Location
Greater London, England, United Kingdom
Itrs Insights is seeking a Professional Services Consultant for their London HQ. This role involves managing client projects and improving service offerings. Candidates should have at least 12 months of ITRS Geneos expertise, 3 years ...

AIOps Product Lead: Observability & Automation

Location
Greater London, England, United Kingdom
S&P Global, Inc. is seeking an experienced AIOps Product Manager to own the enterprise AIOps roadmap, translating complex operational needs into concrete product requirements and backlog items. You will lead high-impact capabilities like ...

Azure Cloud Engineer

Hiring Organisation
Method-Resourcing
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £650 per day
Azure Observability/Grafana Engineer - AI Monitoring - Contract Method Resourcing are supporting a major organisation as they expand their use of AI and intelligent automation across the business. As the number of AI applications and agents increases, the organisation wants to establish a central monitoring and observability capability that provides … being used, by whom, at what cost, and how effectively it is performing. We are looking for an experienced Azure Observability/Monitoring Engineer to initially design the monitoring solution before supporting its build and implementation. The Project The engagement will begin with an initial 4-week design and discovery ...

Cloud SRE Lead: Platform Reliability & Automation

Location
Halifax, England, United Kingdom
Engineer to strengthen reliability across Azure and Google Cloud Platform. You will lead a team of SREs, set engineering standards and drive improvements in observability, incident response and platform reliability. Collaborate with Product Owners, Engineering Leads and platform teams to ensure resilient, scalable cloud services while enabling feature delivery. #J ...

Senior Director, Data & AI Platform Engineering

Location
United Kingdom
workflows. You will partner with Platform Product Management to ensure scalability, reliability, security, and broad adoption across product domains. You will drive production-grade observability, governance, and AI safeguards while aligning with #J-18808-Ljbffr ...

Senior Backend Engineer, Pricing Platform (London)

Location
Greater London, England, United Kingdom
will design and build pricing features end-to-end, partnering with product teams to ship impactfull updates. You will contribute to architecture, improve reliability, observability, and scale across distributed services in a fintech environment, while mentoring teammates and owning the roadmap for pricing infrastructure. #J-18808-Ljbffr ...

Hybrid Linux Automation Engineer – Travel Expensed

Location
Milton, Scotland, United Kingdom
automation across large-scale enterprise infrastructure, including Linux, VMware, and F5 environments. You will build automation to accelerate patching and changes, strengthen validation, improve observability with Prometheus, Grafana, and Airflow, and work with Python and Ansible within a collaborative engineering team. #J-18808-Ljbffr ...

Senior SRE - Cloud Reliability & Automation

Location
Knutsford, England, United Kingdom
embedding SRE practices and maturity across diverse stakeholder groups. You will apply advanced programming, automation, and data‐driven approaches to reduce incident impact, improve observability, and accelerate delivery. #J-18808-Ljbffr ...

AI Platform & SRE Transformation Lead

Location
United Kingdom
engineering across LLM/ML operations and governance, partnering with senior stakeholders to move from experimentation to production. You will guide enterprise AI platforms, observability, and reliability while shaping operating models and governance essential for safe, scalable AI-enabled #J-18808-Ljbffr ...

Senior Data Platform Engineer: Lakehouse & Data Catalog

Location
Greater London, England, United Kingdom
lead by example with clean, well-tested code. Collaborate with cross-functional partners to clarify requirements, deliver reliable services, mentor peers, and drive observability and production reliability across Expedia Group's Lakehouse initiatives. #J-18808-Ljbffr ...

DevOps Engineer – Scalable Cloud & Live Ops

Location
Greater London, England, United Kingdom
enable continuous delivery, improve reliability, and support high-profile clients and live event services, leveraging cloud-first architectures, IaC, CI/CD, and comprehensive observability tools. #J-18808-Ljbffr ...

Kubernetes Platform Engineer (AWS) – London On-Site

Location
Greater London, England, United Kingdom
Kubernetes-based platform on AWS, collaborating with Security, Architecture and Developer Experience teams. The role demands hands-on expertise in production Kubernetes, cloud infrastructure, observability, and automation, with 5 days in the London office and a short notice period where applicable. #J-18808-Ljbffr ...

Monitoring Network Engineer – Global Finance Infra

Location
United Kingdom
Hamilton Barnes Associates Limited is seeking a Network Monitoring Engineer to design, maintain, and enhance observability across enterprise network infrastructure to ensure real-time operational visibility. You will build dashboards and metrics pipelines using Prometheus and Grafana, proactively detecting performance anomalies and diagnosing root causes, while partnering with cross-functional ...

MLOps & Infra Engineer

Location
Greater London, England, United Kingdom
software on HPC platforms, enabling distributed systems across the team. You’ll work with Kubernetes, Terraform, and CI/CD practices, shaping cloud infrastructure, observability, and ML workflow orchestration within a flexible hybrid UK setup. #J-18808-Ljbffr ...

Staff Cloud Native Engineer: Kubernetes AI Infra Leader

Location
United Kingdom
with core networking components on GPU-backed infrastructure. You will extend control plane capabilities, own significant components end-to-end, and raise reliability and observability across platforms. Strong Go skills and hands-on Kubernetes internals are essential for collaboration with platform teams. #J-18808-Ljbffr ...

Distinguished Engineer (Head of Service Reliability Engineering) - HMRC - SCS1

Location
Manchester, England, United Kingdom
Head of Service Reliability Engineering, you will lead a distinct, independent capability spanning services, products and platforms. You will make reliability, operability and observability integral to engineering from the outset. With enterprise-wide reach and influence, you will set the direction for SRE, raise service maturity and build a lasting … adoption of SRE practices by using SLOs, error budgets and reliability metrics to drive measurable improvements in service performance and operational decision-making. Observability: Establishes and governs enterprise observability capabilities, using telemetry, dependency mapping and modern monitoring platforms to improve operational insight and decision quality. Resilience and Recovery: Designs ...

Principal Engineer - Member Experience Platform

Location
Skipton, England, United Kingdom
Quality), and bar‐raising across squads: you shorten lead times, increase deployment frequency, hold change‐failure rate low, and improve MTTR through release‐linked observability - turning fast, safe flow into the default way of working. Operating at platform scale, you define cross‐cutting architecture and delivery standards (API/event … contracts, resilience, observability, language/dependency baselines) and drive adoption through the Golden Path: policy‐as‐code CI/CD, progressive delivery (feature flags, canary/blue‐green), automated rollback/forward‐fix, ephemer... data‐ready environments, and guardrails that make security and compliance by design. You partner with Platform ...

GenAI Full-Stack Engineer (Python/TypeScript) – Hybrid

Location
Belfast City District, Northern Ireland, United Kingdom
/Azure cloud architecture, delivering production GenAI systems used across global operations. This hybrid role focuses on feature implementation, system integration, cloud architecture and observability tooling, ensuring secure coding and scalable deployments. #J-18808-Ljbffr ...

Senior Backend Engineer – AI-Driven FinTech Platform

Location
Greater London, England, United Kingdom
that supports AI-driven advisory workflows and regulatory compliance. You will contribute to event-driven architectures, robust data models and reliable integrations, ensuring correctness, observability and scalable operations for complex financial services platforms. #J-18808-Ljbffr ...