2,551 to 2,575 of 4,362 Permanent Observability Jobs

Hybrid SRE: AI-Driven Reliability & Observability

Location
Manchester, England, United Kingdom
role involves developing tools, participating in incident resolution, and implementing best practices. Ideal candidates will possess strong software engineering skills and experience with contemporary observability tools. Bonus benefits include eye care and life assurance. #J-18808-Ljbffr ...

Senior Platform Engineer - iCloud Observability

Location
Greater London, England, United Kingdom
London is seeking an experienced Software Engineer to help develop the next generation of Apple's cloud services platform and infrastructure, focusing on iCloud observability and distributed data systems. You will design and implement fault-tolerant backend services, collaborate across iCloud teams, and own deployment, monitoring, and performance at scale ...

Platform SRE for Observability — Satellite Networks

Location
Greater London, England, United Kingdom
Aalyria Technologies in London is seeking a senior SRE to design and build a centralized observability platform for satellite‐level systems. You will shape the strategy, implement best practices, and automate the stack using Terraform and ArgoCD. You will own SLOs/SLIs, contribute to incident response, and partner with ...

Platform Engineer: Observability, SIEM & Automation

Location
Greater London, England, United Kingdom
Berenberg in London is seeking a Platform Engineer for MONSO, focusing on monitoring, observability and SIEM, and enabling self-service infrastructure through as-code platforms on Kubernetes. You will build and scale SIEM tooling, integrate data pipelines, and collaborate with CyberSecurity and development teams to automate detection and response, while ...

Senior SRE - Platform Reliability & Observability Lead

Location
Cambridge, England, United Kingdom
seeking a Senior Site Reliability Engineer to lead reliability, performance and continuous improvement of the Bango Platform. You will own end-to-end reliability, observability and incident response across infrastructure and delivery pipelines, serving as a technical centre of gravity for the SRE function. You will shape the bench across ...

SRE & Reliability Lead — AI-Ops & Observability

Location
Carshalton, England, United Kingdom
services used by internal and external customers. You will drive reliability improvements, advance automation and AI-Ops capabilities, and lead a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities include translating priorities into clear plans, line managing team leaders, ensuring RCAs and post-mortems are completed ...

Senior Lead SRE - Reliability & Observability Leader

Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase in the United Kingdom is seeking a Senior Lead Site Reliability Engineer to join an agile team focused on reliability, observability, and performance across critical platforms. You will mentor engineers, lead incident response, and shape SRE strategy while delivering scalable, secure production systems. The role demands deep expertise ...

Senior SRE & Observability Leader, AI-Driven Ops

Location
Birmingham, England, United Kingdom
OneAdvanced is seeking an Senior Manager, Site Reliability Engineering in Birmingham to lead the monitoring and observability roadmap across Health, Legal, Education, and Workforce products. You will guide a 15–20 engineer team, champion SRE practices, and push AI-driven monitoring and automation to reduce toil. You will partner with ...

Lead Observability Engineer

Hiring Organisation
Tria Recruitment
Location
London, UK
Employment Type
Full-time
Location: London, onsite 3 days per week (Sheffield as an alternative) Rate: £tbd/day inside IR35 Duration: 6 months+ Are you a Senior Observability Engineer/SRE Lead, with demonstrable experience of assessing and defining observability and monitoring roadmaps within enterprise scale environments? If so, apply now for this … contract opportunity. The Lead Observability Engineer/SRE Lead will be required to assess a complex hybrid estate, understand how services, platforms, infrastructure and networks should be monitored, and work across multiple internal teams, partners and suppliers to build a consolidated view of existing telemetry, monitoring and alerting capabilities. ...

Senior Java Engineer — Platform & Observability (Hybrid)

Location
Greater London, England, United Kingdom
London is seeking a Senior Java Software Engineer to join the Platform Team. You will help build the core platform that ingests and transforms observability data for multiple applications, working on a hybrid schedule from our London office. You will participate in all phases of the product lifecycle, mentor peers ...

Senior Software Engineer II — Reliability & Observability

Location
Greater London, England, United Kingdom
self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services. You will own incident management tooling, evolve observability infrastructure with SLOs and real-time signals, and contribute to AI-driven automation that reduces toil and speeds delivery. #J-18808-Ljbffr ...

Staff Software Engineer, Observability & Profiling

Location
Greater London, England, United Kingdom
policy experts, and business leaders working together to build beneficial AI systems. About the role the company is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at the company depends on—from … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Azure Observability - AI Monitoring - Contract

Hiring Organisation
Method Resourcing
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£600.00 - £650.00 per day
Azure Observability/Grafana Engineer - AI Monitoring - Contract Method Resourcing are supporting a major organisation as they expand their use of AI and intelligent automation across the business. As the number of AI applications and agents increases, the organisation wants to establish a central monitoring and observability capability that provides … being used, by whom, at what cost, and how effectively it is performing. We are looking for an experienced Azure Observability/Monitoring Engineer to initially design the monitoring solution before supporting its build and implementation. The Project The engagement will begin with an initial 4-week design and discovery ...

Staff Software Engineer, Observability & Profiling

Location
Greater London, England, United Kingdom
engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at Anthropic depends on—from metrics … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Senior Cloud SRE: Automation, Reliability & Observability

Location
Swindon, England, United Kingdom
Octave in Swindon is seeking a Cloud Engineer with a focus on SRE to design, build, and support cloud platforms and services. You’ll collaborate with development, architecture, security, and operations to drive automation, reliability ...

Remote UK SRE: Observability & Reliability Lead

Location
United Kingdom
Orex Nova, Inc. is seeking an experienced SRE/Platform Engineer to help keep our systems fast, reliable, and observable. This fully remote role covers the UK and requires strong incident response experience. You will ...

Platform Engineering Director - SaaS & Observability

Location
Greater London, England, United Kingdom
ITRS is seeking an experienced Director of Platform Engineering to lead our Analytics SaaS platform, focusing on Kubernetes-based services, cloud-native operations, and reliable service outcomes. You will guide hands-on SaaS engineering, set ...

Senior Network SRE: Automation, Reliability & Observability

Location
Greater London, England, United Kingdom
A leading IT solutions provider in London is seeking a Senior Network Site Reliability Engineer (SRE) with extensive experience in network engineering and automation. The ideal candidate will design and maintain high-availability network solutions ...

Junior Platform Engineer Telemetry, SIEM, Observability

Hiring Organisation
apto solutions
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£40,000
engineering seat with real customers, real supervision and a clear path up. ABOUT APTO SOLUTIONS Apto Solutions is a Bristol business that runs telemetry, observability and SIEM platforms for other organisations . Banks, airports, energy companies, engineering firms and government bodies depend on those platforms to tell them whether their ...

Software Developer – Deploy & Run – Remote: UK Skilled Worker Visa Sponsor

Location
United Kingdom
tooling in Python Enhance CI/CD pipelines using GitHub Actions Use generative AI to build and maintain internal services, including agentic services Improve observability with Datadog dashboards, metrics, logging, tracing, and alerting Partner with product engineering teams to improve developer experience Reduce engineers’ cognitive load by standardizing workflows … modern cloud-native deployment patterns Understanding of GitOps principles Experience with ArgoCD, GitHub Actions, or similar deployment and automation platforms Knowledge of reliability, observability, operational simplicity, monitoring, logging, metrics, and incident response Clear and concise written and oral communication Ability to define requirements in RFC/PRD documents or ticketing ...

Microsoft AI CoPilot Developer

Location
Leeds, England, United Kingdom
with enterprise systems via Microsoft Graph, APIs, and Power Platform connectors Implement secure-by-design and responsible AI practices: guardrails, controls, monitoring, auditability Build observability: logging, telemetry, and LLM monitoring for quality and incident triage Create reusable assets - prompt libraries, agent templates, connectors, test harnesses, and documentation Conduct rapid prototyping … cloud-native engineering, and secure deployment patterns Experience with agent engineering: orchestration, lifecycle management, versioning, drift detection Familiarity with DevOps, CI/CD, IaC, observability, and modern engineering pipelines Ability to debug unexpected AI behaviour - hallucinations, variability, reliability issues Strong documentation skills and ability to produce reusable code assets ...

Senior AI Engineer, Azure DevX & MCP Pipelines

Location
Leeds, England, United Kingdom
platform on Microsoft Azure, focusing on provisioning infrastructure, MCP integration, and robust agent deployment pipelines. The role emphasizes IaC automation, CI/CD, and observability, with a strong DevOps culture and opportunities to work with cutting-edge AI tooling in a collaborative environment. #J-18808-Ljbffr ...

AI Support Analyst

Location
Churwell, England, United Kingdom
keeping it running, 24/7. As an AI Support Analyst, you’ll be working at the intersection of cutting-edge AI-driven observability and real-world operational decision-making. You’ll help shape how DAZN uses machine intelligence to detect issues earlier, reduce incidents, and continuously improve the reliability … Global Operations Centre, responsible for proactively identifying risks, anomalies and emerging issues across DAZN’s OTT and broadcast platforms using AI-powered monitoring and observability tools. This role bridges human judgement with machine intelligence. You will actively interrogate AI systems through advanced prompting and analytical review, validating insights against live ...