3,451 to 3,475 of 4,093 Permanent Observability Jobs

Remote SRE: Cloud Reliability Engineer (AI Tools)

Location
Hemel Hempstead, England, United Kingdom
holiday operator, is seeking an experienced Site Reliability Engineer to join our Product Technology team. This role focuses on cloud reliability, CI/CD, observability and database resilience across diverse engines, with occasional travel to Hemel Hempstead and off-site events. You will work with engineering teams to design, implement ...

Senior Data Engineer: Scale Pipelines & AI-Driven Data

Location
Greater London, England, United Kingdom
hands-on role, you’ll mentor peers through code reviews and knowledge sharing while promoting best practices and an emphasis on CI/CD, observability, and IaC. The team operates in a hybrid model across the UK. #J-18808-Ljbffr ...

SRE Associate: Build Reliable Cloud Platforms

Location
Birmingham, England, United Kingdom
seeking a Site Reliability Engineer to join Core Engineering. You will help build, run and maintain high-performing, distributed systems, focusing on reliability, observability and automation across critical services. Responsibilities include monitoring production services, capacity planning, incident response and driving improvements to SLIs/SLOs. The role requires ...

Remote SRE: Cloud Reliability & AI-Driven Ops

Location
Hemel Hempstead, England, United Kingdom
Haven is seeking a hands-on Site Reliability Engineer to join our Product Technology team. This remote-first role involves shaping CI/CD, observability, and incident response while collaborating with engineers and tech leads to ensure reliable, scalable platforms for guests and colleagues. You’ll tackle infrastructure design, tooling ...

Operations Team Lead: Scale Reliable Production

Location
Guildford, England, United Kingdom
Complexio is seeking an Operations Team Lead to own and scale production across live customer-facing systems. You will drive reliability, observability, and continuous improvement, shaping processes, leading incidents, and building a high-performing team. This hands-on role requires deep SRE/DevOps expertise, leadership, and a calm, clear ...

Senior Go Software Engineer — Low-Latency Trading Platform

Location
Greater London, England, United Kingdom
develops our trading platform. You will architect and implement features to meet performance, stability, and product goals while upholding technical standards, CI/CD, observability and security. Our Golang-centric stack handles high-throughput, low-latency data with multiple trading venues and global data feeds. The role is based ...

Senior Full-Stack Engineer — Backend-Heavy with AI Impact

Location
Greater London, England, United Kingdom
deliver measurable impact at scale. This is a remote, internationally oriented role with opportunities across 175 countries. You’ll contribute to architectural direction, champion observability, and foster engineering excellence while balancing short-term delivery with long-term #J-18808-Ljbffr ...

Senior AI/ML Platform Reliability Engineer

Location
Auchentibber, Scotland, United Kingdom
platforms, with a focus on reliability and security. As part of the Reliability Engineering team, you will own non-functional requirements, build tooling for observability and resilience, and partner across teams to unblock high-impact AI use cases. This role emphasizes production-grade code, on-call readiness, and shaping ...

Senior Real-Time Full-Stack Engineer (Voice & AI)

Location
Greater London, England, United Kingdom
front ends, Scala and Python services, and cloud infrastructure. You will own realtime voice integrations, streaming pipelines, and multi-tenant architecture while ensuring safety, observability, and robust go-lives across partner integrations. The role emphasizes senior ownership, low-latency delivery, and collaboration with product, design, and operations to ship iteratively ...

Backend Engineer — AI-Powered Expense Platform

Location
Greater London, England, United Kingdom
will ship backend and frontend features, build AI-backed agents, and improve extraction, policy, and audit capabilities at scale. The role emphasizes production quality, observability, and cross-team collaboration across time zones. Ideal candidates have 5+ years in production software, strong CS fundamentals, and experience with Java/Spring Boot ...

Lead Backend Engineer — AI-Driven Insurance Platform

Location
Greater London, England, United Kingdom
real-time insurance platform. You will guide several engineers, shape backend architecture, and drive end-to-end delivery with a strong emphasis on simplicity, observability, and reliability. You will collaborate across squads, mentor peers, and advocate AI-first approaches to deliver high-impact solutions in a fast-paced, hybrid London ...

Backend SDE II — Real-Time Data & Event-Driven Systems

Location
Welwyn Garden City, England, United Kingdom
delivering personalised experiences at scale while collaborating with more senior engineers. You will gain hands-on experience with Kafka, Kubernetes, CI/CD, and observability tools, contributing to distributed systems and production readiness in a fast-paced environment. #J-18808-Ljbffr ...

Platform Engineer

Location
City Of London, England, United Kingdom
platform, working across IAM, Kubernetes, networking, logging, monitoring, and cloud architecture. Key experience: AWS Platform Engineering Kubernetes/EKS IAM & Cloud Security CloudTrail, GuardDuty & Observability Cloud Architecture & Infrastructure Documentation Security or compliance focused engineering projects London, 2 days onsite per week Up to £450/day, Outside IR35 #J ...

Principal Engineer - Core Banking Platform

Location
Skipton, England, United Kingdom
lead time, high deployment frequency, low change-failure rate and rapid recovery. You’ll equip teams with paved roads, Golden Path pipelines, strong observability and secure environments that make fast, safe change the norm. You’ll also grow skills across squads, align engineering practices and create a trusted environment where … release strategies and resilience-first patterns that allow teams to deliver change safely and confidently as we modernise the platform. Architect for reliability and observability, shaping API/event contracts, posting and settlement patterns, product and account boundaries and platform dependency baselines. Shift left on quality and security, embedding contract ...

Site Reliability Engineer / Production Support

Location
Greater London, England, United Kingdom
person when incidents occur, escalating only to Head Of when required. Run on-call and incident response; ensure fast detection, triage, and restoration. Maintain observability standards (logs, metrics, traces) and alert quality (low noise, high signal). Understand at a working level all key system flows, the services, partners … first - every manual task is a candidate for automation. You actively hunt for toil and eliminate it. Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source. WHAT YOU BRING Strong SRE or production support experience with accountability for incident response ...

Senior Dedicated Support Engineer - 24/7 Incident Response

Location
Greater London, England, United Kingdom
incident resolution, automation, and cross-team collaboration to ensure high-quality service experiences. The position requires strong coding in Python, experience with AWS and observability tools, and willingness to join a 24x7 on-call rotation. Bachelor's degree is preferred, with 5–7 years of relevant work experience. #J ...

Senior Technical Product Manager - API Experience & Security

Hiring Organisation
Wise
Location
London, UK
Employment Type
Full-time
webhook infrastructure and tooling of Wise Platform API integrations. You'll focus on reducing partner-impacting incidents (especially auth- and webhook-related), improving observability and self-serve troubleshooting, and modernising the auth and security foundations so the platform can scale safely over the next few years. You will partner … partnership with the Security team, credential management, authentication strategies, and secure connection protocols that balance security with developer experience. Build and improve webhook reliability, observability, and usability. Deliver API and webhook logs and diagnostics tooling that helps partners and internal teams investigate and recover from issues faster. Ensure the auth ...

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Location
United Kingdom
operate agent workflow platform capabilities aligned to product-defined standards and interfaces, including traceability, state handling, and convergence patterns* Implement production-grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows* Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large-scale SaaS systems with production operations accountability* Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management* Strong hands-on technical leadership background in distributed systems and platform engineering* Deep experience with data platform engineering at scale, including ingestion ...

Regional Director, Sales, UKI

Location
Greater London, England, United Kingdom
Riverbed. Empower the Experience Riverbed, the leader in AIOps for observability, helps organizations optimize their users’ experiences by leveraging AI automation for the prevention, identification, and resolution of IT issues. With over 20 years of experience in data collection and AI and machine learning, Riverbed’s open and AI-powered … observability platform and solutions optimize digital experiences and greatly improve IT efficiency. Riverbed also offers industry-leading Acceleration solutions that provide fast, agile, secure acceleration of any app, over any network, to users anywhere. Together with our thousands of market-leading customers globally – including 95% of the FORTUNE ...

AI & Data Architect

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
will take deep ownership of one or more critical architecture domains — such as agentic application design, AI security and trust, AI operations and observability, data and knowledge engineering, or model platforms and inference — serving as the lead authority in your domain across client engagements. You will develop and maintain specialized … technical authority on your domains, you will lead architecture decisions and be accountable for ensuring systems meet rigorous non-functional requirements across security, observability, governance, performance, and scalability. You will produce and own the architecture artifacts that shape delivery — including architecture decision records (ADRs), component and data flow diagrams ...

Senior Product Manager - Platform

Location
Belfast City District, Northern Ireland, United Kingdom
THIS ROLE We're hiring a Senior Platform Product Manager to own two of our most critical platform areas: our API gateway and our observability platform. The API gateway is the public entry point for all Apex APIs, handling hundreds of millions of requests per month (and growing). Observability … with a track record of owning cloud platform infrastructure Good understanding of API concepts - latency, rate limiting, response codes Experience with Datadog or comparable observability tooling Demonstrated ability to set product vision and strategy, and communicate it credibly to technical and non-technical stakeholders Excellent prioritization judgment in ambiguous, high ...

Forward Deployed Strategist

Location
Greater London, England, United Kingdom
Dynamo’s products within enterprise environments Translate customer requirements into deployment strategies and implementation plans Help customers design scalable workflows around AI governance, evaluation, observability, and real-time guardrails Navigate ambiguity and unblock cross-functional execution across internal and customer teams Balance speed, technical feasibility, governance, and business impact during … from ambiguity to execution Nice To Have Experience with Generative AI, LLM deployments, AI governance, AI evaluation, or guardrail systems Familiarity with enterprise AI observability, monitoring, or security workflows Experience working with regulated industries such as financial services, healthcare, or government Background in management consulting, enterprise SaaS, or technical implementation ...

AI Engineer

Location
Rochdale, England, United Kingdom
shipping LLM features to real users. Fluent in Python or TypeScript; comfortable across both. Understand embeddings, retrieval, prompt design and evaluation. Rigorous about testing, observability and cost. Right to work in the UK. #J-18808-Ljbffr ...

AI AGENTS ENGINEER

Location
Manchester, England, United Kingdom
status, support history, and customer‐specific metadata. Design agent workflows that sit inside real task surfaces rather than generic chatbot experiences. Build evaluation and observability for agent behaviour, including tool‐call history, task success metrics, regression tests, failure modes, traceability, and guardrails. Collaborate with ML, product, hardware, client, and deployment … models or multimodal AI over images, video, inspection evidence, diagrams, screenshots, or technical records. Experience with RAG over structured and unstructured data. Familiarity with observability, log analysis, incident response, support tooling, runbooks, or developer tools. Experience building workflow UIs where AI assists a specific operational task. Familiarity with manufacturing, quality ...

Clickhouse Solutions Architect

Location
Slough, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...