3,076 to 3,100 of 4,044 Permanent Observability Jobs

Hybrid Senior Backend Engineer - Golang & AWS Serverless

Location
Metropolitan Borough of Solihull, England, United Kingdom
Platform teams. You will own API design, implement event-driven components, and contribute to CI/CD pipelines while maintaining high standards of observability, testing, and performance. #J-18808-Ljbffr ...

Principal Java Engineer: Low-Latency Payments, Hybrid UK

Location
Greater London, England, United Kingdom
Hybrid role in the UK with modern tech stack and real-time services. You will own design decisions, mentor peers, and drive performance and observability enhancements across services. The ideal candidate has 10+ years in Java, SpringBoot, microservices, and cloud deployments (OpenShift/Kubernetes, AWS). #J-18808-Ljbffr ...

Senior Platform Engineer: Kubernetes & CI/CD Leader

Location
Cambridge, England, United Kingdom
deploy and scale workloads. You’ll own day-to-day delivery across our on-premise estate, driving platform stability, self-service tooling, and robust observability to help engineering ship value safely and rapidly to production. #J-18808-Ljbffr ...

Senior Backend Engineer: Architecture, Azure & .NET (Hybrid)

Location
Nottingham, England, United Kingdom
features with resilience, security, and cost-efficiency in an agile environment. You’ll lead code reviews, mentor engineers, and contribute to CI/CD, observability, and testing, while collaborating across product and tech teams to deliver real-world outcomes. #J-18808-Ljbffr ...

Agent Platform Engineer: AI-Driven Automation & API Design

Location
Greater London, England, United Kingdom
across relational and NoSQL databases. You will collaborate across engineering, product, and design to deliver a scalable platform from concept to production, contributing to observability and reliability for crypto market safety. The ideal candidate has 5+ years of engineering experience, strong Python skills, and hands-on experience with LangSmith ...

Backend Engineer – AWS CDK & Edge Platform (London)

Location
City Of London, England, United Kingdom
contract with hybrid work arrangements, offering exposure to high-scale distributed systems for a global streaming platform. The team focuses on automated traffic routing, observability, and AI tool hosting. You will develop resilient backend microservices, enable AI tool integrations, and improve performance with low-latency client routing within an experienced ...

Senior Full Stack & DevOps Engineer – Hybrid E‐commerce

Location
Rotherham, England, United Kingdom
responsibilities in a hybrid Rotherham-based role. You will build customer-facing front ends, APIs and services while owning CI/CD, testing, and observability infrastructure. The position emphasizes AI-assisted development tools and strong collaboration with stakeholders across marketing, accounts and internal systems. Candidates should have 5+ years ...

Senior Azure Platform Engineer — Remote (London, 1 day/wk)

Location
Greater London, England, United Kingdom
Azure platforms, and collaborate with multiple teams to share and reuse best practices and technology stacks. You’ll lead CI/CD, automation, and observability efforts, contribute to the codebase, and perform peer reviews while owning platform engineering tooling and processes. #J-18808-Ljbffr ...

Senior Java Engineer – Real-Time Banking Platform

Location
Greater London, England, United Kingdom
production and cloud-native services using Java/Spring Boot. You’ll design, build and operate robust distributed systems with emphasis on reliability, observability and operational excellence, and mentor peers while collaborating across autonomous teams to align engineering with #J-18808-Ljbffr ...

AI-Driven SRE & Software Engineer

Location
Manchester, England, United Kingdom
bet365 Group seeks a Site Reliability Engineer to improve reliability, observability and performance across critical systems. You will implement service instrumentation with OpenTelemetry-inspired practices, enhance logging, and drive AI-native approaches to reduce toil. You will collaborate across teams, contribute to SLIs/SLOs, and help shape a culture ...

Senior Azure Platform Engineer — Multi-Tenant AKS

Location
England, United Kingdom
lead a shift toward cloud-native, asynchronous architectures. You will act as SME, bootstrap platform team, define operating model, upskill engineers, and implement OpenTelemetry observability, GitOps pipelines, and policy-as-code. This is a high-impact, fintech-adjacent role with strong growth potential. #J-18808-Ljbffr ...

Hybrid Junior DevOps Engineer | AWS, Terraform, CI/CD

Location
Greater London, England, United Kingdom
offers a hybrid setup with a focus on learning and rapid impact. You’ll collaborate with in-house developers and AI Engineering to improve observability, security, and deployment reliability while advancing your Terraform and AWS skills in a fast-paced scale-up #J-18808-Ljbffr ...

Lead Cloud Architect — Multi-Cloud AI Platform

Location
Greater London, England, United Kingdom
secure, scalable deployments and guide the team through architectural decisions that accelerate delivery in live environments. You will champion DevEx, CI/CD, and observability while collaborating across disciplines to ensure pragmatic, high-quality engineering and rapid #J-18808-Ljbffr ...

Senior Network Engineer — Remote/Hybrid (Multi-Region DC)

Location
Greater London, England, United Kingdom
regions while modernizing automation with Ansible and Terraform. The role offers flexible remote/hybrid work with relocation support and a strong focus on observability using Prometheus and Grafana. Remote-friendly with international scope. #J-18808-Ljbffr ...

Lead AI Platform Engineer - Secure, Scalable Data Infra

Location
Greater London, England, United Kingdom
safe and scalable AI adoption. You will collaborate with AI/ML research teams, replace manual operations with self-service workflows, and implement robust observability and security by default. This role offers ownership of production systems in a regulated financial setting. #J-18808-Ljbffr ...

Technical Account Manager

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
team rewriting the rules of observability. A platform at the bleeding edge, empowering businesses to understand and act on their data through better observability, all in real time. Innovation removes the need for unnecessary indexing, cutting costs and complexity. Truly end to end, the platform encompasses everything; from logs … This is a high-impact role that demands serious technical firepower. You'll need deep experience with Cloud native tooling, hands-on knowledge of observability tools e.g. Grafana, DataDog or Splunk, and the ability to troubleshoot containerised environments like a pro. You have the technical knowledge and the confidence ...

Senior Software Engineer, AI-Powered Expense Platform

Location
Greater London, England, United Kingdom
interfaces with AI components, focusing on accuracy, latency, and scalable data flows. You will work across international teams, mentor others, and ensure robust observability, auditing, and compliance in financial contexts. Based in London or willing to relocate, you will join a fast-paced, tech-driven environment. #J-18808-Ljbffr ...

Applied AI Engineer New

Location
Greater London, England, United Kingdom
built quickly and as separate systems. The next challenge is to bring these approaches together: build reusable agentic infrastructure, establish a robust evaluation and observability layer, and create systems that allow us to automate new workflows quickly and reliably as Dwelly scales. This is not an AI research role. … loops. Move us from one-off AI solutions toward reusable infrastructure where new workflows can be introduced quickly and with predictable reliability. 2. Evaluation & observability Build the evaluation framework that allows us to understand how our agents perform and why they succeed or fail. Make testing, tracing, debugging, and evaluating ...

Director, Technical Account Management

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
growing services revenue; you position TAM as a value driver, not a cost center. Technically fluent across infrastructure, cloud platforms (AWS, Azure, GCP), observability, and monitoring, credible with practitioners and C-level executives alike. A strategic operator who pairs vision with execution: you set direction, build the systems to support … region. Bonus Points: Experience redesigning organizational structures mid-growth: building new team shapes or specializations as a business scales. Track record evangelizing cloud adoption, observability, or security to C-level and board-level stakeholders across EMEA.Hands-on experience with Datadog or other leading cloud monitoring and observability platforms. Multilingual ...

Senior Connectivity Engineer / Network Engineer

Location
Wallingford, England, United Kingdom
WireGuard/Tailscale or equivalent): access‐as‐code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet‐wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end‐to‐end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel‐level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL‐as‐code patterns, and carrier/router API integrations. #J-18808-Ljbffr ...

Principal GenAI Platform Architect (Full-Stack)

Location
York and North Yorkshire, England, United Kingdom
back-end work, designing scalable GenAI features and integrating MCP services. You will lead architecture and rollout of autonomous AI agents, ensure robust observability, and collaborate with product and design teams to turn vision into production-ready systems. The role demands hands-on experience with React/TypeScript ...

Senior/Staff Software Engineer (Nova Core)

Location
Greater London, England, United Kingdom
operations across Nova Cloud deployments. This role focuses on the Nova Core “inner loop”: service architecture, APIs, data models, persistence, authn/authz, observability, and developer experience that other Nova modules and product teams depend on. What you’ll do Own and ship critical Nova Core backend services (e.g., common … engineering and product teams. What success looks like Core Nova services are delivered, adopted, and operated reliably with clear SLIs/SLOs and runbooks. Observability is strong enough that incidents are detected quickly and resolved faster over time (improving MTTD/MTTR). API versioning and compatibility practices reduce integration ...

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Location
Greater London, England, United Kingdom
operate agent workflow platform capabilities aligned to product‐defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production‐grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large‐scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands‐on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Senior/Staff Software Engineer

Location
Greater London, England, United Kingdom
tooling, optimizing for reliability, latency, cost, and debuggability in production. Build and maintain the surrounding infrastructure: data pipelines, evaluation harnesses, prompt and model management, observability, and safety/guardrails. Work across the stack—from backend integrations and APIs to simple UI hooks—to deliver complete AI features, not just model … workflows quickly, then harden what works. Own, downscope, ship, iterate: one clear owner per feature, from prototype to production. Fundamentals done well: evaluation, observability, and safety are part of the first version, not an afterthought. Competitive salary and meaningful equity. Health, dental, and vision coverage. Flexible time off and support ...