3,501 to 3,525 of 3,959 Observability Jobs

Senior Value Engineer – Remote: Growth & ROI Storytelling

Location
Pathhead, Scotland, United Kingdom
persuasive value stories. The role emphasizes storytelling with data, enabling GTM teams, and partnerships to articulate Grafana's value propositions, with a focus on observability, open source roots, and a global remote #J-18808-Ljbffr ...

Service Transition & Reliability Consultant

Location
Nottingham, England, United Kingdom
changed services land in production reliably and with robust operational foundations. You will partner with project teams during design, influence outcomes on reliability and observability, and own service readiness across the lifecycle. The role emphasizes hands-on engagement, continuous improvement, and collaboration with SRE, Product Engineering, and Service Operations ...

Senior Backend Engineer: Own architecture & resilient systems

Location
Bath, England, United Kingdom
production systems, aiming for scalable, resilient services and robust incident handling. Collaborating with Engineering, Product and Operations, you will drive CI/CD, observability and reliability improvements across a suite of business-critical platforms. #J-18808-Ljbffr ...

AI Platform Engineer: Kubernetes, Cloud & AI Agents

Location
Knutsford, England, United Kingdom
Manchester. You will build scalable Kubernetes platforms, automate cloud infrastructure, and deploy AI agents with governance and monitoring. The role emphasizes secure coding, observability, cost-aware cloud usage, and self-service tooling to enable rapid delivery across teams and business units. #J-18808-Ljbffr ...

Lead Software Engineer - Hybrid, Share Options, £80k+

Location
Wigan, England, United Kingdom
production. You will lead a small cross-functional team, delivering features in a fast-paced SaaS environment. You will help modernise microservices, improve observability, and enhance testability using modern architectural approaches. Hybrid working with two days in the office is available. #J-18808-Ljbffr ...

Senior AI-Powered Backend Engineer

Location
United Kingdom
workflows and decision paths. The Agent Systems team designs inference pipelines and back-end services that orchestrate AI across internal APIs, maintaining reliability and observability as the platform scales. You will prototype quickly, validate agent behaviors, and productionize what proves valuable, contributing to a 0→1 agent platform with cross ...

Senior Data Platform SRE — Hybrid, Obs & Reliability Leader

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer to embed within the Data Engineering team. You will own reliability, performance, and operability of the data platform, building observability from the ground up and leading incident response for data outages. You'll balance feature velocity with system stability, applying SRE practices and capacity planning ...

Senior Mobile Engineering Leader (iOS & Android)

Location
United Kingdom
collaborate with product, design, architecture, security, quality, and operations to deliver reliable, scalable mobile experiences and roadmaps. You will shape architecture, improve release quality, observability, and AI-enabled development, while building a high-performing organization through hiring, mentoring, #J-18808-Ljbffr ...

Senior Platform Engineer - Apollo Core & DevEx

Location
England, United Kingdom
Lead Analytics Engineer and reporting to the Head of Data & Engineering. You’ll lead platform and application engineering, drive architectural decisions, and guide security, observability and deployment pipelines. #J-18808-Ljbffr ...

Senior Quality Engineer - Champion Quality in Cloud Apps

Location
Greater London, England, United Kingdom
perform exploratory testing, and collaborate with Product and Development to build the right software for customers, embracing shift-left quality practices. You will enhance observability, apply risk-based testing, and contribute to automation while growing a culture of quality across teams. #J-18808-Ljbffr ...

Senior SRE: Cloud Platform Reliability & Automation

Location
Cambridge, England, United Kingdom
Engineer to ensure the reliability, availability, and performance of our large-scale cloud platforms and SaaS products. Your work will automate operational tasks, improve observability, and contribute to incident management while collaborating with development teams to build more reliable and scalable applications across regions. #J-18808-Ljbffr ...

Service Transition Lead: Design, Reliability & Readiness

Location
United Kingdom
changed services land in production reliably and with strong operational foundations. You will partner with project teams from design through deployment, focusing on reliability, observability and governance. The role emphasizes hands-on readiness, collaboration with SRE, Product Engineering and Service Operations, and continual improvement of service transition processes within ...

Cloud Operations Engineer - Reliability & Automation

Location
England, United Kingdom
premises and AWS cloud services. You will work with MySQL, PostgreSQL and SQL Server, ensure security, availability and disaster recovery while driving automation and observability in a collaborative team. You will collaborate with Software Engineers and Technology teams to deliver reliable platforms and support product delivery, with training opportunities ...

Senior SET: Platform Quality Engineer (Hybrid, London)

Location
Greater London, England, United Kingdom
architecture level, automate validation across environments, and manage delivery risk in distributed systems. You will define standards for testing distributed systems, enable observability-driven debugging, automate validation of availability and latency, and contribute to security posture. #J-18808-Ljbffr ...

Senior Data Platform Engineer | ELT & Analytics

Location
Greater London, England, United Kingdom
senior member, you will influence platform architecture, data products, and engineering standards while collaborating with analytics teams to deliver reliable production solutions and observability across key datasets. #J-18808-Ljbffr ...

SRE Lead: Data, Cloud & Developer Experience

Location
Greater London, England, United Kingdom
Blackstone is seeking a Site Reliability Engineer to lead the adoption of SRE practices, improve observability, and ensure reliable services across the firm. The role involves instrumentation, monitoring, and automation to reduce toil and incident impact. The successful candidate will collaborate with development and operations teams to design resilient systems ...

Site Reliability Engineer - Live Ops & Cloud Resilience

Location
Greater London, England, United Kingdom
experienced Site Reliability Engineer to design, build, and operate resilient, secure platforms underpinning our digital and live operations. You’ll focus on reliability, observability, automation, and disaster recovery across hybrid environments, collaborating with engineering, operations, and project stakeholders. The role emphasizes improving service availability, incident response, and continuous improvement ...

Platform Engineering Lead: GPU Infra & Kubernetes

Location
Greater London, England, United Kingdom
Volta is seeking a hands-on Platform Engineering leader to own reliability, observability, and security across a Kubernetes-native compute platform. You will lead a team, set technical direction, and translate product requirements into scalable platform features while collaborating with security and cross-team leads. This role requires 5+ years ...

Remote D365 F&O Data Engineer for Azure Pipelines

Location
United Kingdom
outside IR35. You will build scalable Azure/Databricks data pipelines to support a global F&O deployment and drive data governance and observability practices. Ideal candidates will have strong D365 F&O data experience, deep Azure expertise, and hands-on Databricks experience. #J-18808-Ljbffr ...

Production Reliability Lead

Location
Norwich, England, United Kingdom
team to move from firefighting to proactive reliability engineering. This hands-on role requires leading SRE/DevOps practices, defining SLIs/SLOs, improving observability, and ensuring clear ownership with blameless accountability. #J-18808-Ljbffr ...

AI-Ops SRE & Operations Leader

Location
Greater London, England, United Kingdom
experienced SRE Manager to lead the reliability function for production services used by internal and external customers. You will drive AI-Ops adoption, automation, observability and incident response. You will own end-to-end incident and problem management, coach team leads, ensure RCAs and post-mortems are completed, and balance ...

Field-Ready AI Engineer: Deploy & Scale Agentic AI

Location
Crawley, England, United Kingdom
production across diverse platforms. The role focuses on building production architectures, APIs, CI/CD pipelines and security practices, with emphasis on reliability, observability and scalable performance in scientific and enterprise contexts. #J-18808-Ljbffr ...

Cloud Native SWE II: Kubernetes & OpenTelemetry

Location
Manchester, England, United Kingdom
platform. You will work with a team of engineers defining what it means to be cloud native, contributing to components and integrations across Kubernetes, observability and automation tooling. Under guidance from senior engineers, you will deepen your knowledge of the Cloud Native ecosystem, help bring Couchbase everywhere ...

Operations Team Lead — Production Reliability & Scale

Location
Cambridge, England, United Kingdom
seeking an Operations Team Lead to own production and scale systems. You will lead operational excellence across live customer-facing platforms, ensuring reliability, observability, and proactive improvements. This hands-on role involves shaping processes, guiding incidents, building the team, and moving from firefighting to sustainable reliability engineering. The ideal candidate ...