3,476 to 3,500 of 3,862 Observability Jobs

Platform Engineer: AI Tools & Infra for Developers

Location
Rochdale, England, United Kingdom
Group is seeking an experienced engineer to build internal SDKs and gateways, and to own developer-facing tooling. You will contribute to shared evaluation, observability and cost-attribution tooling, and design multi-tenant data infrastructure with strong isolation. The role emphasizes platform engineering, distributed systems expertise, and ownership of tooling ...

Staff Infrastructure Engineer - Cloud Platform & Security

Location
Greater London, England, United Kingdom
projects from inception to production, write RFCs and ADRs, and mentor teammates while shaping future capabilities of Engine — with a strong emphasis on automation, observability and best practices. #J-18808-Ljbffr ...

Senior DevOps Engineer — Remote, Global Cloud Ops

Location
Belfast City District, Northern Ireland, United Kingdom
users, and data, across multiple regions. You will influence architecture, reliability, automation, and cloud efficiency, collaborating with engineers worldwide. The position emphasizes cost optimization, observability, and scalable deployment practices, with occasional travel to Belfast or remote collaboration. #J-18808-Ljbffr ...

Senior IP Network Engineer & SRE Lead

Location
Birmingham, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...

SRE Associate: Build Reliable, Scalable Systems

Location
Birmingham, England, United Kingdom
Birmingham seeks an Associate Site Reliability Engineer to join Core Engineering. You will help build and operate high-performing, scalable systems, focusing on reliability, observability, and automation across critical services in a fast-paced financial environment. You will collaborate with engineering teams to reduce downtime, improve deployment speed, and implement ...

Investment Tech Engineer — AI & Data Pipelines (Hybrid)

Location
United Kingdom
platform. You will work on cloud-based tools (Streamlit/ReactJS), apply AI-assisted tooling, and help drive CI/CD, data quality, and observability across the lifecycle. Collaborating with Cloud Teams, Enterprise Architecture and Security, you will translate business requirements into robust technical solutions, build data pipelines for analytics ...

Senior IP Network Engineer & SRE Lead

Location
Ipswich, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...

Operations Team Lead: Scale Reliability & Incident Mastery

Location
Stockport, England, United Kingdom
move from reactive firefighting to proactive reliability engineering. You will lead the Operations team, set standards, manage incidents, on-call rotations, and drive observability improvements. A strong SRE/DevOps background and calm, clear communication are essential. #J-18808-Ljbffr ...

MLOps Engineer — Real-Time AI in Production (Hybrid London)

Location
Greater London, England, United Kingdom
production systems, collaborating with research and engineering teams. The role emphasizes designing end-to-end ML pipelines, deploying and monitoring models in production, building observability and CI/CD, and enabling reproducible experimentation. #J-18808-Ljbffr ...

24/7 HPC Infra SRE for AI & GPU Compute

Location
Gloucester, England, United Kingdom
work across network, storage, virtualization and orchestration with hands‐on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance benchmarking. This role champions observability, automation and on‐call reliability, shaping next‐gen HPC platforms within a globally distributed team. #J-18808-Ljbffr ...

Senior ML Engineer - Conversational AI & GenAI | Hybrid

Location
Greater London, England, United Kingdom
full ML lifecycle from design to production, deploying models and building robust, scalable systems in collaboration with product managers. Strong focus on reliability, observability, and GenAI-enabled solutions. #J-18808-Ljbffr ...

Senior AWS Data Engineer: Build Secure Data Platforms

Location
Greater London, England, United Kingdom
solutions, establish engineering standards and patterns, and mentor other engineers. The role blends hands-on execution with technical leadership, enabling data quality, governance and observability while influencing #J-18808-Ljbffr ...

Lead SRE: Build Reliable, Scalable Systems

Location
Greater London, England, United Kingdom
reliability, scalability, and performance of our systems in a hybrid UK environment. You will partner with development and operations teams to automate workflows, improve observability, and drive reliability initiatives across production services. The role involves defining SLOs/SLIs, handling on-call rotations, and contributing to capacity planning while advocating ...

Scala Engineer — Core Streaming Platform (Hybrid)

Location
Greater London, England, United Kingdom
contribute to design discussions, participate in reviews, and help juniors grow. Expect ownership of features end-to-end, with emphasis on reliability, observability, and production readiness. #J-18808-Ljbffr ...

Senior IP Network Engineer & SRE Lead

Location
Greater London, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...

Senior AI/ML Engineer — Agentic LLM Systems (Hybrid, London)

Location
Greater London, England, United Kingdom
services that support AI-driven workflows. You will work with researchers and engineers to diagnose failures, optimize latency and cost, and contribute to observability and monitoring of production AI systems. #J-18808-Ljbffr ...

Senior Solutions Engineer: Your Partner in Technical Wins

Location
Greater London, England, United Kingdom
with Product Management and Engineering to refine demonstrations, participate in trials, and guide customers toward realizing continuous value from our cloud-native security and observability #J-18808-Ljbffr ...

Senior Fullstack Engineer, AI Platform & Scale

Location
United Kingdom
production-grade platform services using Node.js, TypeScript, NestJS and React, and you’ll work with LLMs, LangChain, LangGraph and agent orchestration to solve governance, observability and scalability #J-18808-Ljbffr ...

Finance AI Engineer: Build Production-Grade LLM Apps

Location
City Of London, England, United Kingdom
processes, translating CFO needs into secure, observable software within a delivery pod. The role emphasizes end-to-end engineering of agent-based workflows, evaluation, observability, and integration with ERP/EPM systems, with exposure to senior decision-makers. #J-18808-Ljbffr ...

Senior GTM Automation Architect - AI-Powered Pipelines

Location
Wedmore, England, United Kingdom
enrichment, and routing, while shaping prompting architectures used by sales teams. You will influence architecture, vendor relations, and production readiness, with a focus on observability and data quality. This is a hands-on, production‐oriented role in a remote-first company. #J-18808-Ljbffr ...

Senior AI Engineer: Production Agents & Orchestration

Location
United Kingdom
Principal Applied AI Engineer to design, build, and operate production AI agents and orchestrations for a web-native trading platform. You will own implementation, observability, and security while partnering with Product and Platform teams to deliver measurable AI systems in production. This hands-on role requires writing production code, establishing ...

Lead, Production Reliability & Ops

Location
Portsmouth, England, United Kingdom
Lead to own production and build scalable reliability. You’ll lead live customer-facing systems, shaping processes, incidents, and a team culture focused on observability and uptime. This hands-on role requires strong SRE/DevOps experience, proven leadership, and the ability to drive blameless postmortems and reliability roadmaps across ...

Senior AI Software Engineer - Hybrid, High-Impact Delivery

Location
Greater London, England, United Kingdom
secure production code, and mentor engineers and AI coding agents. The role blends deep software engineering with AI-native practices, overseeing code quality, testing, observability, deployment readiness, and governance while ensuring outputs are explainable and compliant with standards. #J-18808-Ljbffr ...

Senior Backend Engineer - AI Platform & Scalable Infra

Location
Greater London, England, United Kingdom
establish best practices for prompt engineering, AI guardrails, and evaluation methodologies, ensuring safe deployment, monitoring, and iteration of AI applications. A strong focus on observability and cost efficiency guides this role. #J-18808-Ljbffr ...

Embedded Quality Engineer: .NET Automation & CI/CD

Location
Swindon, England, United Kingdom
left practices. Join a hybrid role with 3 days in the office. You will shape test strategy, contribute to performance testing, and work with observability tools to improve release readiness and quality insights. #J-18808-Ljbffr ...