3,001 to 3,025 of 4,043 Permanent Observability Jobs

Observability Platform Engineer — Hybrid London

Location
Greater London, England, United Kingdom
N Consulting Limited is looking for a Platform Engineer to join their team in London, UK. This hybrid position requires hands-on experience with OTel collectors, telemetry pipelines, and strong knowledge of cloud platforms. The ...

Cyber Platform Engineer - Scale & Observability

Location
Belfast, Northern Ireland, United Kingdom
JobTailor in Belfast is seeking a senior engineering leader to design, deliver and operate production-grade capabilities that power cyber operations at scale. You will build reliable, secure and observable services, tools, platforms, or data ...

Site Reliability Engineer: Automation & Observability

Location
Greater London, England, United Kingdom
Apple Inc. is seeking a Site Reliability Engineer in London to join the Apple Services Engineering team. You will help sustain large-scale services powering the App Store, Apple Music, TV, Podcasts and Books for ...

AI-Powered Observability Tech Lead

Location
Greater London, England, United Kingdom
Cisco’s Collaboration Technology Group in London seeks a Technical Leader to drive architectural vision and implement an AI-powered Production Intelligence platform. You will blend Site Reliability Engineering with agentic AI to improve monitoring ...

Lead Site Reliability Engineer - Observability & Resilience

Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase & Co. seeks a Lead Site Reliability Engineer to define the future of reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers ...

Product Manager, XDR & Observability Intelligence

Location
Greater London, England, United Kingdom
Cato Networks is seeking a Product Manager to define and build the intelligence layer of our SASE platform. You will own how Cato converts unified network and security telemetry into high-confidence detections, actionable insights ...

Product Manager - XDR & Observability

Location
Greater London, England, United Kingdom
pain points, validate priorities, and ensure measurable customer impact. Requirements: 4+ years of B2B product management experience in cybersecurity, SIEM, XDR, observability, or IT operations platforms. Cross functional work with mutliple teams across R&D, Marketing, Sales, Customer Success Hands-on familiarity with SOC workflows, including alert triage, investigation processes ...

Lead AI Engineer – Agentic Systems & Production ML

Location
Greater London, England, United Kingdom
GlobalLogic is seeking a Senior/Lead AI Engineer to own the development of agentic frameworks, evaluation pipelines, and observability tooling for safe, scalable enterprise AI. You will lead architectural directions, mentor engineers, and ship production-grade, LLM-backed applications. The role requires 5+ years in software engineering, hands … experience with agentic or LLM-backed systems, and strong Python, observability, and cloud-native skills. Hybrid model offered. #J-18808-Ljbffr ...

MLOps Engineer

Location
Rochdale, England, United Kingdom
build the runway. Training pipelines, model registries, feature stores, evaluation, deployment, observability and cost governance - you turn ML from artisanal to industrial across every AutoThink product. What you'll build A shared ML/LLM platform used by every AutoThink product team. CI/CD for models, prompts and evaluations. … Cost and latency observability across every model call. Secure, multi-tenant inference infrastructure. Requirements 5+ years platform/infra/DevOps, ideally with an ML flavour. Strong Kubernetes, Terraform, CI/CD. Comfortable with GPU workloads and inference optimisation. Security and compliance mindset (SOC2, ISO 27001, UK GDPR). Right ...

Boomi Solutions Lead

Hiring Organisation
Hackajob Ltd
Location
Farnborough, Hampshire, South East, United Kingdom
Employment Type
Permanent
global IT services company. Job Description: Job Title: SME Dynatrace Solutions Location: UK Hybrid About Us We help organizations accelerate digital transformation through intelligent observability and automation. As a strategic partner of Dynatrace, we deliver cutting-edge solutions that empower enterprises to optimize performance, enhance user experience, and drive innovation. … role: We are seeking a dynamic and experienced Dynatrace SME to join our team and lead strategic engagements around Dynatraces observability platform. This role is ideal for someone who thrives in a client-facing environment, understands enterprise IT ecosystems, and excels at uncovering business needs to deliver tailored solutions. ...

Boomi Solutions Lead

Hiring Organisation
Hackajob Ltd
Location
Cove, Devon, UK
make your application promptly. Job Description: Job Title: SME – Dynatrace Solutions Location: UK Hybrid About Us We help organizations accelerate digital transformation through intelligent observability and automation. As a strategic partner of Dynatrace, we deliver cutting-edge solutions that empower enterprises to optimize performance, enhance user experience, and drive innovation. … role: We are seeking a dynamic and experienced Dynatrace SME to join our team and lead strategic engagements around Dynatrace's observability platform. This role is ideal for someone who thrives in a client-facing environment, understands enterprise IT ecosystems, and excels at uncovering business needs to deliver tailored solutions. ...

Fullstack Engineer (DV Clearance) - Cloud Native, 37.5h

Location
Cheltenham, England, United Kingdom
Java applications, with cloud native deployments in OpenShift and Kubernetes. You will work across distributed architectures, implement Web API integrations, and contribute to observability and security practices. A DV clearance or recent status is required. #J-18808-Ljbffr ...

Managing Engineer – Cyber Platform Engineering

Location
Belfast, Northern Ireland, United Kingdom
secure, and observable services, tools, platforms, or data pipelines aligned to shared engineering standards. Own service and platform lifecycle expectations, including reliability, operational readiness, observability, and continuous improvement. Partner across product, platform, and engineering stakeholders to shape architecture, integration design, and delivery sequencing. Drive platform maturity through automation, reuse, standardization … experience leading engineering teams or technical delivery in product-based environments. Experience operating production systems, services, tools, or pipelines with strong expectations around reliability, observability, security, and maintainability. Strong understanding of integration patterns, automation, cloud-native engineering practices, and scalable platform design. Experience building reusable services, workflows, platforms, or data ...

GenAI Platform SRE Lead: Reliability, Automation & Scale

Location
United Kingdom
security and scalability of the Oasis Platform. You will lead a high-performing team at the intersection of Operations and Platform Engineering, driving automation, observability, and DevOps practices to improve platform resilience at scale. The role requires senior expertise in cloud platforms, Kubernetes, SRE principles and hands-on leadership. #J ...

Lead SRE: AWS & Python for Scalable Reliability

Location
United Kingdom
engineering and product teams to embed reliability into the SDLC, drive incident response, build automation, mentor junior engineers, and champion proactive reliability improvements and observability across platforms. #J-18808-Ljbffr ...

Cloud Data Platform Engineer III — AWS & Databricks

Location
Glasgow, Scotland, United Kingdom
deliver scalable, secure, and production-ready infrastructure that powers enterprise analytics. The role emphasizes IaC, automation, and secure coding practices, with focus on observability and cost optimization across the platform. #J-18808-Ljbffr ...

Lead Analytics Engineer

Location
Greater London, England, United Kingdom
their platform: from APIs and pipelines to governance and observability. Greenfield impact – Inherit a live but early function, define best practice across modeling, testing, observability, and governance. Direct product impact – Your models and marts power decisions and features that 2M+ consumers rely on every day. AI at the core – enable … standards. Build the golden layer : design high-trust dbt marts and a documented semantic layer that powers product features Ship with reliability : stand up observability (tests, freshness, lineage, cost), enforce SLAs, and keep pipelines fast and resilient.. Be the multiplier : mentor analysts/engineers as the team grows, partner with ...

Senior AI-First Cloud Software Engineer

Location
Manchester, England, United Kingdom
solve complex business challenges. The role emphasizes architectural leadership, production-grade code, and strong testing, with AI-assisted development to drive automation and observability across scalable platforms. #J-18808-Ljbffr ...

Solutions Engineering Manager

Location
Greater London, England, United Kingdom
Manager, Solutions Engineering to lead our pre-sales Solutions Engineering team in the UK & Ireland (UKI), with a primary focus on the SolarWinds Core & Observability portfolio, encompassing both SaaS and Self-Hosted deployment options. As a front-line SE leader, you will be responsible for managing a team of Solutions … Engineers and Solutions Architects, partnering with Sales Directors and Account Executives, and leading the technical pre-sales strategy for Observability and the broader SolarWinds portfolio across the UKI region. You will act as a coach, culture carrier, and execution owner ensuring your team delivers outstanding customer outcomes while developing their ...

Senior Cloud Platform Engineer | Azure, IaC & Automation

Location
Monkston, England, United Kingdom
services. You will lead cloud infrastructure deployment, automation, and IaC initiatives to deliver secure, scalable and highly available solutions. You will mentor engineers, drive observability and security practices, and support platform modernization across high-availability environments in a fast-paced, collaborative team. #J-18808-Ljbffr ...

Backend Engineer (Go/Python) for Production Microservices

Location
Greater London, England, United Kingdom
decision intelligence platform. You will develop Go and Python microservices, optimize database and API performance, and contribute to scalable system design, deployment tooling, and observability, working with AI, frontend, and product teams to deliver fast, reliable solutions. #J-18808-Ljbffr ...

Principal Machine Learning Engineer

Location
Greater London, England, United Kingdom
data engineering teams to implement scalable data lakehouse oriented feature architectures and enterprise‐grade ML governance. Champion engineering standards for model quality, documentation, observability, and platform resilience. Feature Engineering & Data Architecture Architect highly scalable, production‐ready feature pipelines within Lakehouse environments. Set the technical direction for fallback and resilience strategies … including scoring metrics, latency, error analytics, and SLOs. Partner with platform teams to optimise cost, scale, and reliability of inference endpoints. Monitoring, Drift Detection & Observability Define observability standards for feature drift, concept drift, performance degradation, and data integrity. Lead the creation of dashboards, benchmarks, and automated alerting across ...

Lead Java Developer — Real-Time Risk & Cloud (Hybrid)

Location
Greater London, England, United Kingdom
full lifecycle from design to production support, integrating new analytics and data sets across global teams. The role emphasizes scalable microservices, streaming data, and observability with ELK, Prometheus and Grafana. Hybrid work model and competitive benefits are offered. #J-18808-Ljbffr ...

Senior Cloud Infrastructure Engineer - AWS SaaS Platform

Location
United Kingdom
tooling, working with a Linux-based stack (Ubuntu) and a distributed services model. You’ll help evolve the platform with a focus on reliability, observability, and scalable storage while integrating with internal teams and incident response processes. #J-18808-Ljbffr ...

Lead Java/Go Engineer - Enterprise Automation Platform

Location
Bournemouth, England, United Kingdom
across AaaS, Ansible Automation Platform, AutoM8, and AI-enabled service management. You will work within an agile team, shaping platform tooling, event-driven automation, observability, and cloud readiness to improve reliability and scale across #J-18808-Ljbffr ...