1,351 to 1,375 of 4,717 Permanent Observability Jobs

Staff Platform Engineer

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
capabilities, reusable golden paths, and standardised service templates for efficient software delivery. Provide technical direction across platform domains including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Enable AI-assisted engineering workflows to enhance productivity, automate routine activities, and improve decision-making. Embed governance, security … compliance controls through policy-as-code and auditable platform practices. Improve platform reliability through SLOs, observability frameworks, resilience engineering, and incident-driven improvements. Collaborate with product, engineering, security, and architecture teams to translate business needs into scalable platform solutions. Promote engineering excellence through automation, standardisation, and efficiency-driven practices, including ...

Lead Software Engineer - Platform

Location
Greater London, England, United Kingdom
need to scale accordingly, and our platform foundations need to support faster, more reliable delivery. You’ll be central to making that happen – from observability and developer experience to the infrastructure patterns that underpin everything we ship. What you’ll do Alongside the other Lead Engineers you’ll support … infrastructure. You’ll be responsible for the reliability, scalability, and operability of our systems. That means CI/CD pipelines, infrastructure‐as‐code, observability, incident response, and the day‐to‐day health of production. You’ll make sure we can ship with confidence and sleep at night. Shape technical direction. ...

Senior AI Product Engineer

Location
Greater London, England, United Kingdom
more junior engineers through pair programming, code review, and design feedback. Raise the engineering bar across the team by promoting good practices in testing, observability, and AI system reliability. Influence cross-team decisions on how AI capabilities integrate with the rest of the Elliptic platform. What you will achieve … technical direction of an AI workstream, including architecture, evaluation, and rollout. Established or improved at least one team practice for building AI systems (evals, observability patterns, prompt management, rollout safety). Mentored junior engineers on AI engineering practices and contributed to their growth. Built strong working relationships across product ...

ML Ops Engineer

Hiring Organisation
CMC Markets
Location
London, UK
Employment Type
Full-time
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services. You'll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams. This is not a research role. … production issues across model, application, infrastructure and critical data-dependency layers. Reliability, security and engineering qualityImprove system robustness, scalability and cost efficiency through automation, observability and infrastructure as code. Write production-grade Python for long-running services, deployment tooling and ML workflows. Establish testing, validation, release and incident-management practices ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
London, UK
Employment Type
Full-time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering … solving, prioritization, and decision-making skills. What We're Looking For: Basic Required Qualifications:10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics ...

Junior Azure Engineer

Location
Greater London, England, United Kingdom
containers Implement security and governance controls (RBAC, Azure Policy, Management Groups) Build and support landing zones and foundational cloud environments Manage monitoring and observability using Azure Monitor, Log Analytics, and alerts Support CI/CD pipelines using Azure DevOps or GitHub Actions Automate operational tasks using PowerShell, Azure … experience with Infrastructure-as-Code (Bicep, ARM, or Terraform) Solid understanding of Azure networking and cloud architecture principles Experience with monitoring, logging, and observability tools Ability to troubleshoot and resolve complex cloud issues Experience with automation and scripting (PowerShell, Azure CLI) Strong collaboration and communication skills Experience with Azure Kubernetes ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
scalability. Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions. Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively. Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling, and improve infrastructure efficiency. Improve … Kubernetes operational experience (on-prem and AWS EKS).Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows. Deep understanding of observability and telemetry (monitoring, logging, tracing).Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD.Scripting proficiency in Python, Bash, or Go. Proven ...

Senior / Principal Applied AI Engineer (UK / Europe, Remote)

Location
United Kingdom
Core Platform Engineering. This is a hands-on engineering seat, not an advisory one. You write and own production code, you put evaluation and observability on everything you ship, and you run autonomously in a lean team. You report to the COO and partner closely with the CTO, the Head … commodity internal needs; build or replace it internally when it’s differentiating or becomes cost-prohibitive. Ship outcomes, not architecture debates. Put evaluation, observability, and cost guardrails on everything - golden datasets, eval harnesses, tracing, fallback chains, latency and spend controls. Nothing ships as an unmeasured demo. Operate inside our security ...

Engineering Manager - Grafana Application Security | EMEA | (Remote, UK)

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Software Engineer, Senior

Hiring Organisation
Infor
Location
Orpington, Greater London, UK
Employment Type
Full-time
integration patterns using modern cloud-native approaches. Evaluating and implementing emerging AI integration technologies, frameworks, and engineering practices. Helping establish engineering standards for security, observability, resiliency, and maintainability. Contributing to technical design discussions and architecture reviews. Mentoring and supporting other engineers as AI capabilities become more broadly adopted across … applications and services. Experience with cloud-native development on AWS, Azure, or similar cloud platforms. Strong understanding of software architecture, security, scalability, reliability, and observability principles. Experience working with event-driven architectures and asynchronous integration patterns. Experience working with relational databases and modern data access technologies. Experience developing reusable frameworks ...

Platform Engineer

Hiring Organisation
Springer Nature
Location
London, United Kingdom
Salary
£ 70 K
engage on providing a unified and standardised platform by using modern and open standards. Our approach is encompassed by defined core capabilities, such as observability, continuous integration, security and storage. We therefore closely collaborate in our department so that the core capabilities are not only tightly integrated but also provide … Engineering department at Springer Nature Technology, we provide platform engineering expertise to support core capabilities and use cases across run time to databases to observability, ci/cd, and SRE. You will join a multidisciplinary team with different nationalities, backgrounds and experience levels. We are a very distributed department, sometimes ...

Senior Platform Reliability Engineer

Location
Greater London, England, United Kingdom
platforms, ensuring services meet defined availability, performance, security, and compliance standards. It combines deep operational expertise with strong automation, Infrastructure as Code (IaC), and observability capability to reduce toil, improve recovery, and enable predictable service outcomes. What you will be doing Deliver standards for availability, latency, performance, capacity, and scalability. … closed. Drive infrastructure-as-code and automation across Azure and co-lo environments. Evolve image bakery pipeline for secure, repeatable server images. Embed observability using metrics, logs, traces, and alerting tools. Partner with SRE and helpdesk teams to deliver service. Oversee automated patching, vulnerability remediation, and configuration compliance. Introduce KPIs ...

Software Engineer (Next.js Playwright)

Location
Greater London, England, United Kingdom
maintain applications within Azure. Contribute to CI/CD pipelines using GitHub Actions and Azure DevOps. Monitor application performance and reliability using modern observability tools. Collaborate Across the Business Work closely with clinicians, product managers and fellow engineers to understand user needs. Contribute ideas that improve both the product … Experience with Playwright or other automated testing frameworks. Experience working within Azure. Experience in healthcare, NHS, or regulated environments. Familiarity with application monitoring and observability tools. Experience working in a SaaS or scale-up environment. Mindset Product-minded and user-focused. Pragmatic, collaborative and delivery-oriented. Takes ownership and sees ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Live Operations Associate Roku, Inc.

Location
Cambridge, England, United Kingdom
role sits at the intersection of live operations and platform reliability, so you will work alongside infrastructure, engineering teams, and service providers to support observability, automation, and encoding pipeline health as the FAST ecosystem scales. As a trusted point of contact for service partners, you will manage expectations, communicate clearly … resolved efficiently Managing ongoing FAST channel configurations, including schedule updates, metadata changes, feed ingestion, and EPG alignment Monitoring channel and pipeline health using observability tools; validating ingestion, triaging transcode and CDN delivery issues, and escalating incidents with clear context to engineering or peer technical teams Driving automation of repetitive operational ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Forward Deployed AI Engineer

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
enabled systems. You'll bring deep expertise across modern full-stack technologies (.NET, Azure, SQL, React/Angular), along with experience in distributed systems, observability, and AI tooling such as LLMs, retrieval pipelines, agentic workflows, and platforms such as Anthropic Claude. Experience designing and deploying AI agents, leveraging Model Context … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Senior Specialist, Production Services Application Support Analyst

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Analytical and problem-solving skills with the ability to identify root causes and implement sustainable ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelinesImplement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloadsTune GPU and accelerator capacity, autoscaling, and cost efficiency ...

Applied AI Engineering Lead - VP, Markets Operations

Location
Greater London, England, United Kingdom
information Build and enhance robust AI services and infrastructure using modern engineering practices, including APIs, event‐driven patterns, CI/CD, Infrastructure-as-Code, observability, and automated testing Partner with AI researchers, data scientists, and software engineers to translate emerging AI capabilities into practical, reliable, and compliant enterprise applications Establish … data engineering concepts, ETL and data pipelines, structured and unstructured data, and integration with enterprise data platforms Experience with CI/CD, automated testing, observability, production monitoring, and operational readiness practices Familiarity with Infrastructure-as-Code solutions such as Terraform and cloud or container‐based deployment patterns Working knowledge ...

Senior Software Engineer

Hiring Organisation
Daniel James Resourcing Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
engineers. Where AI enters the product itself, youll help establish the controls required to take it into production responsibly, including validation, guardrails, permissions, observability, evaluation, failure behaviour and appropriate human oversight. What you'll bring Youll be an accomplished Software Engineer who has designed and delivered complex production systems … data modelling Docker, Kubernetes and Infrastructure as Code Event-driven architecture CI/CD, automated testing and modern engineering practices Production ownership, reliability and observability Technical leadership and mentoring other engineers The underlying brief prioritises strong C#/.NET experience but can consider engineers from another modern backend language ...

Senior Associate, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile (Scrum ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real‐time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Lead Software Engineer

Hiring Organisation
StepChange Debt Charity
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
practices while fostering a culture of ownership, collaboration, and continuous improvement. You'll also champion operational excellence by improving developer onboarding, enhancing monitoring and observability, supporting incident management and root cause analysis, and driving improvements in reliability, productivity, and delivery velocity. Alongside this, you'll work with stakeholders across … strong understanding of RESTful APIs, relational databases, Git-based development workflows, and software engineering best practice. Experience with CI/CD pipelines, GitHub Actions, observability tooling, and cloud-native architectures will enable you to drive quality and consistency across the engineering function. You'll thrive in an Agile environment, balancing ...

Senior Associate, Full-Stack Engineer

Location
Westminster, West End, United Kingdom
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile (Scrum ...