1,451 to 1,475 of 5,223 Permanent Observability Jobs

Full Stack Engineer - Specialist

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
stack with a strong Java backend emphasis, working within agile teams that own the full product lifecycle from design and build through to deployment, observability, and iteration. We are particularly interested in candidates holding aMasters in Computer Science or Artificial Intelligence, ideally with some industrial placement or internship experience … ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into ...

Software Engineer - Cloud Compute Platform

Location
Greater London, England, United Kingdom
contribute to include: Workload orchestration across Kubernetes clusters Platform APIs and Kubernetes operators Cloud platform integrations Multi-tenant workload isolation and security Observability, health checks, and operational tooling Workload scheduling, disruption management, and rolling updates What You Will Do: Design and implement services, APIs, and Kubernetes controllers using Go. Contribute … engineers to understand requirements and deliver reliable solutions. Participate in technical design discussions and help evaluate implementation trade-offs. Improve the reliability, scalability, observability, and maintainability of existing systems. Write automated tests, documentation, and operational runbooks. Participate in code reviews and provide constructive feedback to teammates. Help investigate and resolve ...

Oracle OSS Stack Lead

Location
Newbury, England, United Kingdom
monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk, and enterprise monitoring platforms. Act as the senior technical escalation point for business stakeholders, support teams … Experience with Unix/Linux administration and Oracle SQL performance troubleshooting. Knowledge of cloud technologies, containerisation, and Kubernetes environments is advantageous. Familiarity with monitoring, observability, and enterprise operations tooling. Understanding of telecommunications network integrations and Oracle Fusion Middleware integration concepts is beneficial. Strong knowledge of Incident, Problem, Change, and Release ...

Backend Engineer

Location
Manchester, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
London, United Kingdom
Employment Type
Full-Time
Salary
£65,000 - £70,000 per annum
automate deployments using Terraform and Ansible, and build CI/CD pipelines using GitHub Actions. You'll also work across cloud security, networking and observability, while having the opportunity to develop your technical expertise through training and professional certifications. This role would suit an experienced DevOps Engineer looking to work … private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join ...

Network Automation & OSS Designer

Location
Greater London, England, United Kingdom
designs that span multiple network domains. Key Responsibilities/Job Description Design end-to-end network automation solution architectures spanning orchestration, inventory, compliance, and observability layers. Develop reference architectures and detailed solution blueprints for multi-domain network automation (IP/MPLS, optical, mobile RAN/Core, SD-WAN, cloud). … Terraform, Ansible, Helm) for network resource provisioning. Design Kafka-based event streaming and messaging architectures for real‐time network telemetry and automation triggers. Define observability strategies covering metrics, logs, traces, and network telemetry pipelines. Architect AIOps capabilities including closed‐loop automation, anomaly detection, and predictive analytics for network operations. Integrate ...

Backend Engineer

Location
City of Edinburgh, Scotland, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

Backend Engineer

Location
West of England, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

Backend Engineer 1861

Location
Leeds, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
Camden, London, Camden Town, United Kingdom
Employment Type
Permanent
Salary
£65000 - £70000/annum + Remote + Progression
automate deployments using Terraform and Ansible, and build CI/CD pipelines using GitHub Actions. You'll also work across cloud security, networking and observability, while having the opportunity to develop your technical expertise through training and professional certifications. This role would suit an experienced DevOps Engineer looking to work … private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
capabilities, reusable golden paths, and standardised service templates for efficient software delivery. Provide technical direction across platform domains including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Enable AI‐assisted engineering workflows to enhance productivity, automate routine activities, and improve decision‐making. Embed governance, security … compliance controls through policy‐as‐code and auditable platform practices. Improve platform reliability through SLOs, observability frameworks, resilience engineering, and incident‐driven improvements. Collaborate with product, engineering, security, and architecture teams to translate business needs into scalable platform solutions. Promote engineering excellence through automation, standardisation, and efficiency‐driven practices, including ...

platform engineer for legal AI

Location
Greater London, England, United Kingdom
environments; Implement and evolve Infrastructure-as-Code using Terraform; Collaborate with developers to improve CI/CD pipelines, deployment strategies, and developer experience; Enhance observability and reliability through alerting, monitoring, and incident response; Collaborate on cloud optimization projects to improve performance, cost efficiency, and security posture; Mentor and guide team … management; Experience with CI/CD systems such as Buildkite and GitHub Actions; Experience with relational or non-relational datastores in production; Familiarity with observability platforms such as Datadog, Prometheus, and Grafana; Familiarity with Linux systems administration, networking, and troubleshooting; Excellent communication and documentation abilities, with a focus on knowledge ...

Principal DevSecOps Engineer

Hiring Organisation
83zero Limited
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
workflows * Establish secure-by-design engineering practices and enforce security and technical standards * Lead Infrastructure as Code (IaC) practices across teams and environments * Drive observability, monitoring, logging and audit controls * Support incident response, patching, compliance reporting and technical debt remediation * Partner with developers and delivery teams to improve engineering quality … Security & compliance - Trivy, vulnerability management, HashiCorp Vault, cert-manager * Containers & cloud - Docker, AWS EKS, AWS IAM, S3 and network policies * Infrastructure as Code - Terraform * Observability - Grafana, Loki * Automation - Python and Bash * Experience delivering within the UK Government Digital Service (GDS) lifecycle on a public sector engagement Why join ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
scalability, and performance. Develop and maintain automation workflows using Ansible, Salt, and related frameworks to reduce operational toil. Build and operate monitoring, alerting, and observability dashboards using tools such as Grafana and Splunk. Proactively identify network bottlenecks, performance issues, and reliability risks, implementing long‐term fixes rather than reactive solutions. … network protocols (BGP, OSPF, EIGRP, STP, VXLAN, VPNs, QoS, MPLS, etc.). Strong experience with infrastructure automation using Ansible and Salt. Proficiency with observability tooling such as Grafana, Splunk, or equivalent. Solid understanding of SRE practices including SLIs, SLOs, error budgets, and proactive reliability. Strong troubleshooting, analytical, and performance optimization ...

Oracle OSS Stack Lead

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk, and enterprise monitoring platforms. Act as the senior technical escalation point for business stakeholders, support teams … Experience with Unix/Linux administration and Oracle SQL performance troubleshooting. Knowledge of cloud technologies, containerisation, and Kubernetes environments is advantageous. Familiarity with monitoring, observability, and enterprise operations tooling. Understanding of telecommunications network integrations and Oracle Fusion Middleware integration concepts is beneficial. Strong knowledge of Incident, Problem, Change, and Release ...

AWS Data Engineer

Location
Greater London, England, United Kingdom
version control, CI/CD pipelines, automated testing, and modular code principles. Participate in pair programming, code reviews, and architectural design sessions. Data Quality & Observability Build monitoring and observability into data flows. Implement basic data quality checks and contribute to continuous improvements. Stakeholder Collaboration Work closely with business teams ...

Staff Platform Engineer

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
capabilities, reusable golden paths, and standardised service templates for efficient software delivery. Provide technical direction across platform domains including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Enable AI-assisted engineering workflows to enhance productivity, automate routine activities, and improve decision-making. Embed governance, security … compliance controls through policy-as-code and auditable platform practices. Improve platform reliability through SLOs, observability frameworks, resilience engineering, and incident-driven improvements. Collaborate with product, engineering, security, and architecture teams to translate business needs into scalable platform solutions. Promote engineering excellence through automation, standardisation, and efficiency-driven practices, including ...

Lead Software Engineer - Platform

Location
Greater London, England, United Kingdom
need to scale accordingly, and our platform foundations need to support faster, more reliable delivery. You’ll be central to making that happen – from observability and developer experience to the infrastructure patterns that underpin everything we ship. What you’ll do Alongside the other Lead Engineers you’ll support … infrastructure. You’ll be responsible for the reliability, scalability, and operability of our systems. That means CI/CD pipelines, infrastructure‐as‐code, observability, incident response, and the day‐to‐day health of production. You’ll make sure we can ship with confidence and sleep at night. Shape technical direction. ...

Senior AI Product Engineer

Location
Greater London, England, United Kingdom
more junior engineers through pair programming, code review, and design feedback. Raise the engineering bar across the team by promoting good practices in testing, observability, and AI system reliability. Influence cross-team decisions on how AI capabilities integrate with the rest of the Elliptic platform. What you will achieve … technical direction of an AI workstream, including architecture, evaluation, and rollout. Established or improved at least one team practice for building AI systems (evals, observability patterns, prompt management, rollout safety). Mentored junior engineers on AI engineering practices and contributed to their growth. Built strong working relationships across product ...

ML Ops Engineer

Hiring Organisation
CMC Markets
Location
London, UK
Employment Type
Full-time
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services. You'll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams. This is not a research role. … production issues across model, application, infrastructure and critical data-dependency layers. Reliability, security and engineering qualityImprove system robustness, scalability and cost efficiency through automation, observability and infrastructure as code. Write production-grade Python for long-running services, deployment tooling and ML workflows. Establish testing, validation, release and incident-management practices ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
London, UK
Employment Type
Full-time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering … solving, prioritization, and decision-making skills. What We're Looking For: Basic Required Qualifications:10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics ...

Junior Azure Engineer

Location
Greater London, England, United Kingdom
containers Implement security and governance controls (RBAC, Azure Policy, Management Groups) Build and support landing zones and foundational cloud environments Manage monitoring and observability using Azure Monitor, Log Analytics, and alerts Support CI/CD pipelines using Azure DevOps or GitHub Actions Automate operational tasks using PowerShell, Azure … experience with Infrastructure-as-Code (Bicep, ARM, or Terraform) Solid understanding of Azure networking and cloud architecture principles Experience with monitoring, logging, and observability tools Ability to troubleshoot and resolve complex cloud issues Experience with automation and scripting (PowerShell, Azure CLI) Strong collaboration and communication skills Experience with Azure Kubernetes ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
scalability. Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions. Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively. Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling, and improve infrastructure efficiency. Improve … Kubernetes operational experience (on-prem and AWS EKS).Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows. Deep understanding of observability and telemetry (monitoring, logging, tracing).Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD.Scripting proficiency in Python, Bash, or Go. Proven ...

Senior / Principal Applied AI Engineer (UK / Europe, Remote)

Location
United Kingdom
Core Platform Engineering. This is a hands-on engineering seat, not an advisory one. You write and own production code, you put evaluation and observability on everything you ship, and you run autonomously in a lean team. You report to the COO and partner closely with the CTO, the Head … commodity internal needs; build or replace it internally when it’s differentiating or becomes cost-prohibitive. Ship outcomes, not architecture debates. Put evaluation, observability, and cost guardrails on everything - golden datasets, eval harnesses, tracing, fallback chains, latency and spend controls. Nothing ships as an unmeasured demo. Operate inside our security ...

Engineering Manager - Grafana Application Security | EMEA | (Remote, UK)

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...