1,326 to 1,350 of 1,791 Permanent Observability Jobs

AI Native SW Engineering Specialist

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Design and build production‐grade agentic systems end‐to‐end: multi‐agent orchestration, RAG pipelines, policy‐based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … open‐source models – with fallback routing, token, cost, and latency management Implement LLMOps in production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions – workshops, proofs of concept ...

Network Automation Engineer

Hiring Organisation
HCLTech
Location
City of London, London, United Kingdom
Modern Ops and AI-first operating model. The role focuses on Network Infrastructure as Code (NetIaC), CI/CD pipelines, AI-driven operations (AIOps), observability integration, and SRE-led reliability engineering. Key Responsibilities Develop and manage Network Infrastructure as Code (NetIaC) using Python, Ansible, and Terraform for provisioning and lifecycle … ITSM workflows. Drive AI/ML use cases such as WAN capacity forecasting, anomaly detection, predictive analytics, and self-healing networks. Integrate and manage observability platforms (SolarWinds Orion, Elastic, Grafana, ZDX) for proactive monitoring and insights. Provide engineering and support for MCP (Model Context Protocol) and AI agent integrations. Ensure ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Head of SRE to define, lead and evolve our global reliability strategy. This senior leadership role is responsible for driving operational excellence, service reliability, observability, automation and continuous improvement across our technology landscape. Key Responsibilities Drive adoption of SRE principles (SLOs, error budgets, toil reduction). Establish observability and monitoring … with strong scripting and development capabilities using technologies such as Python, PowerShell, Bash, Terraform and Ansible Automation Platform. Other key skills: Robust knowledge of observability and monitoring practices, and experience implementing and managing platforms such as Dynatrace, Prometheus, Grafana, and Splunk. Good understanding of CI/CD tooling and modern ...

Senior Data Platform Engineer (Fixed Term Contract)

Hiring Organisation
Jobleads-UK
Location
Nottingham, England, United Kingdom
safe, controlled deployments Supporting modern Lakehouse architecture, ingestion, processing and data serving layers Embedding security, governance and compliance controls across the platform Driving platform observability, monitoring, and performance optimisation Owning aspects of cost management (FinOps) and platform efficiency Collaborating with Data Engineers, Analytics & AI teams to deliver trusted data products …/CD pipelines Experience building secure, scalable cloud environments A solid understanding of data platforms, pipelines or analytics ecosystems A mindset focused on reliability, observability and continuous improvement Technical skills: AWS (compute, storage, IAM, networking) Terraform and/or AWS CDK CI/CD tooling and pipeline engineering Observability, monitoring ...

Senior Associate Engineer (Digital Product Team)

Hiring Organisation
Canada Life
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
broad, T-shaped engineer, you will also bring a particular interest in the operational side of engineering, Site Reliability Engineering (SRE), observability, and the delivery pipelines that get our work safely to production. This is a blended role: you will spend plenty of your time building product features as well … facilitate a seamless, repeatable route to production for Azure hosted solutions Operating and supporting environments up to and including production, instrumenting services with meaningful observability so we can trust our releases and respond quickly when something goes wrong You will also have a product mindset: you care about customer outcomes ...

Principal AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
golden paths, and standard service templates to simplify service provisioning and operations. Contribute to cloud‐native platform architecture, including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Integrate AI‐assisted engineering workflows to support faster delivery, improved code quality, automation, and data‐driven operational decisions. … Establish governance, security, and compliance guardrails through policy‐as‐code and auditable platform patterns. Improve platform reliability using SLOs, observability practices, resilience engineering, and insights from incidents. Collaborate with product, engineering, security, and architecture teams to align platform capabilities with business priorities and user needs. Drive efficiency and sustainability through ...

Principal AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
reusable golden paths, and standard service templates to simplify service provisioning and operations.Contribute to cloud-native platform architecture, including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code.Integrate AI-assisted engineering workflows to support faster delivery, improved code quality, automation, and data-driven operational decisions.Establish governance … security, and compliance guardrails through policy-as-code and auditable platform patterns.Improve platform reliability using SLOs, observability practices, resilience engineering, and insights from incidents.Collaborate with product, engineering, security, and architecture teams to align platform capabilities with business priorities and user needs.Drive efficiency and sustainability through automation, standardisation, and FinOps-informed ...

Logging and Inventory Cloud Engineer - GCP & AWS

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
scripts and tooling using Python and Shell scripting. Support cloud engineering activities across AWS and Google Cloud Platform (GCP). Drive improvements in cloud observability, monitoring and operational reporting. Manage and optimise containerised workloads where required. Work closely with engineering and platform teams to improve cloud governance and operational controls. … Strong analytical and troubleshooting skills. Experience within large-scale enterprise or financial services environments. Knowledge of cloud governance, compliance and operational controls. Experience with observability and monitoring tooling in cloud-native environments. Exposure to DevOps and CI/CD practices. #J-18808-Ljbffr ...

Principal Site Reliability Engineering Expert Director

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
shaping how reliability, automation, and operational excellence are engineered across the organisation. Operating across domains including traditional infrastructure, cloud engineering, network operations, identity, observability, security, AI-driven operations, and automated data workflows, the role focuses on designing scalable systems, reusable engineering patterns, and standardised controls that reduce operational toil, improve … first, measurable, and repeatable practices. A key part of the role is building and evolving reusable CI/CD and Terraform modules, engineering guardrails, observability patterns, and automation frameworks that can be adopted across multiple teams and domains without requiring each team to solve the same problems independently. The Principal ...

Senior Site Reliability Engineer

Hiring Organisation
Source Group International
Location
Colchester, Essex, United Kingdom
Employment Type
Permanent
Salary
GBP 80,000 - 100,000 Annual
improvement. Automate operational processes and reduce manual toil through engineering. Build and operate cloud-native platforms using Azure, Kubernetes, and Infrastructure as Code. Develop observability through effective monitoring, alerting, and telemetry. Mentor engineers and promote reliability best practices across the organisation. Experience Proven experience as a Senior or experienced Site … make safe changes to PHP and Java or .NET applications. Experience operating large-scale production SaaS systems. Strong knowledge of SRE principles, incident management, observability, and operational excellence. Hands-on experience with Azure, Kubernetes, Infrastructure as Code, and monitoring platforms such as Prometheus, Grafana, or Datadog. Experience influencing engineering teams ...

Lead Site Reliability Engineer – Operations Excellence

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting‐edge AI platforms. You’ll work hands‐on with cloud and Kubernetes‐based deployments, deep observability, and cost‐aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on‐prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

NOC Engineer (AWS)

Hiring Organisation
Spectrum IT Recruitment
Location
Basingstoke, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£60000/annum Bonus, Pension, Healthcare
issues and restoring services quickly and effectively Developing automation to reduce manual operational tasks and improve platform resilience Building and improving monitoring, alerting and observability across cloud environments Working alongside Software, Platform, Cloud and Security Engineers to improve reliability and operational excellence Contributing to post-incident reviews and driving continuous … with exposure to: Linux systems administration AWS cloud infrastructure Kubernetes and Docker Production support and incident management Python, Bash or Go scripting Monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch Networking fundamentals including DNS, TCP/IP and load balancing A passion for automation, continuous improvement ...

NOC Engineer, AWS

Hiring Organisation
Spectrum It Recruitment Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
issues and restoring services quickly and effectively Developing automation to reduce manual operational tasks and improve platform resilience Building and improving monitoring, alerting and observability across cloud environments Working alongside Software, Platform, Cloud and Security Engineers to improve reliability and operational excellence Contributing to post-incident reviews and driving continuous … with exposure to: Linux systems administration AWS cloud infrastructure Kubernetes and Docker Production support and incident management Python, Bash or Go scripting Monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch Networking fundamentals including DNS, TCP/IP and load balancing A passion for automation, continuous improvement ...

DevOps Team Lead - Development Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
ownership and collaboration with other department teams Understand development team needs and design CI/CD solutions for a Java/Springboot stack Drive observability, reliability, and adoption of software development technologies Maintain, architect, and evolve the development environment to improve efficiency and productivity Collaborate with ops teams on technology …/CD practices Experience with containers and container orchestration systems Experience with IaC, cloud computing, and DevSecOps/GitOps practices Experience with an observability stack Nice to Have Java/Springboot projects experience GitOps proficiency with ArgoCD (ApplicationSets, sync policies, multi‐environment promotion) AWS experience (EKS, EC2, IAM, VPC, CloudWatch ...

Azure Platform Engineer

Hiring Organisation
EMBS Engineering
Location
Newbury, West Berkshire, Berkshire, United Kingdom
Employment Type
Permanent
Salary
£65000 - £75000/annum + Benefits
Azure platform templates and engineering patterns to support consistent platform adoption. Build and improve CI/CD pipelines, deployment automation and release processes. Implement observability, monitoring, logging, resilience and operational readiness across the platform. Embed FinOps principles, improving cloud cost visibility and optimisation. Work closely with engineering teams throughout sprint … engineering patterns. Strong understanding of DevOps and DevSecOps practices including CI/CD, source control, automated testing and release management. Experience with Azure observability, monitoring, alerting, logging and platform reliability. Practical knowledge of cloud security, governance and enterprise engineering standards. Experience applying FinOps principles including tagging strategies, right-sizing ...

Technical Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
requirements into scalable solutions and balance long‐term technical strategy with day‐to‐day delivery. You’ll champion engineering best practices including automation, testing, observability, security, operational excellence and AI‐assisted software development, creating an environment where engineers can do their best work. About the USA Tech Team Our mission … initiatives. Champion engineering excellence through modern development practices, testing, automation and continuous delivery. Drive improvements in platform reliability, scalability, security and operational performance. Promote observability, monitoring and operational ownership across the engineering lifecycle. Support engineers through mentoring, coaching and technical leadership. Collaborate with architecture, platform and security teams to ensure ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
areas of responsibility Working within the Platforms team, you’ll build robust platform services, enabling self‐service infrastructure, seamless CI/CD, and strong observability, while also supporting API‐first development, system integration, and the adoption of AI/ML capabilities across the organisation. Turning platform strategy into reality through … security, compliance, and best practices into the platform by design. Contributing to platform standards, patterns, and reusable components to drive consistency across teams. Improving observability, monitoring, and alerting capabilities across systems and services. Supporting onboarding of teams to the platform and enabling adoption through documentation and guidance. Continuously identifying opportunities ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
take responsibility for our Kubernetes estate, multi-cloud footprint (Azure primary, plus AWS and GCP), Istio service mesh, Terraform-managed infrastructure, and LGTM observability stack, and for making all of it secure by default. Our customers trust us with their most sensitive data, so security is built into every layer … self-hosted search cluster. Infrastructure as code: A multi-cloud Terraform estate spanning Azure, AWS, and GCP, built for repeatability, auditability, and least privilege. Observability: Our LGTM stack (Loki, Grafana, Tempo, Mimir), making the platform's behaviour legible to every engineer on the team. Security & compliance: Identity and access, secrets ...

Senior Software Engineer

Hiring Organisation
Fruition Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
multiple products Integrate external identity, KYC and payments platforms (e.g. Senzing, Auth0, Stripe) Own services end-to-end: API design, MongoDB data modelling, testing, observability, deployment Build secure systems (JWT/OIDC, fine-grained authorisation, IDOR protection, audit logging) Write automated tests (Vitest, Playwright) as part of everyday development Mentor …/AML, fintech or other regulated-industry experience Payment provider integration (e.g. Stripe) Monorepo tooling (pnpm, TurboRepo) and CI/CD (GitHub Actions) Observability tooling (Prometheus, Sentry) We are an equal opportunities employer and welcome applications from all suitably qualified persons regardless of their race, sex, disability, religion/belief ...

DevSecOps Engineering Lead CGEMJP00346044

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
management, including allowlist processes and risk acceptance where required Secrets management and identity/access management Policy enforcement for workloads, container images and infrastructure Observability, monitoring, logging and audit controls Partner with developers to embed secure-by-design engineering and ensure compliance with CLIENT security standards. Enable and govern Infrastructure … compliance tooling (e.g. Trivy scanning and vulnerability management, HashiCorp Vault, cert-manager) Containers and orchestration (e.g. Docker, AWS EKS) Infrastructure as Code (e.g. Terraform) Observability (e.g. Grafana, Loki) Scripting and automation (e.g. Python, Bash) Cloud and networking fundamentals (e.g. AWS IAM, S3, network policies) Experience delivering within the UK Government ...

Staff / Senior Software DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
baselines, sound metric aggregation, determinism), and turn CI and test signal into CI‐health and product‐readiness dashboards that drive real decisions. Own Software Observability: Choose the metrics store that scales to many series on daily runs with long‐lived history, making dashboards for observable software. Set Standards: Define … Performance‐analysis support: trustworthy regression baselines, determinism and noise handling, sound metric aggregation (geometric vs arithmetic mean vs median), and fast attribution and bisection. Observability and metrics platforms: CI‐health and readiness dashboards, a metrics store that scales to long‐lived, high‐cardinality time series, and self‐serve access ...

Software Engineer III - Python

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
infrastructure‐as‐code using Terraform within established team patterns across modules, environments, and state management Improve operability of services by adding and using observability tooling including logs, metrics, traces, dashboards, and alerts, and participate in incident response and root‐cause analysis Leverage enterprise‐authorized AI coding assist tools within … implementing application logic and APIs on top of relational data Experience building APIs and microservices using REST or gRPC, including contracts, security basics, and observability Practical experience delivering LLM‐based features as part of software systems, with familiarity with agentic patterns Working knowledge of delivery and operations including CI/ ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Staff Software Engineer UK

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Operations to influence technical roadmaps and delivery. Mentor senior engineers and strengthen technical leadership across the organisation. Lead initiatives that improve engineering quality, testing, observability, and operational excellence. Contribute hands-on to critical systems where your expertise delivers the greatest impact. What it takes to succeed: Extensive experience designing … Experience with microservices, APIs, event-driven architectures, relational databases, and cloud platforms (preferably AWS). Knowledge of CI/CD, Infrastructure as Code, containerisation, observability, and automated testing. Strong communication skills with the ability to influence technical decisions and mentor experienced engineers. Professional fluency in English. Java experience is highly ...