801 to 825 of 6,477 Permanent Observability Jobs

Integration Architect

Location
United Kingdom
Terraform or Bicep . Design integration patterns using REST APIs, microservices, Azure Service Bus, Event Grid and event-driven architectures . Establish monitoring and observability using Azure Monitor, Log Analytics and Application Insights . Present solution designs through architecture governance and review processes. Provide technical leadership and guidance to engineering … implementing APIOps/GitOps approaches for API lifecycle management. Knowledge of Azure API Center or similar API cataloguing/governance platforms. Experience with Azure observability tooling including Application Insights, Azure Monitor and Log Analytics . Knowledge of emerging AI integration patterns such as Model Context Protocol (MCP) or AI Gateway ...

Site Reliability Engineer

Location
Newcastle upon Tyne, England, United Kingdom
Makes This Role Great: Develop and maintain scalable infrastructure as code (IaC) using Terraform to ensure reliable and scalable cloud environments. Implement and enhance observability solutions using tools like New Relic, DataDog, Sumologic and Splunk for monitoring, logging, and alerting. Perform code deployments and manage CI/CD pipelines using … management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized SRE observability experience with New Relic or DataDog. Familiarity with OpenTelemetry, AIOps, MLOps, or SecOps. Location: Newcastle, UK - In-Office (at least 4 days per week ...

Platform Engineer III - (Pipelines & Developer Experience)

Location
Leeds, England, United Kingdom
least privilege, policy‐as‐code, scanning). Improve developer experience: faster feedback loops, quality gates, ephemeral/preview environments, and great documentation. Instrument pipeline observability (Datadog or equivalent) and define SLOs (queue time, lead time, change fail rate, MTTR) to drive reliability. Automate IaC workflows (Terraform/Terragrunt) and integrate … implement and maintain CI/CD release pipelines. Experience with scripting and programming (.NET preferred; familiarity with Go, Python, PowerShell beneficial). Knowledge of observability tooling, chaos testing, and incident management. Strong analytical and problem‐solving abilities, with the capability to closely collaborate with engineering teams. Highly outcome‐oriented, pragmatic ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity ...

Lead SRE - Chase UK

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational ExcellenceDesign and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry).Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skillsFamiliarity with infrastructure ...

Senior MLOps Engineer

Hiring Organisation
AECOM
Location
London, United Kingdom
Salary
£ 70 K
/CD, automation, and model delivery workflows used across the engineering teamDevelop Python services, APIs, and platform tooling that support production ML workloadsOwn reliability, observability, performance, availability, and security across production ML systemsTroubleshoot complex production issues across software, infrastructure, and ML workflowsDrive architecture decisions and partner with ML engineers, data … with Kubernetes and infrastructure as code, such as TerraformExperience with ML lifecycle tooling such as MLflow, Weights & Biases, or similar platformsKnowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, or OpenTelemetryExperience with distributed ML workloads, model serving, GPU infrastructure, or performance optimizationBackground in reinforcement learning, optimization, or generative ...

Senior Full Stack Engineer

Location
Greater London, England, United Kingdom
technical designs and code produced by delivery partners, identifying risks and opportunities for improvement Champion engineering excellence across code quality, testing, continuous delivery, security, observability and platform reliability Collaborate closely with AI, Machine Learning, Product, Infrastructure and DevOps teams to deliver integrated platform capabilities and share knowledge across the engineering … Familiarity with AI orchestration frameworks, intelligent workflows and emerging AI technologies Experience with data platforms and infrastructure that support AI workloads Knowledge of monitoring, observability and continuous delivery practices Experience within telecommunications or other large scale customer facing digital platforms What's in it for you? Competitive salary and bonus ...

Devops Engineer

Location
Greater London, England, United Kingdom
deployment, software repositories, databases, and web servers Own the patching and update lifecycle for managed systems Monitoring & Reliability Implement and maintain monitoring, alerting, and observability across both the existing VM estate and the new container environment Proactively identify risks, bottlenecks, and failure patterns before they impact users Define and track … e.g. Terraform, Ansible, Puppet, or similar) Solid scripting ability in Bash and at least one higher‐level language (Python preferred) Experience with monitoring and observability tooling (e.g. Prometheus, Grafana, Datadog, or similar) Strong incident diagnosis skills—able to work from vague symptoms to root cause using logs, metrics, and reasoning ...

Backend Engineer (Java)

Location
Southampton, England, United Kingdom
Implement and evolve event-driven workflows using Google Cloud Pub/Sub to synchronise profile data across multiple products. Improve the performance, reliability, and observability of the service (logging, metrics, tracing, alerting). Architecture & Technical Direction Work closely with Senior Engineers and the rest of your team to influence … fostering growth and supporting good engineering practices. Experience with Google Cloud Platform (GCP) Experience with Keycloak or other identity/auth frameworks Experience with observability tooling (e.g., OpenTelemetry, Grafana, Prometheus) Even if you don't meet all of the requirements for this role, we encourage you to apply ...

Intermediate Data Engineer

Location
Cheadle, England, United Kingdom
maintain automated deployment, testing and operational workflows using GitLab pipelines, version control, Terraform/OpenTofu and scripting. Contribute to data quality, lineage, governance and observability practices to ensure data is trusted, traceable, secure and supportable throughout its lifecycle. Investigate data quality, platform reliability, pipeline performance and operational issues, identifying practical … willingness to learn, AI-assisted engineering tools such as Kiro, GitHub Copilot or similar technologies to improve engineering effectiveness and delivery outcomes. Familiarity with observability and troubleshooting tooling such as Dynatrace, OpenSearch/Elasticsearch or CloudWatch logs. Ability to communicate clearly, analyse problems, work collaboratively and explain technical concepts ...

Site Reliability Engineer (SRE) – Cloud Engineer

Location
Glasgow, Scotland, United Kingdom
potential extension) Start Date: Immediate Key Responsibilities Design, build and maintain highly available cloud infrastructure across multi-cloud environments. Drive SRE best practices including observability, resilience, reliability engineering and automation. Develop and maintain Infrastructure as Code (IaC) solutions using Terraform. Build and enhance automation tooling using Python. Support cloud platform ...

Python Developer - AI / LLM Platform

Hiring Organisation
Ncounter
Location
Reading, Berkshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £100,000 per annum
while improving existing components through structured refactoring • Integrate LLM APIs and develop practical AI capabilities across live platforms • Take ownership of production issues, testing, observability and ongoing platform reliability • Strengthen application security, deployment practices and technical documentation We are looking for: • 5+ years of commercial Python development experience, ideally including ...

Solutions Architect

Location
Leeds, England, United Kingdom
Strong problem-solving, communication, and leadership skills. Experience with GraphQL, gRPC, and Service Mesh (Istio/Linkerd). Familiarity with AI/ML integrations, observability tools (Dynatrace, Grafana, ELK). Experience or good understanding in Agentic AI #J-18808-Ljbffr ...

Staff Platform Site Reliability Engineer

Hiring Organisation
Index Exchange
Location
London, UK
Employment Type
Full-time
about its architecture. Must Have8+ years in platform engineering, SRE, infrastructure engineering, or DevOps. Deep experience with Linux internals: kernel tuning, network stack, system observability, security. Strong Kubernetes expertise: cluster lifecycle, networking, storage, RBAC, multi-cluster—across bare-metal and cloud (EKS, GKE).Infrastructure-as-code at scale: Terraform, Ansible … driving technical strategy across teams—not just executing within one. Valuable ExperienceDistributed storage systems (e.g. Ceph)Big data infrastructure: Hadoop, Spark, HBase, Kafka. Observability stack design: Prometheus, Grafana, ELK, Mimir, Loki, Tempo. Secrets management (Vault), certificate management, access control at scale. Hybrid cloud architectures: federating public cloud (AWS, GCP) with ...

Senior DevOps

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£80,000
including Kubernetes and Docker Infrastructure as Code using Terraform or similar tools CI/CD pipeline development with GitLab, Jenkins or equivalent Monitoring and observability tools such as Prometheus, Grafana or Elastic Secure infrastructure, automation and platform engineering Agile software delivery within cross-functional teams Building resilient, scalable and highly ...

Senior DevOps

Hiring Organisation
Anson Mccade
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£80,000
including Kubernetes and Docker Infrastructure as Code using Terraform or similar tools CI/CD pipeline development with GitLab, Jenkins or equivalent Monitoring and observability tools such as Prometheus, Grafana or Elastic Secure infrastructure, automation and platform engineering Agile software delivery within cross-functional teams Building resilient, scalable and highly ...

Lead Data Engineer

Location
Southampton, England, United Kingdom
Owners, Architects, and stakeholders to prioritise and deliver work effectively. Drive adoption of modern engineering practices including automation, testing, CI/CD, source control, observability, and infrastructure-as-code. Contribute to roadmap planning by identifying technical opportunities, risks, dependencies, and improvement initiatives. Mentoring & Capability Development Provide coaching, mentoring, and technical … products and platform capabilities across one or more business domains. Experience implementing and promoting data quality frameworks, service level objectives (SLOs), data contracts, and observability tooling. Strong understanding of data governance principles, including data lineage, metadata management, cataloguing, security, and access controls. Experience ensuring compliance with regulatory and organisational requirements ...

Fullstack Software Engineer- National Security

Location
Cheltenham, England, United Kingdom
Azure or GCP Experience with Docker and/or Kubernetes Experience building and maintaining CI/CD pipelines Experience with monitoring, logging and observability platforms Experience with relational and/or NoSQL databases Experience designing and operating highly available distributed systems Experience working in an Agile/Scrum environment Locations ...

AWS Cloud Engineer - Hybrid - MCR or LDN

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
solving skills Willingness to learn new technologies and work across different project environments Microservices and API design Docker, Kubernetes, EKS or ECS AWS monitoring, observability and FinOps CI/CD tools such as Jenkins, Bamboo, TeamCity or Bitbucket Candidates must be willing and eligible to obtain UK security clearance checks ...

Python Developer - AI

Hiring Organisation
83zero Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£55,000
Experience with CI/CD using GitHub, GitLab or Jenkins Agile engineering experience Experience with AI agents, tool calling, embeddings, prompt engineering or LLM observability is beneficial React/TypeScript and Terraform/IaC experience is beneficial Experience taking GenAI POCs into production is highly beneficial Why this role ...

Python Developer - GenAI

Hiring Organisation
83zero Limited
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Salary
£55,000
Experience with CI/CD using GitHub, GitLab or Jenkins *Agile engineering experience *Experience with AI agents, tool calling, embeddings, prompt engineering or LLM observability is beneficial *React/TypeScript and Terraform/IaC experience is beneficial *Experience taking GenAI POCs into production is highly beneficial Why this role ...

Python Developer - AI

Hiring Organisation
83zero Limited
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£55,000
Experience with CI/CD using GitHub, GitLab or Jenkins Agile engineering experience Experience with AI agents, tool calling, embeddings, prompt engineering or LLM observability is beneficial React/TypeScript and Terraform/IaC experience is beneficial Experience taking GenAI POCs into production is highly beneficial Why this role ...

DevOps Engineer

Location
Preston, England, United Kingdom
Collaborate with engineering teams to ensure infrastructure is reproducible, reliable, auditable, and fit for use across environments. Monitor platform health, logging, and metrics using observability tools such as Azure Monitor, Prometheus, and Grafana where applicable. Champion automation, DevOps best practice, platform security, auditing, and cost-aware engineering across the team. ...

Remote Senior Site Reliability Engineer Manager (Remote)

Location
Cambourne, England, United Kingdom
maintaining highly reliable infrastructure and services. Expertise in incident management, including incident response, resolution, and post-mortem analysis. Proficiency in monitoring, alerting, and observability tools such as Prometheus, Grafana, ELK stack or Datadog. Experience with cloud platforms such as AWS, Azure, or GCP, including infrastructure as code tools like Terraform ...