626 to 650 of 1,493 Observability Jobs

Lead Data Product Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
run.* Industrialise onboarding of sources such as ETRM/Endur, Trayport, market data vendors, SAP and operational data feeds through clear product requirements, SLAs, observability needs, runbook expectations and adoption measures.*Translate business context into product outcomes and engineered delivery** Bring deep commodities data knowledge across trading lifecycles, market data ...

Head of Engineering - Quality

Hiring Organisation
Jobleads-UK
Location
Reading, England, United Kingdom
continuous improvement. Reliability and availability principles: Understanding of how quality practices support highly available systems, including change failure reduction, incident prevention, operational readiness, observability, SLOs and post-incident learning. Metrics-driven improvement: Ability to define, collect and interpret quality and delivery metrics such as escaped defects, defect leakage, change failure ...

Senior Architect - Email & Collaboration Security

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
fluency and daily practice, not passing familiarity. ECS Domain Email threat detection and prevention Email security efficacy Collaboration security DMARC analyzer Data platform and observability Analysis and Response End user application integration Internal operations enablement What You Will Do Lead from Problem Space, Not Solution SpaceThe role engages ...

Principal Engineer

Hiring Organisation
Jobleads-UK
Location
United Kingdom
understanding their specific situation and working through it with them. Familiarity with modern software delivery practices: CI/CD, trunk‐based development, automated testing, observability tooling. Enough context to understand how agentic workflows fit into – and sometimes challenge – existing engineering norms. Approach Does the work first, explains it second. Earns ...

AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Vertex AI/Gemini Enterprise and the wider Google Cloud Platform (e.g. Cloud Run, Pub/Sub, Cloud Storage). Experience with an LLM observability or evaluation platform such as Langfuse, LangSmith, Braintrust or Arize Phoenix. Understanding of MCP and integration of tools with agentic systems. Experience grounding LLM output ...

Senior Frontend Engineer - Billing

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
coding standards across the frontend stack. Frontend Strategy: Establishes frontend architecture patterns, including state management, performance budgets, accessibility, and comprehensive testing. Operational Maturity: Leads observability and frontend performance monitoring; defines SLOs where relevant, manages incident responses, and conducts blameless post-mortems. Security & Risk: Oversees frontend security practices, including secrets hygiene ...

Product Engineer (all levels)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
help shape that), here’s a current snapshot: Full‐stack TypeScript, React, Postgres, and Temporal for long‐running orchestration. Infra: AWS, Terraform, strong observability via Sentry and Datadog (full‐stack, not just logs). Data: we lean heavily on Snowflake, Omni, dbt, Fivetran, Amplitude and Segment — and we’ve even ...

Senior Backend Engineer - Order Management & Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Jobtailor in London is seeking a senior backend software engineer to design, build and evolve order management capabilities supporting customer fulfilment journeys across checkout, fulfilment, returns, amendments, and cancellations. You will own features end-to ...

Lead Azure SRE: Reliability, Observability & Automation

Hiring Organisation
Jobleads-UK
Location
Nottingham, England, United Kingdom
The Nottingham Building Society is seeking a Lead Azure Site Reliability Engineer to drive resilience, performance, and availability of Azure platforms. You will champion SRE practices, guide reliability improvements, and lead modern operating standards to ...

Senior DevOps / SRE

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Contract
operational challenges, and is passionate about building resilient, highly available platforms. You'll play a key role in maintaining and improving cloud infrastructure, automation, observability, and system reliability. Key Responsibilities Support and operate cloud-native applications within Google Cloud Platform (GCP) Build and maintain reliable, scalable production systems with … maintain CI/CD pipelines to support efficient and reliable deployments Manage infrastructure through Infrastructure as Code using Terraform Implement best practices around observability, monitoring, incident response, and platform resilience Collaborate with engineering teams to improve operational performance and system reliability Required Experience Strong Google Cloud Platform (GCP) experience Proven ...

Platform Engineer (DevOps / MLOps Focus)

Hiring Organisation
The Portfolio Group
Location
London, United Kingdom
Employment Type
Permanent
Salary
£100000/annum
environments for production workloads. Developing and managing Infrastructure as Code using Terraform. Supporting CI/CD pipelines and platform automation initiatives. Improving platform reliability, observability and scalability. Collaborating closely with Software Engineers, Data Engineers and ML teams to optimise deployment workflows and infrastructure performance. Essential experience: Strong commercial experience with … available, scalable production environments. Nice to have: Experience with Kubeflow and ML platform tooling. Exposure to AI, machine learning or GenAI projects. Experience with observability tooling such as Prometheus, Grafana or OpenTelemetry. Experience working within regulated or enterprise environments. This is a fantastic opportunity to join a high-profile ...

Custody Support Applications Support - Assistant Vice President

Hiring Organisation
Jobleads-UK
Location
Belfast City District, Northern Ireland, United Kingdom
operational excellence of a suite of business-critical custody and settlement applications. The role focuses on distributed systems, cloud-native technologies, microservices, and modern observability platforms supporting securities processing and settlement functions.This position combines traditional application support responsibilities with Site Reliability Engineering (SRE) principles, automation, resiliency engineering, and operational risk … enterprise relational databases* Database performance analysis and tuning* Linux/Unix fundamentals* Application troubleshooting and performance diagnostics* Incident and Problem Management processes* Monitoring and observability platforms**Desirable:*** OpenShift/Kubernetes* Cloud technologies (Google Cloud Platform, Azure, AWS)* Microservices and distributed architectures* Event-driven architectures and messaging platforms (Kafka, MQ)* Elastic ...

Java Developer

Hiring Organisation
Global
Location
Greater London, United Kingdom
Employment Type
Full Time
with modern Java microservices (Java 21, Spring Boot), deployed on Kubernetes (EKS on AWS), using CI/CD pipelines with Jenkins and Terraform, and observability through Prometheus and Grafana, in a collaborative, agile team. As a Java Developer at Global, you will: Key Responsibilities Feature development (35%) : Build and enhance … development practices in code reviews, offering insightful feedback and supporting others' growth. Helped maintain a reliable production environment, using the team's monitoring and observability tooling. Built a strong understanding of the business context and how the team's work supports it. Gained a solid grasp of the team ...

Senior Azure DevOps Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
improve deployment processes, developer experience, and platform reliability Support and optimise Azure SQL environments, ensuring performance, availability, and security Implement monitoring, logging, and observability solutions to improve platform health and performance Champion cloud security, governance, and DevOps best practices across the Azure estate Be actively involved in architectural decisions … DevOps CI/CD pipelines Solid scripting experience using PowerShell Experience supporting or administering Azure SQL (or Microsoft SQL Server) Experience with monitoring and observability tooling Strong understanding of Azure networking, cloud security, identity, and automation Excellent communication skills with the ability to work across engineering and business teams ...

Senior Cloud Platform Engineer (Kubernetes, AWS/GCP, Terraform)

Hiring Organisation
HTC Global Services Inc
Location
Dearborn, Michigan, United States
Employment Type
Permanent
Salary
USD Annual
/CD pipelines. Automate deployment validation, rollback, and migration processes. Implement cloud security best practices, including IAM and workload identity. Build monitoring, alerting, and observability solutions. Develop automation tools and scripts using Python, Go, Bash, or Node.js. Participate in on-call support and incident response. Troubleshoot production infrastructure issues … GitOps workflows. Experience creating and maintaining Helm charts. Strong troubleshooting and problem-solving skills in cloud infrastructure and distributed systems. Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, or similar. Experience developing CI/CD pipelines. Strong infrastructure automation and scripting skills. Preferred Qualifications Experience with ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum It Recruitment Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using … with tools such as Jenkins, GitLab CI/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling ...

DataOps Engineer

Hiring Organisation
Capgemini
Location
Hampshire, United Kingdom
Employment Type
Full Time
data platforms so they are reliable, secure, and easy to change. You will work across engineering and operations to automate delivery, improve data pipeline observability, and embed good governance so analytics and AI workloads can run at scale. You will be part of the Data Platforms team that sits within … deliver DataOps capabilities, with a focus on delivering data applications in an automated approach. You will also implement and manage comprehensive monitoring and observability solutions to ensure data quality across the entire data flow, including supporting delivery in air-gapped and other restricted environments. Delivering data applications in an automated ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
evolve the fully automated CI and CD pipelines. This includes establishing best practices for fast, reliable, and secure build, test, and deployment processes. Observability: Implement and manage robust systems for monitoring (metrics), logging (centralised log aggregation), and distributed tracing to provide deep insights into application and infrastructure health. What … have: Proven experience developing robust, maintainable, and well‐tested automation scripts, services and pipelines to manage infrastructure, deployments, and operational tasks. Operational Tooling and Observability Management: Must have: Experience owning, managing, and maintaining mission‐critical operational tooling. Desirable: Proven background in implementing and managing centralised logging solutions or similar platforms ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
Cardiff, Wales, United Kingdom
workflows for Kubernetes‐based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application‐level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real‐world, not greenfield ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
workflows for Kubernetes‐based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application‐level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real‐world, not greenfield ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
Leeds, England, United Kingdom
workflows for Kubernetes‐based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application‐level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real‐world, not greenfield ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
workflows for Kubernetes‐based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application‐level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real‐world, not greenfield ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
workflows for Kubernetes‐based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application‐level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real‐world, not greenfield ...

Compute Platform Engineer II

Hiring Organisation
GSK
Location
Greater London, United Kingdom
Employment Type
Full Time
/CD-driven platform represents and enables the entire application and analysis lifecycle including interactive development and explorations (notebooks), large-scale batch processing, observability and production application deployments. A Compute Platform Engineer II is a technical contributor who can consistently take a poorly defined business or technical problem, work … Knowledge and use of at least one common programming language: e.g., Python, Go, C++, Scala, Java, including toolchains for documentation, testing, and operations/observability Expertise in modern software development tools/ways of working (e.g. git/GitHub, devops tools, metrics/monitoring, ...) Cloud expertise (e.g., AWS, Google ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
Newcastle upon Tyne, England, United Kingdom
workflows for Kubernetes-based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application-level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real-world, not greenfield ...