926 to 950 of 1,849 Permanent Observability Jobs

Applied Scientist — Generative AI for Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Amazon Development Centre (London) Limited seeks an Applied Scientist specializing in generative AI to shape Prime Video observability. You will build and deploy large models, run experiments, and collaborate with UK/US teams to ...

Lead Site Reliability Engineer - Observability & Resilience

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase & Co. seeks a Lead Site Reliability Engineer to define the future of reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers ...

Strategic Account Director, Capital Markets & Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
ITRS in Greater London is seeking an experienced Account Director to manage and grow revenue within financial services. You will be responsible for building strategic relationships, managing the sales cycle, and collaborating with internal teams ...

Lead AI Software Engineer - TRP Labs London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Guide the development of reusable AI capabilities and shared platform components across areas such as agent orchestration, tool use, retrieval‐based systems, evaluation frameworks, observability, and guardrails. Champion engineering excellence through strong software design, code quality, automated testing, continuous integration, and continuous delivery practices. Ensure AI systems are built with … measurable quality, production readiness, operational observability, and appropriate safety controls. Oversee technical debt and drive continuous improvement across AI platforms, services, and development standards. Identify and pursue opportunities to apply AI in ways that accelerate workflows, improve decision‐making, and create scalable business impact across the firm. Present and demonstrate ...

Software Engineering Team Lead

Hiring Organisation
Jobleads-UK
Location
Birmingham, England, United Kingdom
software lifecycle from design ideation through to production and eventual decommissioning. Our engineering teams work under a true DevOps culture — with infrastructure as code, observability, automated testing, and continuous delivery treated as first-order concerns, not afterthoughts. You'll set architectural direction, partner closely with your Product Manager counterpart … systems and microservice development - we use Azure Service Bus, and welcome experience with similar messaging technologies such as Kafka or RabbitMQ Infrastructure: Kubernetes, Docker Observability: Prometheus, Grafana Engineering culture: DevOps, infrastructure as code, automated testing across all environments including production, continuous delivery Our Engineering Approach Full ownership: Teams own their ...

Software Engineering Team Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
software lifecycle from design ideation through to production and eventual decommissioning. Our engineering teams work under a true DevOps culture — with infrastructure as code, observability, automated testing, and continuous delivery treated as first-order concerns, not afterthoughts. You'll set architectural direction, partner closely with your Product Manager counterpart … systems and microservice development - we use Azure Service Bus, and welcome experience with similar messaging technologies such as Kafka or RabbitMQ Infrastructure: Kubernetes, Docker Observability: Prometheus, Grafana Engineering culture: DevOps, infrastructure as code, automated testing across all environments including production, continuous delivery Our Engineering Approach Full ownership: Teams own their ...

Software Engineering Team Lead

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
software lifecycle from design ideation through to production and eventual decommissioning. Our engineering teams work under a true DevOps culture — with infrastructure as code, observability, automated testing, and continuous delivery treated as first-order concerns, not afterthoughts. You'll set architectural direction, partner closely with your Product Manager counterpart … systems and microservice development - we use Azure Service Bus, and welcome experience with similar messaging technologies such as Kafka or RabbitMQ Infrastructure: Kubernetes, Docker Observability: Prometheus, Grafana Engineering culture: DevOps, infrastructure as code, automated testing across all environments including production, continuous delivery Our Engineering Approach Full ownership: Teams own their ...

Software Engineering Team Lead

Hiring Organisation
Jobleads-UK
Location
Belfast City District, Northern Ireland, United Kingdom
software lifecycle from design ideation through to production and eventual decommissioning. Our engineering teams work under a true DevOps culture — with infrastructure as code, observability, automated testing, and continuous delivery treated as first-order concerns, not afterthoughts. You'll set architectural direction, partner closely with your Product Manager counterpart … systems and microservice development - we use Azure Service Bus, and welcome experience with similar messaging technologies such as Kafka or RabbitMQ Infrastructure: Kubernetes, Docker Observability: Prometheus, Grafana Engineering culture: DevOps, infrastructure as code, automated testing across all environments including production, continuous delivery Our Engineering Approach Full ownership: Teams own their ...

SVP of Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering organization across Paris and global centers (target 150–300+ engineers). Establish CI/CD, automated testing, IaC, feature flagging, canary deployments, and observability-first culture. Drive metrics for deployment frequency, lead time, MTTR, change failure rate; implement platform reliability standards (target 99.95%+ uptime, SOC 2 Type … prompt engineering, function calling, agent frameworks) and graph/knowledge graph technologies. DevOps/SRE practices at scale: CI/CD, IaC (Terraform, Pulumi), observability (Datadog, Grafana), incident management. Leadership Qualities Builder mentality with hands-on orientation; executive presence and strong communication skills; collaborative and bias for action; comfort with ...

Senior Cloud Platform Engineer—AI Infra & Automation

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
storage solutions. This is a hand‐on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure‐as‐Code, observability, high‐performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer … internal users in their use. Turn end‐user and product requirements into deployed services. Help build automation to collect and analyse metrics and other observability data from the cloud services to support clear identification and reporting of any issues. Work with users to provide information on any product‐related issues ...

Senior Platform Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
improve developer experience, deployment processes, and platform reliability Support and optimise Azure SQL environments, ensuring performance, availability, and security Implement monitoring, logging, and observability solutions to improve platform performance and resilience Champion cloud security, governance, and automation best practices across the Azure estate Contribute to the evolution of the company … DevOps CI/CD pipelines Solid scripting skills with PowerShell Experience supporting or administering Azure SQL (or Microsoft SQL Server) Experience with monitoring and observability tooling Good understanding of cloud networking, security, identity, and platform automation Excellent communication skills with a collaborative mindset Someone who enjoys solving complex problems, improving ...

Engineering Lead

Hiring Organisation
17918
Location
London, United Kingdom
guide implementation of scalable integration solutions across APIs, middleware, event-driven platforms, and external systems. Drive best practices across software engineering, DevOps, resiliency, observability, operational excellence, audit readiness, and governance. Conduct architecture reviews, code reviews, technical design assessments, and controls compliance reviews. Provide mentoring, coaching, and hands-on technical guidance … Services, or large-scale transformation programmes. Experience supporting regulated applications, risk technology platforms, or compliance-driven initiatives. Experience with Docker and Kubernetes. Knowledge of observability, monitoring, logging, and control monitoring frameworks. Experience working within Agile delivery environments. Experience supporting geographically distributed teams. Exposure to AI-enabled engineering tools, automation frameworks ...

Senior Platform Engineer

Hiring Organisation
Capgemini
Location
Surrey, United Kingdom
Employment Type
Full Time
GitOps-driven platforms and internal developer platforms that give engineers true self-service. You’ll modernise legacy estates into cloud-native architectures, unify observability across complex environments using OpenTelemetry and modern APM platforms, and drive FinOps-focused optimisation. You’ll contribute to and help shape engineering standards, SRE-aligned operability … guardrails. • Lead technical workshops and design reviews with client stakeholders and engineering teams. • Set and evolve platform standards (security, reliability, operability, CI/CD, observability) and help teams adopt them in practice. • Define SLOs, improve reliability, and lead incident reviews. • Coach engineers through mentoring, pairing and pragmatic hands-on technical ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
operational efficiency. Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service … programming language, preferably Python, for automation and tool development. Tooling & Concepts CI/CD: Experience setting up and maintaining modern CI/CD pipelines. Observability: Practical experience implementing and managing monitoring and logging tools. Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud‐native networking within Kubernetes. ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
17918
Location
United Kingdom
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Machine Learning Systems & Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Ship workloads with Docker and Kubernetes; maintain IaC (Terraform) for the surfaces you own and CI/CD pipelines, including self‐hosted GPU runners. Observability and reliability: Monitoring, logging, and alerting for job performance, data‐pipeline health, and cost (e.g., Prometheus/Grafana, OpenTelemetry); define SLOs and incident response … stores; and object storage with caching layers. Familiarity with ML workflow orchestration and experiment tracking (e.g., Kubeflow Pipelines, MLflow). Experience with monitoring and observability tooling (e.g., Prometheus/Grafana, OpenTelemetry) and CI/CD for infra and ML workflows (e.g., GitHub Actions). At SpAItial, we are committed ...

Cloud Platform Engineer (Senior / Lead)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices—SLOs/SLIs, observability, capacity planning, resilience and blameless incident management—to keep the platform reliable and cost‐efficient. Partner with data engineering to design and optimise data pipelines, data … workloads and data—including secrets management, least‐privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost‐optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. ...

Senior Software Engineer - Agentic AI Platform

Hiring Organisation
Pinnacle Technical Resources
Location
Fort Worth, Texas, United States
Employment Type
Permanent
Salary
USD 60 Annual
maintain agent orchestration services, tool registries, and execution runtimes. Build APIs and microservices for LLM integration, prompt management, and agent lifecycle management. Implement observability, logging, and monitoring for agentic workflows. Write comprehensive tests (unit, integration, end-to-end) to ensure platform reliability. Collaborate with architects on design decisions … like LangChain, LangGraph, or similar orchestration tools. Nice to Have Skills: Experience with event-driven architectures (Kafka, EventBridge). Knowledge of vector databases and observability tools (Datadog, Splunk, OpenTelemetry). Familiarity with agent evaluation frameworks, FastAPI/Flask, async Python, and Redis/caching patterns. Experience in the airline ...