1,476 to 1,500 of 1,885 Observability Jobs

Senior SRE / Platform Engineer- Global Prime Brokerage & Financing Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
optimize CI/CD pipelines to improve software delivery for large-scale distributed systems using Amazon CodeBuild, GitHub Actions & Terraform Enterprise Implement and maintain observability solutions to establish real-time monitoring and proactive incident response using Datadog and AWS CloudWatch Ensure high availability and performance of relational and time-series … with event streaming (i.e. Kafka/Kinesis) Deep understanding of distributed systems architecture, fault tolerance, disaster recovery, and performance tuning Hands-on experience with observability and monitoring tools (e.g. Datadog, ELK, CloudWatch) Proven track record managing infrastructure with Terraform or AWS CDK Strong background in incident management, system reliability ...

Lead Product Manager AIOPs

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering … solving, prioritization, and decision‐making skills. What We’re Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics ...

DevOps Engineer

Hiring Organisation
DGH Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£80000 - £100000/annum
platform reliability. Key Responsibilities - Design, deploy, and manage AI platforms and agent infrastructure - Build and maintain CI/CD pipelines and DevOps workflows - Implement observability, monitoring, and logging solutions - Optimise performance, scalability, and cost efficiency - Support AI teams with infrastructure, deployment, and integration - Ensure platform security, compliance, and high availability …/CD, automation, and DevOps best practices - Experience with Kubernetes/containerisation technologies - Strong programming skills (e.g. Python, Go, Node.js) - Experience with observability tools (e.g. OpenTelemetry, Datadog) - Understanding of security, performance optimisation, and scalability Desirable Skills - Experience working on AI/ML platforms or deployments - Exposure to large-scale distributed ...

Site Reliability Engineering Manager

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
Reliability Engineers. Shape and deliver our Site Reliability Engineering roadmap alongside the Head of Platform. Champion modern engineering practices including SLIs, SLOs, error budgets, observability and automation. Improve the reliability, scalability and performance of our cloud platforms and digital services. Partner with Engineering, Security, Data and Product teams to embed … operational excellence from design through to production. Drive the adoption of our observability platform, helping teams gain deeper insight into the health and performance of their services. Lead incident learning, continuous improvement and automation initiatives that reduce operational toil. Provide technical leadership across AWS, Kubernetes, Infrastructure as Code, CI/ ...

Junior Azure Engineer

Hiring Organisation
COMPUTACENTER (UK) LIMITED
Location
South East London, London, United Kingdom
Employment Type
Permanent
containers Implement security and governance controls (RBAC, Azure Policy, Management Groups) Build and support landing zones and foundational cloud environments Manage monitoring and observability using Azure Monitor, Log Analytics, and alerts Support CI/CD pipelines using Azure DevOps or GitHub Actions Automate operational tasks using PowerShell, Azure … experience with Infrastructure-as-Code (Bicep, ARM, or Terraform) Solid understanding of Azure networking and cloud architecture principles Experience with monitoring, logging, and observability tools Ability to troubleshoot and resolve complex cloud issues Experience with automation and scripting (PowerShell, Azure CLI) Strong collaboration and communication skills Desirable Experience with Azure ...

Senior Cloud Platform Engineer (GCP | Kubernetes | DevSecOps)

Hiring Organisation
Jobleads-UK
Location
Bolsterstone, England, United Kingdom
repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. … Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude ...

Senior Cloud Platform Engineer (GCP | Kubernetes | DevSecOps)

Hiring Organisation
GCS
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £620/day
repeatable, automated deployments. Implement and maintain CI/CD pipelines and GitOps deployment workflows. Manage cloud networking, connectivity and platform security. Implement platform observability including logging, monitoring, metrics and distributed tracing. Automate platform provisioning, configuration management and operational tasks. Support deployment and operation of identity platform components and supporting services. … Pipelines GitOps Linux Administration Networking and Load Balancing Service Mesh Technologies (Istio/Envoy) Container Platforms Secrets Management PKI and Certificate Management Workload Identity Observability (Logging, Monitoring and Tracing) Scripting and Automation DevSecOps Site Reliability Engineering (SRE) Performance and Capacity Management Operational Support Global Deployment Strategies Positive can-do attitude ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
Greater London, United Kingdom
Employment Type
Full Time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools - Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering … solving, prioritization, and decision-making skills. What We're Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics ...

Azure Platform Engineering Consultant

Hiring Organisation
Morgan McKinley
Location
Newbury, Berkshire, England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
platform templates and landing zone patterns. CI/CD & Automation: Build and refine automated deployment pipelines, environment management, and release practices. Platform Quality: Embed observability (monitoring, logging, alerting), resilience, security, and FinOps principles directly into platform assets. Co-Delivery & Knowledge Transfer: Work closely alongside client engineering teams to pair, document … Core compute, networking, storage, identity, security, and platform services. Infrastructure as Code: Strong proficiency with Terraform AND Terragrunt using modular, reusable implementation patterns. DevOps & Observability: Strong experience with CI/CD tools (Azure DevOps/GitHub Actions) and monitoring stacks (Prometheus, Grafana, Azure Monitor, etc.). FinOps: Practical knowledge ...

Principal Cloud Architect

Hiring Organisation
TXP
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600 per day
delivery teams. The successful candidate will provide manager-level technical leadership across DevOps, cloud platforms, Infrastructure as Code, CI/CD, networking, security, observability and reliability engineering. They will help shape enterprise-scale transformation, hybrid cloud strategy and platform services aligned to the Azure Well-Architected Framework, ensuring solutions … compute/storage design. Evaluate platform changes including major provider upgrades (AzureRM/Cloudflare), DR and high availability improvements, cost optimisation strategies, and observability frameworks. Lead technical designs for large-scale refactoring and provider upgrades, environment creation pipelines, secure container registry access, identity integration and Zero Trust patterns, and event ...

AI DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
transitions from local/team execution to a centralised, hosted enterprise service. This person builds and manages the infrastructure, CI/CD, monitoring, and observability that keep the platform available, secure, and scalable across multiple teams and products. They automate deployments, maintain pipelines, and provide operational support, playing a critical … enterprise SLAs, security, and governance requirements.**Key responsibilities*** Own AI-SDLC deployment, operations, and platform reliability (uptime, performance, incident response).* Manage infrastructure, monitoring, observability, and operational/on-call support.* Build and maintain CI/CD pipelines and automation for platform releases.* Drive the shift to a centralised, hosted ...

Member of Technical Staff (AI Infrastructure Engineer)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical … workloads Experience developing APIs and managing distributed systems for both batch and real-time workloads Solid debugging and monitoring skills with expertise in observability tools for containerized environments Preferred Skills Experience with Kubernetes operators and custom controllers for ML workloads Advanced Slurm administration including multi-cluster federation and advanced scheduling ...

Permanent position: Senior Python Backend Engineer (AI Systems)

Hiring Organisation
Nicoll Curtin Technology
Location
United Kingdom
Employment Type
Permanent
Salary
GBP Annual
responsible for building and operating the infrastructure and workflows that allow AI features to execute end-to-end, with a strong focus on reliability, observability, latency, and user experience. Areas of responsibility include: Building production Back End systems that serve AI-powered product features Designing retrieval, inference, and orchestration workflows … workflows Tool calling and workflow orchestration LLM-based applications used by real users Vector databases and semantic search AI evaluation and monitoring frameworks Production observability and reliability engineering Distributed systems and high-throughput services End-to-end ownership of Back End systems in production Technical Environment Python Node.js SQL/ ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent
Salary
£75,000
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Consultant - AI Ops

Hiring Organisation
Consulting Point
Location
England, United Kingdom
closely with clients and multidisciplinary teams to assess, design, and improve cloud operating models, combining advisory capability with hands-on knowledge of IT operations, observability, service management, DevOps, and emerging AI technologies. Key Responsibilities Advise clients on cloud transformation strategy and the impact of cloud operating models on business performance … ideally in a process transformation or re-engineering environment Experience operating as an LSS Black Belt, preferably with certification Strong knowledge of IT operations, observability, and monitoring practices and tools Awareness of predictive and prescriptive analytics in an IT operations context Understanding of AI frameworks and AIOps or MLOps lifecycle ...

Senior Software Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£95,000
cloud Collaborating with engineers, designers, product managers within and outside of the team Implementing maintainable and well-tested solutions iteratively, keeping business impact and observability as a primary focus Understanding the wider context of the business and designing system architectures that meet short and long term business goals Your skillset … Algorithms, data structures Observability Web services, REST, HTTP Containers, cloud Testing, reliability, monitoring Strong knowledge in building and owning application end-to-end, from inception to maintenance, to retirement Designing systems that scale Expertise in some of our main programming languages - TypeScript, Java, Golang, Rust, Python Desirable Experience with Kubernetes ...

Java Developer with AI

Hiring Organisation
BC Forward
Location
Chicago, Illinois, United States
Employment Type
Permanent
Salary
USD 4,827 Annual
Prototype, validate, and productionize AI-enabled features with attention to latency, accuracy, cost, monitoring, fallback design, and reliability. Design for high availability, scalability, resiliency, observability, performance, security, privacy, and compliance. Integrate with enterprise systems, eventing platforms, data providers, communication providers, and vendor-managed services. Troubleshoot and resolve complex production issues … automation. Understanding of responsible and secure AI practices including data privacy, access control, guardrails, evaluation, monitoring, hallucination risk, human review, and auditability. Experience with observability and analytics tools such as Dynatrace, ELK Stack, or CloudWatch. Experience in agile environments with CI/CD, automated testing, code quality, deployment readiness ...

Senior Vice President, Senior Software Engineering, Developer Experience

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
build automation. Optimize developer workflows by streamlining build orchestration, dependency management, and CI/CD pipelines to accelerate software delivery while reducing friction. Drive observability and resilience by integrating telemetry, analytics, and automated remediation into build systems to proactively identify and resolve bottlenecks and failures. Champion secure software supply chain … Experience (DevEx), Platform Engineering, or Build/Release Engineering. Experience with artifact repositories (e.g., Artifactory, Nexus) and secure software supply chain practices. Knowledge of observability systems, telemetry frameworks, and performance optimization for build pipelines. Exposure to AI/ML-driven build acceleration and failure prediction tooling. Prior hands-on software ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
scalability. Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions. Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively. Lead cost optimisation: monitor spend, right‐size workloads, tune autoscaling, and improve infrastructure efficiency. Improve … operational experience (on‐prem and AWS EKS). Hands‐on experience defining and operating SLOs/SLIs, alerting, and incident workflows. Deep understanding of observability and telemetry (monitoring, logging, tracing). Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD. Scripting proficiency in Python, Bash ...

AI Ops Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
scalable across merchant payment use cases. You will also be accountable for the end‐to‐end engineering of GenAI and ML platforms, embedding governance, observability and operational resilience by design, hile enabling teams to deploy and run AI solutions with clarity, assurance and accountability at scale. To be successful … Bedrock for foundation models, agent orchestration patterns, Lambda and Step Functions, alongside demonstrated Python engineering capability and secure microservices and API design. AI governance, observability and cost optimisation, embedding governance by design through policy as code, alignment to model risk framework expectations, lifecycle traceability and audit‐ready evidence, supported ...

Network Automation Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self-service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with Network, Platform and Security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … protocols Experience with infrastructure-as-code and automation tools such as Ansible, Terraform and Jinja2 Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Desirable: Experience ...

Java Lead Software Engineer — Digital Markets Execution Technology, Execute

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
code written by others Leads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomes Establishes reliability goals and implements observability, resilience patterns, and operational readiness practices Leads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation … teams Preferred qualifications, capabilities, and skills Exposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace) Experience with observability stacks and resilience engineering for low‐latency/latency‐sensitive platforms Familiarity with Python Experience operating services in regulated environments with strong auditability and controls ...

Forward Deployed AI Engineer

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
enabled systems. You’ll bring deep expertise across modern full-stack technologies (.NET, Azure, SQL, React/Angular), along with experience in distributed systems, observability, and AI tooling such as LLMs, retrieval pipelines, and agentic workflows. Acting as a bridge between business and technology, you’ll work across product, data … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Principal Software Engineer (London)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
other teams shift faster and more reliably Taken AI and agentic systems from prototype to production, with a strong understanding of infrastructure, authentication, and observability Delivered on high‐stakes consulting engagements across multiple language paradigms, stacks, ecosystems, and client industries Built high‐quality, maintainable software collaboratively, incrementally, and through … legacy systems with short and long‐term business needs Led and delivered solutions to architecture‐level problems including scalability, security, reliability, performance, maintainability, and observability Facilitated alignment across technical and non‐technical stakeholders to move initiatives forward through ambiguity and complexity Provided mentorship and team support at scale, while sharing ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Head of Platform Engineering, you'll define and deliver robust platform services, establish the patterns for self-service infrastructure, CI/CD, and observability, and drive the integration of AI/ML capabilities across the organisation. You'll translate platform strategy into a coherent technical roadmap, setting standards that teams … these are sustained as the platform evolves. Defining and maintaining platform standards, patterns, and reusable components that drive consistency across teams. Leading improvements to observability, monitoring, and alerting capabilities across systems and services. Mentoring and developing engineers across the organisation, raising the collective standard of platform engineering practice. Identifying ...