526 to 550 of 4,888 Permanent Observability Jobs

Python Technical Lead FinTech

Hiring Organisation
Run-Time Group Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
Python Technical Lead Salary: £90-120K work model: Hybrid Were looking for a Python Technical Lead to drive the architecture, development, and delivery of high-performance financial systems. Youll lead a team of engineers ...

Site Reliability & Observability Engineer – Datadog / Azure

Hiring Organisation
MYO Talent
Location
West Midlands, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £600.00 per day
Site Reliability & Observability Engineer/Datadog – Synthetic Monitoring, APM, RUM, Log Management, SLO’s, Alerting/Azure/Azure DevOps/Cloudflare/6-month contract/Hybrid – West Midlands/Remote/£450 – 600 per day Inside IR35. One of our leading clients is seeking a Lead Site Reliability … Observability Engineer to build and operate a world-class monitoring, synthetic testing, and reliability platform. Location – West Midlands/Remote – 5 days per week with 1-2 days per week onsite Duration – 6 months + Day rate – £450 – 600 per day Inside IR35 This role will lead the implementation ...

SRE Azure Cloud

Hiring Organisation
Sierra Business Solution
Location
Pleasanton, California, United States
Employment Type
Permanent
Salary
USD Annual
Ideal Candidate: A hands-on Senior SRE with expertise in Stage/Production deployments, Azure-based microservices, observability (Grafana, Prometheus, Graylog, Azure Monitor), troubleshooting using Application Insights, database operations, and customer escalation management, focused on maintaining highly reliable and available production services. Position Summary We are seeking a highly skilled … ensure the reliability, availability, performance, and operational stability of customer-facing applications and services. This role is focused on Stage/Production deployments, Monitoring & Observability, Troubleshooting, Incident Management, and Customer Escalation support across Azure-based microservices environments. The ideal candidate will possess strong experience in cloud operations, microservices, production support ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
play a key role in building highly reliable, scalable, and observable infrastructure. This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience. You'll have genuine ownership and the opportunity to influence … infrastructure using Terraform and Infrastructure as Code principles Develop and optimise CI/CD pipelines using GitHub Actions Build and improve comprehensive monitoring and observability across the platform Implement and maintain effective logging, metrics, tracing, alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space. This is a greenfield/brownfield … expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). ...

Site Reliability Engineer (Remote/ EU based) - Chinese Speakers

Hiring Organisation
TrueWatch
Location
Dublin, City of Dublin, Republic of Ireland
Employment Type
Permanent
Salary
£77691 - £86324/annum bonus, benefits
business provides a modern unified observability platform, helping organisations monitor complex cloud environments through data collection, visualisation and security insights. This is a strong opportunity to join at an early stage of the European growth journey, where you'll have real visibility and impact as the team scales. What … Maintain system reliability, availability, and performance for cloud infrastructure and services to ensure continuous operations Monitor production environments and manage observability tools to track metrics, logs, and alerts for proactive issue detection Support incident response by troubleshooting issues, conducting root cause analysis, and leading post-incident reviews to prevent recurrence ...

Senior Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
someone who can design, build and operate scalable systems, not just manage cloud infrastructure. You'll work across infrastructure, automation, CI/CD, observability and production reliability, while writing high-quality software in Go and helping improve the way engineering teams build and run services at scale. What will … making builds, tests and deployments faster and safer Build automation and internal tooling that removes manual toil for engineering teams Improve observability across metrics, logging, tracing and production monitoring Define and manage SLOs across latency, availability and error rates Lead incident response, triage complex production issues and ship long-term ...

Site Reliability Engineer (SRE)

Location
Cambridge, England, United Kingdom
availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications. Join Altium as a Senior Site Reliability Engineer … ensure the reliability and performance of the Altium Cloud Platforms. Key Responsibilities: Understanding how an Altium Cloud Platform works Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection. Develop and implement reliability frameworks and patterns that standardize and elevate ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU Limited
Location
Milton Keynes, Buckinghamshire, United Kingdom
Salary
£ 70 K
focused on their Azure platform, another focused on their AWS platform. Both roles require hands on experience with IaC, Containerisation and Monitoring/Observability tools. Experience required for the Senior Site Reliability Engineer: Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within … experience with both Azure, and on-premise virtual machines.Experience with Infrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.Ability ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
focused on their Azure platform, another focused on their AWS platform. Both roles require hands on experience with IaC, Containerisation and Monitoring/Observability tools. Experience required for the Senior Site Reliability Engineer: Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within … experience with both Azure, and on-premise virtual machines. Experience with Infrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent
Salary
£75,000
focused on their Azure platform, another focused on their AWS platform. Both roles require hands on experience with IaC, Containerisation and Monitoring/Observability tools. Experience required for the Senior Site Reliability Engineer: Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within … experience with both Azure, and on-premise virtual machines. Experience with Infrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways ...

Staff Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
role for someone who can design and build scalable systems, not just manage cloud infrastructure. You'll work across infrastructure, automation, CI/CD, observability and production reliability, while writing high-quality software in Go and helping set the technical standard for how engineering teams build and operate services …/CD pipelines, making builds, tests and deployments faster and safer Build automation and internal tooling that removes manual toil for engineering teams Improve observability across metrics, logging, tracing and production monitoring Define and manage SLOs across latency, availability and error rates Lead incident response, resolve complex production issues ...

Senior Forward Deployment Engineer

Location
Slough, England, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

AI Native DevOps Platform Engineer

Location
Greater London, England, United Kingdom
using AI‐first engineering practices. Working closely with Product Engineering Teams and Technical Leadership, you will build the cloud platforms, infrastructure, deployment pipelines, automation, observability frameworks, and engineering tooling that enable the rapid delivery of both AI‐powered and traditional cloud‐native applications. We are building an AI‐native engineering … platform tooling. Drive engineering productivity through AI, automation, self‐service capabilities and platform standardisation. Improve release processes and operational excellence across teams. Reliability & Observability Implement monitoring, logging, tracing, and alerting solutions. Establish platform SRE principles and operational standards. Proactively identify and resolve reliability, security, and performance issues. Lead incident response ...

Lead AI Software Engineer

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle. What you will do Lead build execution within a squad, platform capability, or engineering domain. Translate … generated and engineer‐written code to ensure correctness, maintainability, security, and alignment with specifications. Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification. Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership. Ensure AI‐generated outputs are explainable, traceable, secure ...

Lead Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, Nottinghamshire, United Kingdom
Salary
£ 70 K
Role Profile:We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a Technical Lead SRE, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. … projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design.Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.Design and evolve monitoring and alerting solutions that ...

Lead Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, UK
Employment Type
Full-time
Role Profile: We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division. As a Technical Lead SRE, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing … projects building environments, monitoring, alerting, and ensuring operational readiness from day one. Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design. Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs. Design and evolve monitoring ...

Senior Backend Engineer | AI Platform

Location
Greater London, England, United Kingdom
high degree of autonomy and ownership, as you'll be responsible for designing scalable solutions that empower multiple engineering teams while ensuring reliability, observability, and cost efficiency. What are we looking for: 5+ years of experience in Software Engineering, Backend Engineering, or Platform Engineering. Strong experience building and maintaining backend … LangChain, LangGraph, CrewAI, or similar. Experience working with cloud platforms such as Google Cloud Platform (preferred), AWS, or Azure. Strong understanding of system reliability, observability, monitoring, and incident management. Experience with Infrastructure as Code and cloud-native architectures. Previous experience working within a Platform Engineering team is a strong plus. ...

Senior Cloud SRE Lead: AWS & Kubernetes, Infra & CI/CD

Location
Greater London, England, United Kingdom
will build and maintain CI/CD pipelines using GitHub Actions, Docker, and Helm, and set up Prometheus/Grafana/Datadog for observability and incident response. Strong scripting in Python/Bash/Go is essential. #J-18808-Ljbffr ...

Lead AI Software Engineer London, United Kingdom Value Stream Engineering Posted 12 hours ago

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle.## **What you will do*** Lead build execution within a squad, platform capability, or engineering domain.* Translate … generated and engineer-written code to ensure correctness, maintainability, security, and alignment with specifications.* Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification.* Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership.* Ensure AI-generated outputs are explainable, traceable, secure ...

Staff Software Engineer-AI

Location
Greater London, England, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging ...

Staff AI Engineer, Payments Intelligence

Location
Greater London, England, United Kingdom
70+ languages, adapting to context and intent. Evaluation, safety, and guardrails — define how the unit measures agent quality and safety, with rigorous evaluation, observability, and guardrails so agents behave reliably in a compliance‐sensitive, multi‐market environment. Intelligence and data products — architect the client‐intelligence systems that turn payment data … systems at scale. Model serving and LLM inference, with the cloud infrastructure to run AI workloads in production (AWS or similar). Evaluation and observability for AI — hands‐on with tools like LangFuse, LangSmith, Braintrust, or MLflow — plus solid automated testing and CI/CD. Proven technical leadership — mentoring engineers ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to take ownership of the cloud infrastructure and DevOps practices powering our BIM Platform, working across Azure, Kubernetes, CI/CD, observability, security and developer tooling. The Role This is a hands‐on senior engineering role with broad ownership across our cloud and platform infrastructure. … Code using Terraform, ARM templates and Helm Design and improve CI/CD pipelines using GitHub Actions, enabling fast, safe and repeatable deployments Build observability across our infrastructure and services through monitoring, logging, alerting and distributed tracing Improve platform reliability, scalability and performance through capacity planning, autoscaling and resource optimisation ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...