3,251 to 3,275 of 4,291 Permanent Observability Jobs

Lead SRE: AWS & Python for Scalable Reliability

Location
Glasgow, Scotland, United Kingdom
availability, performance, and resilience of production systems serving millions globally. You will embed reliability in the software lifecycle, guide cross-functional teams, and champion observability across monitoring and incident practices. The role emphasizes leadership in incident response, SLOs, and tooling, with a focus on engineering excellence and scalable, secure operations ...

Senior Platform Engineer – Remote UK (Cloud/SRE)

Location
Greater London, England, United Kingdom
Hudl in London, United Kingdom, is seeking a Senior Engineer to join our Platform Engineering team. You’ll work on site reliability, cloud infrastructure, observability and production operations to keep Hudl’s platform highly available, scalable and secure. You’ll lead with technical excellence, mentor engineers and drive innovation using ...

Remote NOC Engineer: Cloud Reliability & Automation (UK)

Location
West of England, England, United Kingdom
Ideal candidates will have Linux administration, AWS experience, and hands-on Terraform/Docker work, plus scripting in Python, Bash or Go and strong observability tooling knowledge. #J-18808-Ljbffr ...

Azure Platform Architect: Multi-Tenant AKS

Location
Greater London, England, United Kingdom
upskilling engineers. You will lead platform operating models, SPI communications, and best-practice cloud-native patterns. You will shape the architecture, governance, and observability stack for scalable, multi-tenant workloads in production, with a strong emphasis on IaC, GitOps, and secure, compliant design. #J-18808-Ljbffr ...

Senior Ruby on Rails Engineer | React & Python Leader

Location
Greater London, England, United Kingdom
ensure deliverables are simple, maintainable, and scalable. The role emphasizes owning code quality, API contracts, and system documentation, with responsibility for performance monitoring and observability to maintain reliability across services. #J-18808-Ljbffr ...

Software Reliability Engineer - DevOps & CI/CD Leader

Location
Greater London, England, United Kingdom
coding and configuration, guiding development teams toward rapid and secure releases. The position emphasizes practical DevOps and SRE principles, with focus on testing, automation, observability, and modern runtimes. Strong collaboration across teams is essential for success. #J-18808-Ljbffr ...

Cloud Data Platform Engineer — Real-Time Streaming

Location
Greater London, England, United Kingdom
focus on streaming data, cloud infrastructure, and automation. The role involves creating CI/CD pipelines, integrating data clouds with databases, and enforcing observability and governance across data pipelines. The ideal candidate has 3+ years in data platform or infrastructure roles, strong AWS and Terraform skills, and experience with Kafka ...

Remote Cloud Reliability Architect (Java/C#)

Location
United Kingdom
Bristol or London, aligning with SRE and backend engineering standards. You'll partner with Tech Operations and broader engineering teams to improve deployment safety, observability, and capacity planning, while mentoring staff and delivering maintainable code in Java or C#. #J-18808-Ljbffr ...

Cloud FinOps Analyst

Hiring Organisation
Manufacturing Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£60,000
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Cloud Platform Operations Lead: Azure, Kubernetes & SRE

Location
Greater London, England, United Kingdom
will lead hands-on engineering and inspire a team of Platform Operations engineers. You'll shape platform strategy, manage incident response, and advance observability, automation, and governance across infrastructure, applications, and data. Remote/hybrid working options available with global teams. #J-18808-Ljbffr ...

Lead Backend Engineer - Java & Go for Enterprise Automation

Location
Bournemouth, England, United Kingdom
EPAS products like AaaS, Ansible Automation Platform, and AutoM8. You will lead architecture work, drive automation strategy, and mentor teams in secure coding and observability practices. The role emphasizes enterprise-grade automation, AI-assisted development, and cross-product platform engineering to support Day 2 operations and platform #J-18808-Ljbffr ...

Senior SRE: Front‐Office Trading Reliability

Location
Greater London, England, United Kingdom
within its Trading Technology group in London. You will embed with the software engineering team that builds front‐office trading platforms, shaping SRE patterns, observability, and resilience across globally distributed systems and AI‐accelerated incident response. You will partner with traders and senior stakeholders, lead incident responses, implement reliable code ...

Senior Data Engineer — AI-Driven Data Platform

Location
Greater London, England, United Kingdom
Product to ensure data infrastructure supports analytics, experimentation, and decision-making. This hands-on role emphasizes code quality, CI/CD, security, and observability, with a focus on knowledge sharing and mentoring across the team. #J-18808-Ljbffr ...

(senior) Devops Engineer (m/w/d)

Hiring Organisation
iVentureGroup GmbH
Location
Hammerbrook, Hamburg, Germany
Employment Type
Permanent
Salary
EUR Annual
Verantwortung für unseren operativen IT-Betrieb (24/7), während du gleichzeitig moderne Plattform-Initiativen vorantreibst. Ob Kubernetes-Cluster, CI/CD-Pipelines oder Observability - du bist in deinem Element, wenn du Systeme stabil hältst und . click apply for full job details ...

AWS Cloud Engineer: Serverless, CI/CD & Secure APIs

Location
Nottingham, England, United Kingdom
automation, and clean code. You will collaborate with product, platform, and DevOps teams, contribute to CI/CD pipelines, and ensure robust security and observability across environments. Nottingham-based with 3 days in the office weekly. #J-18808-Ljbffr ...

Automation QA Engineer (SDET) – AI-Driven CI/CD

Location
Greater London, England, United Kingdom
quality-focused software engineer to build AI-assisted quality workflows and scalable test automation. You will work across C#, TypeScript, APIs, data pipelines, and observability to reduce manual checks and improve release confidence. You will contribute to test strategy, migrate automation from Selenium to Playwright, and integrate tests into ...

Senior DevOps Engineer — Reusable CI/CD Pipelines

Location
Greater London, England, United Kingdom
collaborating with application engineers and stakeholders to migrate from legacy tooling and establish modern CI/CD practices. The role focuses on automation, security, observability and developer experience, with hands-on implementation across #J-18808-Ljbffr ...

Remote AWS SRE: Build Resilient Cloud Platforms

Location
England, United Kingdom
join a globally operating AI-driven cloud platform team. This fully remote UK role involves maintaining production systems on AWS, implementing automation and observability, and partnering with software, platform, cloud and security engineers to improve reliability. You will handle 24/7 incidents, build resilient cloud services, and drive continuous ...

Azure Data Platform DevOps Engineer – SC Eligible

Location
Leeds, England, United Kingdom
SFIA Level 4) to support Azure-based data platforms for a UK public sector client. You’ll contribute to IaC, CI/CD, and observability while collaborating with senior engineers and data teams. The role emphasizes automation, platform reliability, and adherence to security and change governance within a hybrid work ...

Senior SRE: Hybrid Cloud Reliability Engineer

Location
Horsell, England, United Kingdom
call rotations and on-site client engagements. You will strengthen SRE capabilities through training, mentoring, and hands-on practice with OpenShift, Kubernetes, and observability stacks. #J-18808-Ljbffr ...

Platform Engineer: Cloud Infra & CI/CD Champion

Location
Leeds, England, United Kingdom
/CD automation and strong security governance. The role supports cross-functional teams and requires on‐call participation. You will contribute to observability, disaster recovery and cost-efficient patterns while collaborating with delivery teams in a hybrid Leeds-based setup, with monthly office presence. #J-18808-Ljbffr ...

Lead DevOps Engineer: Site Reliability (Hybrid)

Location
Telford, England, United Kingdom
DevOps Engineer to strengthen the Customer Digital Platform. You’ll drive practical improvements, Azure API platform integration, cloud infra, CI/CD automation and observability, working with internal squads, partners and third parties. The role emphasizes deployment safety, release readiness and operational readiness, applying SRE principles to improve availability ...

Operations and SRE Manager

Location
United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement.**Responsibilities:*** Lead the implementation of the team’s strategic direction, translating priorities into clear operational plans, backlogs … improvement actions are owned, tracked and completed.* Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective.* Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support.* Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Senior AI Infra & LLM Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
Software Engineer to shape reliable AI production systems. You will build and operate large language model serving infrastructure, with cloud and Kubernetes deployments, deep observability, and cost-aware tuning. Own reliability, performance, and security across end-to-end LLM endpoints in production. You will lead incident response, capacity planning ...

Principal DevOps Engineer — AWS, Kubernetes & SaaS (Remote)

Location
Cambridge, England, United Kingdom
engineering culture. You’ll collaborate with developers, data engineers and product teams to ensure safe, rapid releases, while shaping the technical roadmap and improving observability and incident response. #J-18808-Ljbffr ...