3,051 to 3,075 of 4,044 Permanent Observability Jobs

AWS Cloud Engineer: Serverless, CI/CD & Secure APIs

Location
Nottingham, England, United Kingdom
automation, and clean code. You will collaborate with product, platform, and DevOps teams, contribute to CI/CD pipelines, and ensure robust security and observability across environments. Nottingham-based with 3 days in the office weekly. #J-18808-Ljbffr ...

Senior DevOps Engineer — Reusable CI/CD Pipelines

Location
Greater London, England, United Kingdom
collaborating with application engineers and stakeholders to migrate from legacy tooling and establish modern CI/CD practices. The role focuses on automation, security, observability and developer experience, with hands-on implementation across #J-18808-Ljbffr ...

Remote AWS SRE: Build Resilient Cloud Platforms

Location
England, United Kingdom
join a globally operating AI-driven cloud platform team. This fully remote UK role involves maintaining production systems on AWS, implementing automation and observability, and partnering with software, platform, cloud and security engineers to improve reliability. You will handle 24/7 incidents, build resilient cloud services, and drive continuous ...

Senior SRE: Hybrid Cloud Reliability Engineer

Location
Horsell, England, United Kingdom
call rotations and on-site client engagements. You will strengthen SRE capabilities through training, mentoring, and hands-on practice with OpenShift, Kubernetes, and observability stacks. #J-18808-Ljbffr ...

Azure Data Platform DevOps Engineer – SC Eligible

Location
Leeds, England, United Kingdom
SFIA Level 4) to support Azure-based data platforms for a UK public sector client. You’ll contribute to IaC, CI/CD, and observability while collaborating with senior engineers and data teams. The role emphasizes automation, platform reliability, and adherence to security and change governance within a hybrid work ...

Platform Engineer: Cloud Infra & CI/CD Champion

Location
Leeds, England, United Kingdom
/CD automation and strong security governance. The role supports cross-functional teams and requires on‐call participation. You will contribute to observability, disaster recovery and cost-efficient patterns while collaborating with delivery teams in a hybrid Leeds-based setup, with monthly office presence. #J-18808-Ljbffr ...

Lead DevOps Engineer: Site Reliability (Hybrid)

Location
Telford, England, United Kingdom
DevOps Engineer to strengthen the Customer Digital Platform. You’ll drive practical improvements, Azure API platform integration, cloud infra, CI/CD automation and observability, working with internal squads, partners and third parties. The role emphasizes deployment safety, release readiness and operational readiness, applying SRE principles to improve availability ...

Operations and SRE Manager

Location
United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement.**Responsibilities:*** Lead the implementation of the team’s strategic direction, translating priorities into clear operational plans, backlogs … improvement actions are owned, tracked and completed.* Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective.* Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support.* Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Senior AI Infra & LLM Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
Software Engineer to shape reliable AI production systems. You will build and operate large language model serving infrastructure, with cloud and Kubernetes deployments, deep observability, and cost-aware tuning. Own reliability, performance, and security across end-to-end LLM endpoints in production. You will lead incident response, capacity planning ...

Senior MariaDB Platform Engineer

Location
United Kingdom
MariaDB/MySQL deployments, tuning, and reliability at scale, working with Kubernetes and IaC to enable safe, auditable workflows. The role focuses on automation, observability, and secure practices across distributed systems. This is a high-ownership, modern platform role where you will influence on-call reliability, performance #J-18808-Ljbffr ...

Principal DevOps Engineer — AWS, Kubernetes & SaaS (Remote)

Location
Cambridge, England, United Kingdom
engineering culture. You’ll collaborate with developers, data engineers and product teams to ensure safe, rapid releases, while shaping the technical roadmap and improving observability and incident response. #J-18808-Ljbffr ...

SRE & Operations Manager — AI-Driven Reliability

Location
Sutton, England, United Kingdom
ICIS is seeking an SRE Manager to lead the reliability function for production services and guide AI-Ops initiatives. You will drive observability, incident response, and continuous improvement across internal and external customers. The role requires strong people leadership, hands-on incident involvement, and the ability to align Ops with ...

DevEx Platform Engineer — Build Tooling & Reliability

Location
Greater London, England, United Kingdom
engineer in London to own end-to-end service delivery for internal developer platforms. You will work with US/UK teams on tooling, observability, and workflows to empower engineers across the firm. Role requires hands-on experience with Kubernetes, CI/CD, and infrastructure as code, plus strong communication ...

Senior Java Platform Engineer - Hybrid London

Location
Greater London, England, United Kingdom
Java Software Engineer in London to join the Platform Team. You will work on a core platform of distributed services that ingest and transform observability data for visualisation, analytics and integrations. This is a permanent, full-time role with a hybrid working model from our London office. You will contribute ...

Tech Lead

Location
Greater London, England, United Kingdom
Product, Sales, Marketing, Support, Operations, Legal, and Compliance Mentor and coach engineers, supporting their technical growth and confidence Set standards for code quality, testing, observability, and operational excellence Collaborate with Product and Design to shape solutions, challenge assumptions, and manage trade-offs Lead technical discussions, reviews, and incident investigations Contribute … Modelling Testing Strategies System Design React PostgreSQL Soft Skills Mentoring Communication Collaboration Influence Coaching Industry Keywords Production Software Regulated Environment Operational Excellence Code Quality Observability Tools & Technologies GCP React Native Temporal AI-Assisted Engineering Tools #J-18808-Ljbffr ...

Operational and Technology Platforms Lead - ITSM - CMDB - ITIL

Hiring Organisation
Tria
Location
London, United Kingdom
Employment Type
Permanent
Define and drive platform strategy, governance and roadmaps. Standardise and rationalise platforms across multiple business divisions. Lead the development of ITSM, ITAM, CMDB and observability capabilities. Deliver greater visibility of applications, assets, licences and technology risks. Manage technology vendors, budgets and platform performance. Build and lead the Operational & Technology Platforms … platforms, technical teams or technology functions. Experience working within complex, federated or multi-business organisations. Strong knowledge of ITSM, ITAM, CMDB, endpoint management and observability tooling. A track record of platform transformation, standardisation and continuous improvement. Experience managing technology platforms within engineering, industrial, manufacturing, energy, defence or other asset-intensive ...

Senior Python Platform Lead for AI‐Driven Data Mesh

Location
Glasgow, Scotland, United Kingdom
secure, scalable delivery. You will mentor peers, collaborate across AI, product, and data science teams, and help raise the engineering bar with robust testing, observability, and enterprise-grade practices. #J-18808-Ljbffr ...

Senior Software Engineer - FinTech (TypeScript/AWS) Hybrid

Location
Greater London, England, United Kingdom
office presence. You will own services end-to-end, contribute to a fast-moving, high-trust culture, and mentor teammates while focusing on testing, observability and performance in a regulated industry. #J-18808-Ljbffr ...

Operations and SRE Manager

Location
Sutton, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement.**Responsibilities:*** Lead the implementation of the team’s strategic direction, translating priorities into clear operational plans, backlogs … improvement actions are owned, tracked and completed.* Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective.* Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support.* Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Staff Software Engineer: Cloud Networking (Remote)

Location
Hursley, England, United Kingdom
operational excellence, working across AWS, Azure, and GCP to deliver scalable networking solutions. You will architect and drive cross-team projects, ensuring reliability, observability, and cost efficiency while collaborating with product, security, and platform teams to enable seamless integration for customers. #J-18808-Ljbffr ...

Cloud Platforms Engineer: Self-Service Infra & Automation

Location
Greater London, England, United Kingdom
Azure and GCP, with a focus on self-service, guardrails and reusable components. You will own building blocks, run the cloud estate, improve pipelines, observability, security and cost governance, while collaborating with engineering and data teams. This role requires enthusiasm for learning new platforms, strong cloud and IaC experience ...

Platform Engineer: AWS/Kubernetes & Self-Service

Location
Greater London, England, United Kingdom
infrastructure. This hands-on role emphasizes on-call reliability, platforms tooling, and safe deployments. You’ll work on Terraform/Terragrunt, Kubernetes, GitOps, and observability while enhancing self-service capabilities for engineering teams #J-18808-Ljbffr ...

Senior Node.js Engineer (TypeScript) – AWS, FinTech

Location
Greater London, England, United Kingdom
design through to production support. In a highly autonomous, cross-functional squad, you’ll work with Node.js, TypeScript and AWS, delivering resilient APIs and observability-focused systems. If you enjoy ownership and modern cloud architectures, we’d love to hear from you. #J-18808-Ljbffr ...

Senior SRE Lead for AI/ML Data Platform

Location
Cumbernauld, Scotland, United Kingdom
engineers, implement best-practice SRE processes, and collaborate with cross-functional teams to deliver secure, scalable solutions. The role emphasizes automation, observability, and incident response. #J-18808-Ljbffr ...

Azure CloudOps Engineer — Self-Healing, IaC & CI/CD

Location
Greater London, England, United Kingdom
seeking an Azure CloudOps Engineer to design, deploy, and operate resilient cloud platforms across Azure and hybrid environments. The role emphasizes automation, IaC, observability, and AI-assisted operations to improve reliability and efficiency. You will build automated pipelines, implement self-healing mechanisms, and collaborate with cross-functional teams to drive ...