3,426 to 3,450 of 4,093 Permanent Observability Jobs

Senior Applied AI Engineer — Enterprise AI Platform

Location
Greater London, England, United Kingdom
lead the delivery of enterprise AI products and platform services. You will design, build, and productionise AI systems, govern them through testing and observability, and mentor fellow engineers to raise engineering standards. You will work with Azure AI Foundry, Accio, Nexus, and other components, shaping how the team operates ...

Platform Engineer: Cloud-Native Networking Leader

Location
Greater London, England, United Kingdom
from the office five days a week in London. You will design and operate cloud-based networks, build Kubernetes operators in Golang, and improve observability and security across production systems. A strong background in cloud platforms and networking is essential. #J-18808-Ljbffr ...

Senior DevOps Engineer - Remote, FinOps & Compliance

Location
Greater London, England, United Kingdom
contribute to secure multi-account governance, IAM/SOC 2 alignment, and cost management while supporting a distributed engineering team. You will shape reliability, observability, and automation using Terraform, CDK, or Pulumi, in a fully remote setup with a global team and a strong focus on security and operational excellence. ...

Web App Developer - React/Python, Azure Cloud, Hybrid

Location
Greater London, England, United Kingdom
Vite frontends, contributing to reliability, performance, and developer experience. In this hybrid London role, you’ll collaborate with engineering teams on CI/CD, observability, security and scalable cloud deployments, with a strong focus on delivering robust software and reusable engineering patterns. #J-18808-Ljbffr ...

Engineering Lead for Acquisition & AI Growth

Location
Greater London, England, United Kingdom
raise standards while shaping how AI tools are used across the team. You’ll own end-to-end delivery, drive CI/CD and observability improvements, and balance reliability with cost across AWS services. Hybrid work is offered in a dynamic, AI-driven environment. #J-18808-Ljbffr ...

AWS Platform Engineer - Build the Next-Gen Cloud Platform

Location
Langley Mill, England, United Kingdom
contract. You will design reusable AWS capabilities and automation that empower analytics, AI and digital services across the organisation. You’ll establish security controls, observability and scalable deployment patterns, collaborating with Architecture, Security and Data Engineering teams to deliver secure, reliable infrastructure at scale. #J-18808-Ljbffr ...

Associate Site Reliability Engineer, SRE Platforms

Location
Greater London, England, United Kingdom
help build, run and continuously improve the Consolidated Trade Ledger platform in London. This role combines software and systems engineering to boost reliability, observability and incident response for a cloud-native service. The candidate will implement automation, maintain canary and blue/green deployment approaches, and contribute to scalable design ...

Lead Engineer: Greenfield Systems & Tech Strategy (Hybrid)

Location
Cardiff, Wales, United Kingdom
role focuses on Java (Spring Boot) on the backend, React/TypeScript on the frontend, PostgreSQL multitenancy, and AWS cloud hosting. You’ll drive observability and secure design across the stack. #J-18808-Ljbffr ...

Senior Data Platform Engineer - Lakehouse & Catalog

Location
Greater London, England, United Kingdom
Engineer to join the Data Lake and Catalog team in London. You will design, build, and operate scalable data lake services, focusing on reliability, observability, and end-to-end platform quality. You will mentor engineers, partner with cross-functional teams, and help evolve the lakehouse architecture using Hive, Spark ...

Senior Python Backend Engineer – Fintech API Systems

Location
Greater London, England, United Kingdom
Python, FastAPI, SQLAlchemy, and related technologies. Collaborate with cross-functional teams, own complex problems, and deliver production-ready code with strong emphasis on maintainability, observability, and reliability. #J-18808-Ljbffr ...

Senior Platform Engineer: Cloud Infra Leader (Remote)

Location
Greater London, England, United Kingdom
shared platform layer across B2B and B2C products in a high-growth fintech environment. You will build and improve scalable, secure cloud infrastructure, enhance observability and CI/CD, and collaborate with product and engineering teams to enable rapid, reliable releases. #J-18808-Ljbffr ...

Senior HPC Network Engineer — RDMA, Leaf-Spine, Equity

Location
Greater London, England, United Kingdom
architecture through day-2 operations, including RDMA fabrics, leaf-spine design, and per-tenant isolation. You’ll automate provisioning with Python and Ansible, build observability dashboards, and mentor colleagues while expanding the data-centre and office networks. #J-18808-Ljbffr ...

Forward Deployed Engineer

Hiring Organisation
Noir
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£100,000 - £140,000 per annum
Services - London/Hybrid (Tech stack: Forward Deployed Engineer, AI, Agentic AI, LLMs, LLM Integration, AI Agents, Retrieval, Tool Use, Context Orchestration, Evaluations, Guardrails, Observability, Human-in-the-Loop, Automation, Cloud, Data Management, System Integration, Forward Deployed Engineer) Are you a hands-on AI Engineer, Forward Deployed Engineer or technology … environments, combined with a strong understanding of modern AI systems and agentic workflows. Experience across LLM integration, retrieval, tool use, context orchestration, evaluations, guardrails, observability and human-in-the-loop workflows will be highly valuable. Experience within financial services, wealth management, banking, investment, risk, compliance, KYC or operations will ...

Python Engineer - AI-Driven Energy Platform & Modelling

Location
Greater London, England, United Kingdom
systems from day one, owning APIs, data pipelines, AWS infrastructure, and tooling. You’ll work across the stack, shipping features, hardening pipelines, and improving observability, with autonomy and a bias for practical design. Experience with AI tools and LLM-powered applications is a plus. #J-18808-Ljbffr ...

Distinguished Engineer - Head of Service Reliability Engineering

Hiring Organisation
Global Resourcing
Location
London, UK
Employment Type
Full-time
Head of Service Reliability Engineering, you will lead a distinct, independent capability spanning services, products, and platforms. You will make reliability, operability, and observability integral to engineering from the outset. With enterprise-wide reach and influence, you will set the direction for SRE, raise service maturity, and build a lasting … Leads the adoption of SRE practices, using SLOs, error budgets, and reliability metrics to drive measurable improvements in service performance and operational decision-making. Observability: Establishes and governs enterprise observability capabilities, using telemetry, dependency mapping, and modern monitoring platforms to improve operational insight and decision quality. Resilience and Recovery: Designs ...

Senior Custody Apps Support AVP — Global Hybrid

Location
Belfast City District, Northern Ireland, United Kingdom
seeking a Global Custody Production Support professional to ensure stability and reliability of custody and settlement platforms. You will apply SRE principles, automation, and observability to reduce toil and drive continuous service improvements. The role focuses on distributed systems, cloud-native tech, and cross-functional teamwork in a fast-paced ...

Principal Java Engineer - Payments, Low-Latency Real-Time

Location
Milton Keynes, England, United Kingdom
this hybrid role, you will deploy and monitor on OpenShift and AWS, integrate with MongoDB, and mentor peers while driving architectural decisions and observability enhancements across services. #J-18808-Ljbffr ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Birchanger, Hertfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 40,000 - 50,000 Annual
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, modern technology estate. Responsibilities … improve speed, accuracy and consistency Supporting major changes, deployments and post-incident reviews with data-driven evidence Qualifications Strong experience with monitoring and observability tools (LogicMonitor, Azure Monitor, App Insights, Log Analytics, Defender for Cloud) Excellent understanding of cloud performance, IaaS/PaaS, networking fundamentals, API performance and capacity modelling ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, South East, United Kingdom
Employment Type
Permanent
Salary
£50,000
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, modern technology estate. Responsibilities … improve speed, accuracy and consistency Supporting major changes, deployments and post-incident reviews with data-driven evidence Qualifications Strong experience with monitoring and observability tools (LogicMonitor, Azure Monitor, App Insights, Log Analytics, Defender for Cloud) Excellent understanding of cloud performance, IaaS/PaaS, networking fundamentals, API performance and capacity modelling ...

Senior Platform Engineer – Production Reliability & On-Call

Location
Greater London, England, United Kingdom
Platform/SRE role in London. You will own production reliability, participate in on‐call and incident response, and drive improvements across observability, deployments, and operational tooling. This is an ops‐heavy, hands‐on position suitable for mid‐level to senior SREs who enjoy ownership and direct impact on real ...

AI Platform Engineer for Production ML & Kubernetes

Location
Greater London, England, United Kingdom
failures in distributed systems, and guide customers through complex Kubernetes and GPU workloads in production. We seek engineers who can diagnose performance bottlenecks, improve observability, and build runbooks. This hybrid London role requires in-office presence at least 2 days/week and supports large-scale AI #J-18808-Ljbffr ...

Hybrid Technical Lead - Full-Stack Java & Microservices

Location
Lancaster, England, United Kingdom
quality engineering on a modern, scalable software platform. This full-stack leadership role demands hands-on Java expertise, Spring frameworks, microservices, Kafka, and modern observability tooling. You will mentor engineers, shape architectures, and lead agile ceremonies in a fast-paced, collaborative environment. The role supports a hybrid working model with ...

Senior Manager, Mobile Engineering

Location
City of Westminster, England, United Kingdom
Operations partners, this leader will help shape the future of Expedia's mobile ecosystem, driving engineering best practices, modern mobile architecture, release readiness, observability, and the effective adoption of AI-enabled development capabilities. The team plays a critical role in delivering seamless customer experiences while enabling the business to bring … features to market quickly and safely. Current priorities include evolving the next generation of mobile platform architecture, improving app performance and release quality, strengthening observability and operational resilience, and reducing the risk of customer-impacting issues in production. This is an opportunity to lead a highly experienced team and influence ...

Senior AI Engineer for Production-Grade LLMs

Location
Greater London, England, United Kingdom
features inside cross-functional Innovation Squads. The role focuses on production-grade AI systems, retrieval pipelines, agent workflows, and enterprise integration, with emphasis on observability and guardrails. You will collaborate with host-function stakeholders, mentor capability-building, and ensure solutions fit real workflows, not assumptions. Strong CI/ ...