3,476 to 3,500 of 4,168 Observability Jobs

Associate Site Reliability Engineer, SRE Platforms

Location
Greater London, England, United Kingdom
help build, run and continuously improve the Consolidated Trade Ledger platform in London. This role combines software and systems engineering to boost reliability, observability and incident response for a cloud-native service. The candidate will implement automation, maintain canary and blue/green deployment approaches, and contribute to scalable design ...

Senior Data Platform Engineer - Lakehouse & Catalog

Location
Greater London, England, United Kingdom
Engineer to join the Data Lake and Catalog team in London. You will design, build, and operate scalable data lake services, focusing on reliability, observability, and end-to-end platform quality. You will mentor engineers, partner with cross-functional teams, and help evolve the lakehouse architecture using Hive, Spark ...

Lead Engineer: Greenfield Systems & Tech Strategy (Hybrid)

Location
Cardiff, Wales, United Kingdom
role focuses on Java (Spring Boot) on the backend, React/TypeScript on the frontend, PostgreSQL multitenancy, and AWS cloud hosting. You’ll drive observability and secure design across the stack. #J-18808-Ljbffr ...

Senior Python Backend Engineer – Fintech API Systems

Location
Greater London, England, United Kingdom
Python, FastAPI, SQLAlchemy, and related technologies. Collaborate with cross-functional teams, own complex problems, and deliver production-ready code with strong emphasis on maintainability, observability, and reliability. #J-18808-Ljbffr ...

Senior Platform Engineer: Cloud Infra Leader (Remote)

Location
Greater London, England, United Kingdom
shared platform layer across B2B and B2C products in a high-growth fintech environment. You will build and improve scalable, secure cloud infrastructure, enhance observability and CI/CD, and collaborate with product and engineering teams to enable rapid, reliable releases. #J-18808-Ljbffr ...

Senior HPC Network Engineer — RDMA, Leaf-Spine, Equity

Location
Greater London, England, United Kingdom
architecture through day-2 operations, including RDMA fabrics, leaf-spine design, and per-tenant isolation. You’ll automate provisioning with Python and Ansible, build observability dashboards, and mentor colleagues while expanding the data-centre and office networks. #J-18808-Ljbffr ...

Forward Deployed Engineer

Hiring Organisation
Noir
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£100,000 - £140,000 per annum
Services - London/Hybrid (Tech stack: Forward Deployed Engineer, AI, Agentic AI, LLMs, LLM Integration, AI Agents, Retrieval, Tool Use, Context Orchestration, Evaluations, Guardrails, Observability, Human-in-the-Loop, Automation, Cloud, Data Management, System Integration, Forward Deployed Engineer) Are you a hands-on AI Engineer, Forward Deployed Engineer or technology … environments, combined with a strong understanding of modern AI systems and agentic workflows. Experience across LLM integration, retrieval, tool use, context orchestration, evaluations, guardrails, observability and human-in-the-loop workflows will be highly valuable. Experience within financial services, wealth management, banking, investment, risk, compliance, KYC or operations will ...

Python Engineer - AI-Driven Energy Platform & Modelling

Location
Greater London, England, United Kingdom
systems from day one, owning APIs, data pipelines, AWS infrastructure, and tooling. You’ll work across the stack, shipping features, hardening pipelines, and improving observability, with autonomy and a bias for practical design. Experience with AI tools and LLM-powered applications is a plus. #J-18808-Ljbffr ...

Distinguished Engineer - Head of Service Reliability Engineering

Hiring Organisation
Global Resourcing
Location
London, UK
Employment Type
Full-time
Head of Service Reliability Engineering, you will lead a distinct, independent capability spanning services, products, and platforms. You will make reliability, operability, and observability integral to engineering from the outset. With enterprise-wide reach and influence, you will set the direction for SRE, raise service maturity, and build a lasting … Leads the adoption of SRE practices, using SLOs, error budgets, and reliability metrics to drive measurable improvements in service performance and operational decision-making. Observability: Establishes and governs enterprise observability capabilities, using telemetry, dependency mapping, and modern monitoring platforms to improve operational insight and decision quality. Resilience and Recovery: Designs ...

Senior Custody Apps Support AVP — Global Hybrid

Location
Belfast City District, Northern Ireland, United Kingdom
seeking a Global Custody Production Support professional to ensure stability and reliability of custody and settlement platforms. You will apply SRE principles, automation, and observability to reduce toil and drive continuous service improvements. The role focuses on distributed systems, cloud-native tech, and cross-functional teamwork in a fast-paced ...

Principal Java Engineer - Payments, Low-Latency Real-Time

Location
Milton Keynes, England, United Kingdom
this hybrid role, you will deploy and monitor on OpenShift and AWS, integrate with MongoDB, and mentor peers while driving architectural decisions and observability enhancements across services. #J-18808-Ljbffr ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Birchanger, Hertfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 40,000 - 50,000 Annual
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, modern technology estate. Responsibilities … improve speed, accuracy and consistency Supporting major changes, deployments and post-incident reviews with data-driven evidence Qualifications Strong experience with monitoring and observability tools (LogicMonitor, Azure Monitor, App Insights, Log Analytics, Defender for Cloud) Excellent understanding of cloud performance, IaaS/PaaS, networking fundamentals, API performance and capacity modelling ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, South East, United Kingdom
Employment Type
Permanent
Salary
£50,000
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, modern technology estate. Responsibilities … improve speed, accuracy and consistency Supporting major changes, deployments and post-incident reviews with data-driven evidence Qualifications Strong experience with monitoring and observability tools (LogicMonitor, Azure Monitor, App Insights, Log Analytics, Defender for Cloud) Excellent understanding of cloud performance, IaaS/PaaS, networking fundamentals, API performance and capacity modelling ...

Senior Platform Engineer – Production Reliability & On-Call

Location
Greater London, England, United Kingdom
Platform/SRE role in London. You will own production reliability, participate in on‐call and incident response, and drive improvements across observability, deployments, and operational tooling. This is an ops‐heavy, hands‐on position suitable for mid‐level to senior SREs who enjoy ownership and direct impact on real ...

AI Platform Engineer for Production ML & Kubernetes

Location
Greater London, England, United Kingdom
failures in distributed systems, and guide customers through complex Kubernetes and GPU workloads in production. We seek engineers who can diagnose performance bottlenecks, improve observability, and build runbooks. This hybrid London role requires in-office presence at least 2 days/week and supports large-scale AI #J-18808-Ljbffr ...

Hybrid Technical Lead - Full-Stack Java & Microservices

Location
Lancaster, England, United Kingdom
quality engineering on a modern, scalable software platform. This full-stack leadership role demands hands-on Java expertise, Spring frameworks, microservices, Kafka, and modern observability tooling. You will mentor engineers, shape architectures, and lead agile ceremonies in a fast-paced, collaborative environment. The role supports a hybrid working model with ...

Senior Manager, Mobile Engineering

Location
City of Westminster, England, United Kingdom
Operations partners, this leader will help shape the future of Expedia's mobile ecosystem, driving engineering best practices, modern mobile architecture, release readiness, observability, and the effective adoption of AI-enabled development capabilities. The team plays a critical role in delivering seamless customer experiences while enabling the business to bring … features to market quickly and safely. Current priorities include evolving the next generation of mobile platform architecture, improving app performance and release quality, strengthening observability and operational resilience, and reducing the risk of customer-impacting issues in production. This is an opportunity to lead a highly experienced team and influence ...

Remote SRE: Cloud Reliability Engineer (AI Tools)

Location
Hemel Hempstead, England, United Kingdom
holiday operator, is seeking an experienced Site Reliability Engineer to join our Product Technology team. This role focuses on cloud reliability, CI/CD, observability and database resilience across diverse engines, with occasional travel to Hemel Hempstead and off-site events. You will work with engineering teams to design, implement ...

Senior AI Engineer for Production-Grade LLMs

Location
Greater London, England, United Kingdom
features inside cross-functional Innovation Squads. The role focuses on production-grade AI systems, retrieval pipelines, agent workflows, and enterprise integration, with emphasis on observability and guardrails. You will collaborate with host-function stakeholders, mentor capability-building, and ensure solutions fit real workflows, not assumptions. Strong CI/ ...

Senior Data Engineer: Scale Pipelines & AI-Driven Data

Location
Greater London, England, United Kingdom
hands-on role, you’ll mentor peers through code reviews and knowledge sharing while promoting best practices and an emphasis on CI/CD, observability, and IaC. The team operates in a hybrid model across the UK. #J-18808-Ljbffr ...

SRE Associate: Build Reliable Cloud Platforms

Location
Birmingham, England, United Kingdom
seeking a Site Reliability Engineer to join Core Engineering. You will help build, run and maintain high-performing, distributed systems, focusing on reliability, observability and automation across critical services. Responsibilities include monitoring production services, capacity planning, incident response and driving improvements to SLIs/SLOs. The role requires ...

Remote SRE: Cloud Reliability & AI-Driven Ops

Location
Hemel Hempstead, England, United Kingdom
Haven is seeking a hands-on Site Reliability Engineer to join our Product Technology team. This remote-first role involves shaping CI/CD, observability, and incident response while collaborating with engineers and tech leads to ensure reliable, scalable platforms for guests and colleagues. You’ll tackle infrastructure design, tooling ...

Operations Team Lead: Scale Reliable Production

Location
Guildford, England, United Kingdom
Complexio is seeking an Operations Team Lead to own and scale production across live customer-facing systems. You will drive reliability, observability, and continuous improvement, shaping processes, leading incidents, and building a high-performing team. This hands-on role requires deep SRE/DevOps expertise, leadership, and a calm, clear ...

Senior Go Software Engineer — Low-Latency Trading Platform

Location
Greater London, England, United Kingdom
develops our trading platform. You will architect and implement features to meet performance, stability, and product goals while upholding technical standards, CI/CD, observability and security. Our Golang-centric stack handles high-throughput, low-latency data with multiple trading venues and global data feeds. The role is based ...

Senior Full-Stack Engineer — Backend-Heavy with AI Impact

Location
Greater London, England, United Kingdom
deliver measurable impact at scale. This is a remote, internationally oriented role with opportunities across 175 countries. You’ll contribute to architectural direction, champion observability, and foster engineering excellence while balancing short-term delivery with long-term #J-18808-Ljbffr ...