3,976 to 4,000 of 4,178 Observability Jobs

Production Reliability Team Lead

Location
Normanton on Trent, England, United Kingdom
Lead to own production, building a scalable system for reliability. You will lead operational excellence across all live customer-facing systems to ensure reliability, observability, and continuous improvement. This hands-on leadership role requires shaping processes, leading incidents, growing the team, and transitioning from reactive firefighting to proactive reliability engineering. ...

Azure Technical Architect - Cloud Security (Remote)

Location
United Kingdom
with delivery teams to drive secure, scalable solutions. You will design and evolve enterprise-scale foundations, ensure best practices in identity, networking, security, and observability, and maintain ADRs and platform documentation to support rapid, compliant #J-18808-Ljbffr ...

Staff Engineer - Embedded Accountancy Platform Lead

Location
Norwich, England, United Kingdom
collaborate with external partners to deliver scalable financial tooling. As a platform-focused leader, you will ensure robust integration patterns and high standards for observability, performance, and security across the product ecosystem. #J-18808-Ljbffr ...

Senior IT Operations Analyst & Incident Lead

Location
City of Edinburgh, Scotland, United Kingdom
environments. Based in Edinburgh, the role offers exposure to a global technology-driven transformation within a major banking brand, with a focus on service observability and risk-aware operations. #J-18808-Ljbffr ...

Mid-Market AI Solutions Account Executive

Location
United Kingdom
assigned territory. You will own the full sales cycle—outreach to close—advocating a value-based approach for our AI-powered search, observability, and security platform. You’ll engage stakeholders, articulate the business value, and navigate complex deals. #J-18808-Ljbffr ...

Linux Automation Engineer — Hybrid (Glasgow)

Location
Paisley, Scotland, United Kingdom
London, to support cross-site initiatives. You will develop automation for patching, upgrades and changes across Linux, VMware and F5 environments, and contribute to observability, validation, and continuous improvement of the platform. #J-18808-Ljbffr ...

Cloud HPC Infrastructure Engineer (Kubernetes)

Location
Cambridge, England, United Kingdom
with burst capability for peak demand. You will simplify access and ensure reliability for the data science team. You will own the compute platform, observability, data infrastructure, and security, embracing automation, robust tooling, and zero‐trust networking to deliver a #J-18808-Ljbffr ...

Senior Infrastructure Engineer: Go & Distributed Systems

Location
United Kingdom
SmartNIC/DPU environments powering modern AI infrastructure. This is a highly hands-on technical leadership role with ownership across architecture, implementation, reliability, observability, and operational excellence. A deep systems engineering background and strong Linux networking knowledge are required to drive #J-18808-Ljbffr ...

Senior Mobile Engineering Leader

Location
Greater London, England, United Kingdom
team to deliver high-quality mobile experiences. You will partner with Product, Design, Architecture, Security, and Operations to shape roadmaps, manage dependencies, and improve observability and release quality, while safely integrating AI-enabled capabilities into #J-18808-Ljbffr ...

Mid‐Market AI Search & Cloud Sales Executive

Location
United Kingdom
footprint in existing customers and new logos. You will own the full sales cycle—from outreach to close—driving adoption of AI-powered search, Observability and Security solutions while articulating value to stakeholders. You bring a proven SaaS Mid-Market track record, cloud fluency, and a passion for building trusted ...

Site Reliability Engineer — Production & Incident Response

Location
Greater London, England, United Kingdom
incident response and oversee the offshore Production Support team in London. This hybrid position will require you to ensure system reliability and maintain high observability standards. Ideal candidates will bring significant experience in SRE, embracing automation tools to enhance performance. You will be instrumental in protecting client interests and contributing ...

Staff Product Manager, OpenTelemetry Experience (Remote UK)

Location
United Kingdom
coordinating across multiple product teams. You will work remotely with globally distributed teams, applying customer research and data to drive activation and adoption of observability solutions. Strong AI and tooling knowledge will be leveraged to accelerate outcomes. #J-18808-Ljbffr ...

Data Scientist - ML Research & Production for Payments

Location
Greater London, England, United Kingdom
estimators that boost payment performance across our merchants portfolio. You will collaborate with Data Scientists, Product, and Engineering to deploy robust features and maintain observability for safe model launches. You will design experiments, handle time‐based data leakage, and communicate results clearly, leveraging Python for model training and productionizing features ...

Operations Team Lead - Scale Production & Reliability

Location
Wolverhampton, England, United Kingdom
building a scalable, reliable system that can grow with the business. You’ll lead operational excellence across all live customer‐facing systems, ensuring reliability, observability, and continuous improvement. You will shape processes, manage incidents, and build a high‐performing team that moves from firefighting to proactive reliability engineering. #J ...

Solutions Architect - Ecommerce

Hiring Organisation
Outsource
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
governance across major initiatives. Reviewing technical designs produced by engineering teams. Working across complex applications, platforms and integrations. Ensuring solutions consider scalability, resilience, performance, observability and operational requirements. Supporting major technology transformation programmes. You'll ideally have: Proven Solution Architecture experience within a large enterprise environment. Experience designing customer-facing … complex, integrated technology landscapes. Strong stakeholder management and communication skills. Experience making architectural decisions and influencing technical direction. A good understanding of scalability, resilience, observability and operational considerations. The ability to bridge the gap between business requirements and technology solutions. Why Join? You'll be joining a mature architecture function ...

SRE Manager: Platform Reliability Leader

Location
Greater London, England, United Kingdom
team, drive reliability and scalability, and shape the platform roadmap. The role combines technical leadership with people management to elevate our cloud services and observability practices across the business. You’ll partner with engineering, security and product teams, champion best practices, and deliver resilient systems that support millions of customers ...

Remote SRE-NOC Engineer: Automate, Observe, Resolve

Location
United Kingdom
NiCE is seeking an experienced SRE – NOC to join our 24/7 operations. You will own incident response, runbooks, observability, and automation to improve service reliability across Linux-based systems and cloud platforms. You will collaborate with engineering, security, and product teams, design robust monitoring, and implement self‐healing ...

Lead, Production Reliability & Operations

Location
South Shields, England, United Kingdom
Operations Team Lead to own production and build scalable systems. You will lead operational excellence across all live customer-facing platforms, aiming for reliability, observability, and continuous improvement. This hands-on role requires shaping processes, incident management, and building a high-performing team. You will drive a culture of ownership ...

SRE Associate: Build Scalable, Reliable Systems

Location
Birmingham, England, United Kingdom
will join Core Engineering to build, run, and scale high-performing, fault-tolerant systems and automate operations to reduce toil. The role emphasizes observability, reliability, and collaboration with software teams. Responsibilities include proactive production health monitoring, capacity planning, incident management, and developing self-healing capabilities. #J-18808-Ljbffr ...

Lead Analytics Engineer — Build a Trusted Data Layer

Location
Greater London, England, United Kingdom
decisions for a scale-up with a trans-continental expansion. The role emphasizes building the golden layer with dbt, Python, and BigQuery, while ensuring observability, reliability, and cost controls. You will mentor growing teams and collaborate across Product, Engineering and Brand. #J-18808-Ljbffr ...

Kubernetes & HPC Infra Engineer

Location
Cambridge, England, United Kingdom
spans on‐prem and cloud, unified under Kubernetes, with emphasis on security and a reliable, self‐service platform. You will own the compute platform, observability, data infrastructure, and security, enabling scalable, automated workflows while maintaining strong safeguards and ease of use for the team. #J-18808-Ljbffr ...

Production & Reliability Lead

Location
West of England, England, United Kingdom
Complexio is seeking an Operations Team Lead to own production. You’ll lead operational excellence across all live customer-facing systems, ensuring reliability, observability, and continuous improvement. This hands-on role involves shaping processes, leading incidents, building the team, and moving from firefighting to proactive reliability engineering. ...

Production Reliability Lead

Location
Slough, England, United Kingdom
change coordination. In this hands-on leadership role, you’ll build a scalable operating model, foster strong runbooks, and drive continuous improvement across observability, MTTR, and incident prevention #J-18808-Ljbffr ...