76 to 100 of 119 Observability Jobs in Yorkshire

Operations Team Lead (Production & Reliability)

Location
Hull and East Yorkshire, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

IT Service Operations Manager 24 / 7

Hiring Organisation
ASDA
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
able to switch between technical detail and executive summary. Strong supplier management experience, holding partners to account for delivery. Working knowledge of monitoring and observability tools (e.g. Dynatrace, Splunk, Grafana, ServiceNow ITOM). Experience with ServiceNow or comparable ITSM tooling as an operational user and leader. Working Conditions 24/ ...

Senior Platform Owner - Payments Platform

Location
Skipton, England, United Kingdom
investment cases and budgets, making platform cost, risk and value clear to senior stakeholders. Build the future Payments Platform team and champion modern engineering, observability and continuous improvement. We work in a hybrid way, balancing flexibility with collaboration. For this role, you'll typically spend three days a week ...

IT Service Operations Manager 24/7

Location
Leeds, England, United Kingdom
able to switch between technical detail and executive summary. Strong supplier management experience, holding partners to account for delivery. Working knowledge of monitoring and observability tools (e.g. Dynatrace, Splunk, Grafana, ServiceNow ITOM). Experience with ServiceNow or comparable ITSM tooling as an operational user and leader. Working Conditions 24/ ...

Private Cloud Architect

Hiring Organisation
Experis
Location
England, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£411/day
engineering platforms used across security and data-focused teams. DevOps & SRE Build, maintain, and optimise CI/CD pipelines. Implement monitoring, alerting, logging, and observability solutions. Troubleshoot production incidents and contribute to operational support activities. Continuously improve service reliability, security, scalability, and performance. Essential Skills & Experience Strong software engineering background … Lake technologies. Experience supporting Data Engineering platforms. Understanding of API security concepts including OAuth2, JWT, and mTLS. Knowledge of Site Reliability Engineering (SRE) and observability practices. If you receive suspicious outreach claiming to be from us, please contact us via the ManpowerGroup website. ...

DevOps / SRE Engineer (Sheffield)

Location
Sheffield, England, United Kingdom
DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud‐native platform. You’ll work across the business to improve how we build, deploy and monitor our systems , while playing a key role … platforms Solid experience with Terraform and IaC automation Experience participating in or managing production incidents and on‐call Strong grasp of monitoring, alerting, and observability principles Ability to diagnose and fix complex distributed systems issues Demonstrated use of GenAI tools (ChatGPT, GitHub Copilot, Claude) in engineering workflows Excellent communication ...

Platform Engineer III - (Pipelines & Developer Experience)

Location
Leeds, England, United Kingdom
least privilege, policy‐as‐code, scanning). Improve developer experience: faster feedback loops, quality gates, ephemeral/preview environments, and great documentation. Instrument pipeline observability (Datadog or equivalent) and define SLOs (queue time, lead time, change fail rate, MTTR) to drive reliability. Automate IaC workflows (Terraform/Terragrunt) and integrate … implement and maintain CI/CD release pipelines. Experience with scripting and programming (.NET preferred; familiarity with Go, Python, PowerShell beneficial). Knowledge of observability tooling, chaos testing, and incident management. Strong analytical and problem‐solving abilities, with the capability to closely collaborate with engineering teams. Highly outcome‐oriented, pragmatic ...

Principal Software Engineer - Full Stack - AI

Location
York and North Yorkshire, England, United Kingdom
full stack, guiding the development of responsive frontend applications (React/TypeScript) and robust, scalable backend services (Python, Java, Kotlin, or Node.js). LLM Observability & Reliability: Establish robust LLM observability, evaluations, and caching, implementing latency optimisations and comprehensive monitoring (logging, usage tracking, agent behaviour). Operational Excellence: Champion high availability … performance optimisation, and observability across frontends and backend microservices, focusing on practices that maintain platform reliability and optimise MTTD and MTTR. Mentorship & Collaboration: Elevate the engineering organisation by mentoring senior and junior engineers, conducting rigorous code and system design reviews, and partnering with product managers to translate product visions into ...

SC Cleared Tester (Performance)

Hiring Organisation
VIQU IT
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£450 - £475/day Inside IR35
role of Performance Tester you will enhance and shape the scalability and reliability - creating strategies and solutions in cloud and container based environments, leveraging observability tools and diagnostics to optimise application performance. Essential Criteria • Ability to design, execute, and analyse performance, load, stress, and volume tests. • Familiarity with monitoring … observability tools (Grafana, Splunk) and diagnostics platforms (New Relic, Dynatrace). • Understanding of containerisation technologies (Docker, Kubernetes) and big data platforms (Databricks). • Ability to analyse complex performance issues and provide actionable recommendations. • Strong scripting skills (e.g., Python, JavaScript) for test automation and data analysis. • Knowledge of Performance Test Strategy ...

Senior .NET Backend Developer

Location
York and North Yorkshire, England, United Kingdom
design, implementation, testing, review, refactoring, documentation and migration, while retaining clear ownership of the outcome. Set a high standard for maintainability, automated testing, security, observability and pull-request review. Investigate complex technical and production issues and drive them through to resolution. What we are looking for Strong experience designing … clinically important data. PostgreSQL, Redis, Elasticsearch or other data and caching technologies. GraphQL, including schema design and gateway patterns. Grafana, OpenTelemetry, Prometheus or equivalent observability tooling. CI/CD, production services, microservices and message-driven systems such as RabbitMQ. How we work We value engineers who own outcomes and remain ...

Lead Cloud Engineer

Location
Leeds, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

SRE Lead: Scalable, Reliable Betting Platform (Hybrid)

Location
Leeds, England, United Kingdom
evoke is seeking a Site Reliability Engineer to join our betting and gaming platforms, focusing on reliability, scalability and performance. You will work with observability, automation and engineering to deliver robust customer experiences across services. You will lead incident response, perform postmortems and drive resilience improvements. The role requires hands ...

Full-Stack Software Engineer — Hybrid, Health Benefits

Location
Sheffield, England, United Kingdom
work with a collaborative group of engineers, define scope with the Technical Lead, and contribute to design decisions while delivering reliable, observability-driven software through frequent, #J-18808-Ljbffr ...

Oracle OCI Lead Engineer

Location
Leeds, England, United Kingdom
with industry standards. They will need to be able to design, implement and maintain OCI infrastructure, specifically platform services like identity and access management, observability, and core infrastructure, including virtual machines, storage solutions and networking components. Responsibilities include technical leadership, architectural reviews, platform support and mentoring junior engineers. Responsibilities include … SAML federation, Cloud Guard, Vault, and KMS. Network expertise – dynamic routing gateways, transit routing, domain name services, IPSec tunnels, remote peering connections, and FastConnect. Observability and monitoring – logging, monitoring, alarms, events, notifications, and external third-party feeds (e.g. Splunk). Compute, database, and storage knowledge – compute instances, patching, hardening, functions ...

Enterprise Ecommerce Solution Architect

Location
West Yorkshire, England, United Kingdom
senior leadership to drive digital transformation. Responsibilities include translating business needs into robust architectures, governing major initiatives, and reviewing designs for scalability, resilience and observability across complex platforms. #J-18808-Ljbffr ...

Head of Production Reliability & Operations

Location
Sheffield, England, United Kingdom
incidents, and build a scalable team to move from firefighting to proactive reliability engineering. The role demands hands-on leadership, setting standards, and driving observability, incident management, and on-call practices. You will partner with engineering and product to ensure uptime, performance, and predictable releases. #J-18808-Ljbffr ...

Hybrid Enterprise Integration Product Manager

Location
Sheffield, England, United Kingdom
Sheffield is seeking an Enterprise Integration Product Manager for a hybrid, contractor role. You will own a multi-quarter roadmap across middleware, messaging, and observability, driving platform standards and governance, while aligning with regulatory and regional constraints. The role requires deep IBM MQ knowledge, experience with ACE/ ...

Production Reliability Lead

Location
Leeds, England, United Kingdom
lead incidents, build a scalable team, and move us from reactive firefighting to proactive reliability engineering. You’ll own the live systems: ensure stability, observability, and safe releases, while building runbooks, incident playbooks, and a culture of ownership. Collaborate across product, engineering and support to improve MTTR, establish SLIs/ ...

Incident Commander - 24/7 Digital Reliability Lead

Location
Leeds, England, United Kingdom
roles, and coordinating cross-functional teams to drive rapid resolution for high-priority incidents. The ideal candidate has 5+ years in IT operations or observability with strong analytics, alerting, and dashboard skills, plus a background in gaming or digital commerce. #J-18808-Ljbffr ...

M365 Copilot Specialist — Enterprise Incident & IAM Expert

Location
Sheffield, England, United Kingdom
across M365 services. You will handle complex escalations in the Admin Centre, Entra, Conditional Access, SharePoint/OneDrive permissions, and Teams, while contributing to observability and monitoring efforts. The role requires 5–8+ years in M365 support, strong Copilot skills, PowerShell and Graph API for troubleshooting, and a proven track ...

Backend Engineer

Location
Leeds, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

Technical Lead - GCP

Location
Sheffield, England, United Kingdom
operational resilience. Lead solution design for new integrations and platform capabilities, balancing pace with long-term maintainability. Set engineering standards (coding, testing, documentation, observability, security) and ensure consistent adoption across the team. Provide hands‐on contribution where needed (critical ETLs, framework components, design spikes, performance fixes). ETL orchestration & data … Implement end‐to‐end data lineage (source → platform → vendor), including metadata capture and a visual representation that supports auditability and faster incident resolution. Enhance observability across pipelines and services: logging, metrics, alerting, and dashboards; drive measurable improvements in MTTR and change failure rate. Stakeholder management & ways of working Partner with ...

Senior Engineering Manager

Location
Leeds, England, United Kingdom
Job ContextEngineering • Leeds • Full-Time • HybridAre you a people-first engineering leader who thrives on turning complex technical challenges into elegant, user-centric software As our Senior Engineering Manager, you will sit at the intersection ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Senior Software Engineer

Hiring Organisation
Ask4.com
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
Salary
£60,000
Kubernetes configurations for production and non-production environments Integrate with network management systems, message brokers, and third-party APIs Instrument and monitor applications using observability tooling (metrics, logs, and traces Grafana, Prometheus, or similar) Provide technical mentorship to mid-level engineers and act as an escalation point for complex problems … Experience with message brokers and event-driven systems (NATS, RabbitMQ, Kafka, or similar) Exposure to OpenWiFi, OpenWrt or similar open-source network controller frameworks Observability experience. Use of metrics, logs, and traces using tools such as Grafana, Prometheus and Sentry Experience with AI agent development or LLM integration Agile/ ...