351 to 375 of 445 Observability Jobs in the North of England

Vice President, Production Services Application Support

Location
Manchester, England, United Kingdom
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe ...

Lead Cloud Site Reliability Engineer

Location
Manchester, England, United Kingdom
looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards, improve platform reliability and reduce operational complexity. Working closely with Product Owners … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Lead Cloud Site Reliability Engineer

Location
Manchester, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Senior Software Engineer

Hiring Organisation
Ask4.com
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
Salary
£60,000
Kubernetes configurations for production and non-production environments Integrate with network management systems, message brokers, and third-party APIs Instrument and monitor applications using observability tooling (metrics, logs, and traces Grafana, Prometheus, or similar) Provide technical mentorship to mid-level engineers and act as an escalation point for complex problems … Experience with message brokers and event-driven systems (NATS, RabbitMQ, Kafka, or similar) Exposure to OpenWiFi, OpenWrt or similar open-source network controller frameworks Observability experience. Use of metrics, logs, and traces using tools such as Grafana, Prometheus and Sentry Experience with AI agent development or LLM integration Agile/ ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Deep understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Strong analytical and problem-solving skills with the ability to identify root causes ...

Lead Cloud Site Reliability Engineer

Hiring Organisation
Lloyds Banking Group
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 80 K
excellence to deliver secure, resilient and scalable services for millions of customers.We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments.You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … environment that supports learning, collaboration and continuous improvement.Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery.Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk.Lead incident and problem management activities, promoting effective root cause analysis ...

Lead Cloud Site Reliability Engineer

Hiring Organisation
Lloyds Banking Group
Location
Leeds, West Yorkshire, United Kingdom
Salary
£ 100 K
excellence to deliver secure, resilient and scalable services for millions of customers.We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments.You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … environment that supports learning, collaboration and continuous improvement.Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery.Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk.Lead incident and problem management activities, promoting effective root cause analysis ...

Senior Platform Engineer

Location
Salford, England, United Kingdom
strong communicator who enjoys sharing knowledge, conducting code reviews, and helping the wider team level up. Problem Solver: You thrive on optimizing observability and performance to keep a cloud footprint lean and high-performing. What will you be doing? You will be responsible for the "how" of our engineering delivery … Building self-service tools and documentation that empower squads to own their services from end to end. Driving Operational Excellence: Spearheading cost-optimization and observability initiatives to ensure our cloud spend remains efficient as we scale. Defining "Golden Paths": Contributing to the standardized engineering patterns that define excellence at Finova ...

Principal AI Engineer - Microsoft Azure AI Foundry (Contract)

Location
Leeds, England, United Kingdom
tooling. Define reusable architecture patterns for AI workloads, including development, testing and production environments. Establish platform standards covering resource structure, environments, deployment patterns, observability and operational management. Work closely with AI Engineers to design and deliver AI Agents and agentic workflows that meet business and technical requirements. Governance, Risk & Compliance … provisioning and configuration using Bicep, Terraform or ARM templates. Build automated processes for prompt‐flow evaluation, model deployment, testing and release management. Establish platform observability covering availability, performance, usage, cost and AI workload health. Experience 5+ years' experience in Azure cloud architecture, engineering or platform engineering. 1–2+ years' experience ...

Senior Cloud SRE Lead: Reliability & Observability

Location
Manchester, England, United Kingdom
Lloyds Banking Group is seeking a Site Reliability Engineer Lead to strengthen reliability, observability and operational excellence across Azure and GCP environments. You will lead a team of SREs, shaping engineering standards and reducing operational complexity, while collaborating with product and platform teams to improve design, operation and continuous improvement ...

Data Architect

Location
Manchester, England, United Kingdom
We believe in the power of ingenuity to build a positive human future.We challenge where it matters and own the outcome.As strategies, technologies, and innovation collide, we create opportunity from complexity. Our teams of interdisciplinary ...

Head of Digital Technology

Location
Manchester, England, United Kingdom
We’re Pret: proud makers of freshly made food, organic coffee, and big ideas. Across 750+ shops and 20+ countries, our teams are shaping the future of Pret through innovation, inclusion, great customer service and ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
engineered into every service throughout its lifecycle. Working alongside the Director of Site Reliability Operations, this leader will define the engineering standards, automation, observability, production readiness, and resilience capabilities that enable world‐class operational performance. While Site Reliability Operations owns the day‐to‐day operation of production services, the Site … global Site Reliability Engineering organization. This role owns the engineering strategy, governance, architecture, and technical practices that improve service reliability through software engineering, observability, automation, resilience engineering, production engineering, and operational readiness. Rather than operating production systems on a day‐to‐day basis, the Site Reliability Engineering organization develops ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
daily and process more than 1.5 million bets per hour at peak. Job Description As a Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. You will have strong software engineering skills, approaching system reliability … observability as a software problem — protecting, providing for, and progressing the performance and availability of our critical systems. Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices. You will leverage AI tools and LLM platforms in your daily work ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 05/08/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
urgency, recover priority incidents under pressure, and maintain core support coverage across on-site and offshore support hours. Use SQL scripting, automation, monitoring, and observability tools to improve operational resilience, service health, reliability, and incident response. To be successful in this role, were seeking the following: Excellent SQL scripting skills. … solutions for alert correlation, anomaly detection, predictive monitoring, and service optimisation. Strong understanding of Site Reliability Engineering (SRE) principles, including service health, reliability, availability, observability, incident reduction, and continuous service improvement. Experience with SRE practices such as monitoring and alert tuning, incident management, post-incident reviews, root cause analysis ...

Senior Architect Private Cloud

Hiring Organisation
Randstad Technologies Recruitment
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £550/day Inside IR35 via Umbrella
patterns, and platform guardrails. Kubernetes Platform Leadership: Own the end-to-end Kubernetes ecosystem, including cluster topology, multi-tenancy, ingress, service mesh, secrets management, observability, and seamless workload onboarding. Event-Driven Architecture: Direct the adoption and governance of Apache Kafka, covering partitioning strategies, schema management, resilience patterns, capacity planning … configuration management. Desirable Skills: Familiarity with Service Mesh (e.g., Istio), API Gateways, Policy as Code (e.g., OPA), developer portals (Platform-as-a-Product), and Observability stacks (OpenTelemetry, Prometheus, ELK). Randstad Technologies is acting as an Employment Business in relation to this vacancy. ...

Senior Vice President, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise-scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. To be successful in this role, were seeking … platform design and optimisation. Proven track record of delivering production-grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands-on technical leader who can actively contribute to solution design and critical builds, while defining ...

Senior Software Engineer

Location
Manchester, England, United Kingdom
transformation and delivery Contributing to the evolution of a modern cloud-native data platform Defining and promoting standards around data quality, governance, metadata and observability Working closely with software engineering, platform and product teams to improve data accessibility and reliability Supporting production data systems, monitoring and operational excellence Contributing … data processing and large-scale datasets Understanding of REST APIs and data integration patterns Strong software engineering principles, including testing, maintainability and security Monitoring, observability and operational support for production data platforms Event-driven architectures, messaging and streaming technologies OpenSearch, graph databases or knowledge graph technologies Data discovery, catalogue ...

Senior Vice President, Full-Stack Engineer

Location
Bolton, England, United Kingdom
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise‐scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. To be successful in this role … platform design and optimisation. Proven track record of delivering production‐grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands‐on technical leader who can actively contribute to solution design and critical builds, while defining ...

Lead AI Engineer (Leeds / Edinburgh)

Hiring Organisation
Hays
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
solutions Enhance CI/CD pipelines for AI systems, introducing automation, traceability and controlled release processes to ensure safe and repeatable deployments. Define monitoring, observability and operational strategies, improving visibility across quality, performance, safety, cost and reliability to support effective issue diagnosis and resolution Set standards for prompt engineering, evaluation … using LLM APIs/Bedrock, including RAG architectures, vector databases and orchestration frameworks DevOps skills including CI/CD pipelines, containerisation and monitoring/observability, with experience defining release and operating practices for AI services Experience Leading Engineering teams: providing technical guidance, aligning on standards/patterns, and adapting plans ...

Lead AI Engineer (Leeds / Edinburgh)

Hiring Organisation
Hays
Location
Leeds, West Yorkshire, United Kingdom
Salary
£ 80 K
complex AI solutionsEnhance CI/CD pipelines for AI systems, introducing automation, traceability and controlled release processes to ensure safe and repeatable deployments.Define monitoring, observability and operational strategies, improving visibility across quality, performance, safety, cost and reliability to support effective issue diagnosis and resolutionSet standards for prompt engineering, evaluation … using LLM APIs/Bedrock, including RAG architectures, vector databases and orchestration frameworks DevOps skills including CI/CD pipelines, containerisation and monitoring/observability, with experience defining release and operating practices for AI services Experience Leading Engineering teams: providing technical guidance, aligning on standards/patterns, and adapting plans ...

Lead Integration and Cloud Solution Architect

Location
Bradford, England, United Kingdom
Kubernetes concepts API-led connectivity and reusable integration patterns Hybrid-cloud and multi-cloud architecture OAuth2, OpenID Connect, JWT and Mutual TLS Monitoring, observability and operational support models CI/CD pipelines and DevOps practices Significant experience as a Solution Architect, Integration Architect or Cloud Architect in enterprise-scale environments. … architecture. Strong Microsoft Azure and AWS architecture experience across hybrid-cloud and multi-cloud environments. Strong knowledge of API security, cloud networking, identity, resilience, observability, CI/CD and DevOps practices. Confident stakeholder management, influencing, communication and mentoring skills, with experience working within architecture governance frameworks. Useful, but not essential ...

Agentic AI Forward Deployed Engineer

Hiring Organisation
HCLTech
Location
Leeds, England, United Kingdom
harnesses and test suites that measure agent correctness, safety and regression before anything ships. Own AgentOps/DevSecOps: CI/CD for agents, versioning, observability and telemetry, shift-left security, and Responsible AI governance baked in from day one. Run a continuous, adaptable feedback loop: feed production telemetry, evals … Eval-driven development: designing evaluation harnesses and measuring agent quality, safety and reliability. Standards-based integration and DevSecOps: APIs, secure auth, CI/CD, observability and AgentOps. Ability to conceptualize a business problem as an agent quickly, and operate effectively in ambiguous, customer-embedded settings. Client-facing maturity: translates fluidly ...