276 to 300 of 338 Observability Jobs in the North of England

Senior Engineering Manager

Location
Leeds, England, United Kingdom
Job ContextEngineering • Leeds • Full-Time • HybridAre you a people-first engineering leader who thrives on turning complex technical challenges into elegant, user-centric software As our Senior Engineering Manager, you will sit at the intersection ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health ...

Vice President, Production Services Application Support

Location
Manchester, England, United Kingdom
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Senior Software Engineer

Hiring Organisation
Ask4.com
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
Salary
£60,000
Kubernetes configurations for production and non-production environments Integrate with network management systems, message brokers, and third-party APIs Instrument and monitor applications using observability tooling (metrics, logs, and traces Grafana, Prometheus, or similar) Provide technical mentorship to mid-level engineers and act as an escalation point for complex problems … Experience with message brokers and event-driven systems (NATS, RabbitMQ, Kafka, or similar) Exposure to OpenWiFi, OpenWrt or similar open-source network controller frameworks Observability experience. Use of metrics, logs, and traces using tools such as Grafana, Prometheus and Sentry Experience with AI agent development or LLM integration Agile/ ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Deep understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Strong analytical and problem-solving skills with the ability to identify root causes ...

Lead AI Engineer

Location
Manchester, England, United Kingdom
traffic, within real latency budgets and reliability realities. Hold a high engineering bar on AWS and TypeScript — clean CI/CD, infrastructure as code, observability, testing, and LLMOps for running model‐backed systems in production. Shape delivery against the roadmap with the VP, product managers, data analysts and Staff Engineers … powered applications. Extensive experience building and operating cloud‐native applications on AWS, with strong knowledge of CI/CD, infrastructure as code, observability and modern engineering practices. Experience delivering high‐performance, real‐time systems that operate reliably at scale with demanding latency requirements. Demonstrated experience leading the successful transition ...

Principal AI Engineer - Microsoft Azure AI Foundry (Contract)

Hiring Organisation
Hackajob Ltd
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
tooling. Define reusable architecture patterns for AI workloads, including development, testing and production environments. Establish platform standards covering resource structure, environments, deployment patterns, observability and operational management. Work closely with AI Engineers to design and deliver AI Agents and agentic workflows that meet business and technical requirements. Governance, Risk & Compliance … provisioning and configuration using Bicep, Terraform or ARM templates. Build automated processes for prompt-flow evaluation, model deployment, testing and release management. Establish platform observability covering availability, performance, usage, cost and AI workload health. Experience 5+ years' experience in Azure cloud architecture, engineering or platform engineering. 12+ years' experience specifically ...

Head of Digital Technology

Location
Manchester, England, United Kingdom
We’re Pret: proud makers of freshly made food, organic coffee, and big ideas. Across 750+ shops and 20+ countries, our teams are shaping the future of Pret through innovation, inclusion, great customer service and ...

Data Architect

Location
Manchester, England, United Kingdom
We believe in the power of ingenuity to build a positive human future.We challenge where it matters and own the outcome.As strategies, technologies, and innovation collide, we create opportunity from complexity. Our teams of interdisciplinary ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
engineered into every service throughout its lifecycle. Working alongside the Director of Site Reliability Operations, this leader will define the engineering standards, automation, observability, production readiness, and resilience capabilities that enable world‐class operational performance. While Site Reliability Operations owns the day‐to‐day operation of production services, the Site … global Site Reliability Engineering organization. This role owns the engineering strategy, governance, architecture, and technical practices that improve service reliability through software engineering, observability, automation, resilience engineering, production engineering, and operational readiness. Rather than operating production systems on a day‐to‐day basis, the Site Reliability Engineering organization develops ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 05/08/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 30/09/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

Site Reliability Engineer - Public cloud

Hiring Organisation
Intuition IT Solutions Ltd
Location
Leeds, Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP Daily
remain secure, resilient, performant and available. Working within a multidisciplinary engineering team, you'll apply Site Reliability Engineering principles to reduce operational toil, improve observability, automate repetitive activities and enhance platform reliability across the estate. The role aligns to LBG Grade D engineering expectations, combining hands-on technical delivery with … strong focus on operational excellence, continuous improvement and customer outcomes. WHAT YOU'LL BE DOING Deliver reliability improvements through automation, observability and engineering best practices. Support CI/CD pipeline development and deployment automation activities. Experience supporting enterprise database technologies including: o Oracle Database o IBM DB2 o SAP HANA ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
urgency, recover priority incidents under pressure, and maintain core support coverage across on-site and offshore support hours. Use SQL scripting, automation, monitoring, and observability tools to improve operational resilience, service health, reliability, and incident response. To be successful in this role, were seeking the following: Excellent SQL scripting skills. … solutions for alert correlation, anomaly detection, predictive monitoring, and service optimisation. Strong understanding of Site Reliability Engineering (SRE) principles, including service health, reliability, availability, observability, incident reduction, and continuous service improvement. Experience with SRE practices such as monitoring and alert tuning, incident management, post-incident reviews, root cause analysis ...

Senior Architect Private Cloud

Hiring Organisation
Randstad Technologies Recruitment
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £550/day Inside IR35 via Umbrella
patterns, and platform guardrails. Kubernetes Platform Leadership: Own the end-to-end Kubernetes ecosystem, including cluster topology, multi-tenancy, ingress, service mesh, secrets management, observability, and seamless workload onboarding. Event-Driven Architecture: Direct the adoption and governance of Apache Kafka, covering partitioning strategies, schema management, resilience patterns, capacity planning … configuration management. Desirable Skills: Familiarity with Service Mesh (e.g., Istio), API Gateways, Policy as Code (e.g., OPA), developer portals (Platform-as-a-Product), and Observability stacks (OpenTelemetry, Prometheus, ELK). Randstad Technologies is acting as an Employment Business in relation to this vacancy. ...

Frontend Engineer

Hiring Organisation
Ventula Consulting
Location
Manchester, Lancashire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 600 Daily
improve application performance, reliability, and scalability. Participate in code reviews and contribute to Front End architecture and engineering best practices. Monitor application health using observability and performance monitoring tools. Contribute to CI/CD pipelines and production deployments. Must Have 3+ years of professional Front End development experience. Strong knowledge … quality. Experience with Git and collaborative development workflows. Experience with CI/CD pipelines. Experience debugging and supporting production applications. Experience with monitoring and observability tools. Experience working with Docker and Kubernetes. Ability to quickly understand large existing codebases. Ability to work independently with minimal onboarding. Strong focus on delivery ...

Senior Vice President, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise-scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. To be successful in this role, were seeking … platform design and optimisation. Proven track record of delivering production-grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands-on technical leader who can actively contribute to solution design and critical builds, while defining ...

Agentic AI Forward Deployed Engineer

Hiring Organisation
HCLTech
Location
Leeds, UK
harnesses and test suites that measure agent correctness, safety and regression before anything ships. Own AgentOps/DevSecOps: CI/CD for agents, versioning, observability and telemetry, shift-left security, and Responsible AI governance baked in from day one. Run a continuous, adaptable feedback loop: feed production telemetry, evals … Eval-driven development: designing evaluation harnesses and measuring agent quality, safety and reliability. Standards-based integration and DevSecOps: APIs, secure auth, CI/CD, observability and AgentOps. Ability to conceptualize a business problem as an agent quickly, and operate effectively in ambiguous, customer-embedded settings. Client-facing maturity: translates fluidly ...

Head of Observability & Security Engineering

Location
Manchester, England, United Kingdom
Head of Observability & Security Engineering Employer: Co-op Group Location: Pay: Meets national minimum wage Contract Type: Permanent Hours: Full time Disability Confident: No Closing Date: 20/08/2026 About this job Head of Observability & Security Engineering Up to £100.000 plus private healthcare, company car or car allowance … working so we can deliver even better services for our Co-op, our colleagues, members and customers. We're looking for a Head of Observability and Security Engineering to join our Technology Operations and Assurance team and help us build and run the platforms that keep our systems secure, resilient ...

Senior Vice President, Full-Stack Engineer

Location
Manchester, England, United Kingdom
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise-scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. Qualifications Bachelor’s or Master’s degree … platform design and optimisation. Proven track record of delivering production-grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands-on technical leader who can actively contribute to solution design and critical builds, while defining ...

Data Scientist

Hiring Organisation
Hackajob Ltd
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
Financial Crime Riskdriving end-to-end, production-grade AI solutions from governance and senior decision-making through to LLM/agent workflow design, evaluation, observability, and hypercare, ensuring solutions are adopted in investigations and meet regulated-bank standards. As an HSBC employee in the UK, youll have access to tailored … data integration patterns (batch/stream), data quality controls -LLM/agent patterns: prompt/version management, tool/function calling, RAG, safety guardrails -Observability and production support (logging/metrics/tracing, incident triage) Opening up a world of opportunity. Being open to different points of view is important ...

Senior Site Reliability Engineer

Location
Knutsford, England, United Kingdom
ability to articulate these relationships through narrative, diagrams, and documentation. Keen interest in researching, evaluating and directly engaging with technologies to improve predictability, observability and performance, combined with enthusiasm for teaching others and lifelong learning. Some other highly valued skills include: A deep understanding of systems engineering, including operating systems … tools, and infrastructure‐as‐code. Interest and experience in innovative uses for artificial intelligence, solving technology problems faster and at greater scale. Experience with observability tools and techniques for instrumentation, gathering information and extending the boundaries of the known. You may be assessed on the key critical skills relevant ...

Lead Automation Engineer

Location
Leeds, England, United Kingdom
About Waystone Waystone is a leading asset-servicing solutions provider of institutional governance, administration, risk and compliance services to financial institutions. With over 25 years’ experience and a comprehensive range of specialist services to its ...