151 to 175 of 194 Observability Jobs in Manchester

Site Reliability Engineer (SRE)

Hiring Organisation
Spencer Rose Ltd
Location
Manchester, Lancashire, United Kingdom
Employment Type
Contract
Contract Rate
GBP Daily
operational excellence of cloud-hosted services on Google Cloud Platform. This is a hands-on engineering role spanning SRE practices, production Kubernetes, infrastructure automation, observability, CI/CD, incident response and continuous service improvement. About the role The Senior Site Reliability Engineer will work with Cloud Platform, Software Engineering, Product … shared services. Define and operate service level indicators, service level objectives and error-budget practices that connect technical health to customer impact. Design actionable observability using Dynatrace, including instrumentation, dashboards, distributed tracing, service health views and SLO-based alerting. Build modular, reusable and maintainable Terraform code for secure cloud infrastructure ...

Senior Specialist, Production Services Application Support Analyst

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Analytical and problem-solving skills with the ability to identify root causes and implement sustainable ...

Vice President, DevOps Production Services

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
enterprise applications and ensure platform stability, resiliency, and availability. Monitor application health, system performance, batch jobs, interfaces, and alerts using enterprise monitoring and observability tools. Investigate, troubleshoot, and resolve production incidents within defined SLAs. Perform root cause analysis (RCA) for recurring issues and drive permanent fixes. Analyze production logs, identify … Cloud experience preferred. Knowledge of automation/scripting using Python, Shell, or PowerShell. Exposure to DevOps/SRE practices, CI/CD pipelines, and observability tooling. Strong communication skills with the ability to provide concise incident and executive status updates. Our culture allows us to run our company better ...

Lead Software Engineer

Location
Manchester, England, United Kingdom
high standards through code reviews, mentoring engineers, and contributing to CI/CD and testing best practices Applying site reliability engineering principles to improve observability, resilience, and incident response Monitoring, profiling, and optimising application performance, including capacity planning to meet SLAs Identifying technical debt and recommending improvements to systems, processes … containerisation Strong debugging, performance optimisation, and problem-solving skills across distributed systems Experience mentoring engineers and contributing to technical leadership within teams Knowledge of observability, monitoring, and site reliability engineering practices Experience with Java and/or telecoms technologies (e.g. VoIP, WebRTC) is highly beneficial What do we offer ...

Agentic Platform Engineer · Manchester, UK ·

Location
Manchester, England, United Kingdom
Protocol (MCP). Build production-grade agent services using Python, cloud-native architectures, event-driven design, automation and Infrastructure as Code. Implement robust evaluation, observability and continuous improvement capabilities, including testing, tracing, telemetry and performance optimisation. Embed security, governance and responsible AI principles through least-privilege access, policy enforcement, auditability … workflow state, retrieval-augmented generation and human-in-the-loop patterns. Experience building secure integrations with enterprise APIs, repositories, cloud services, CI/CD, observability or ITSM platforms. Experience creating evaluation frameworks for agent quality, task completion, safety, reliability, latency and cost. Strong understanding of agent security, including workload identity ...

Engineering Manager (Remote - UK)

Hiring Organisation
Reonomy
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
building and evolving AWS-native and cloud-based platforms, alongside legacy systemsSet and uphold strong engineering standards across code quality, testing, CI/CD, observability and documentationStay close to technical decisions through design reviews, architecture discussions and hands-on coachingBalance new feature delivery with technical debt, reliability, security and long … distributed systemsProficiency in at least one modern programming languageStrong grasp of system design and software engineering fundamentalsExperience with Infrastructure as Code, CI/CD, observability and secure production systemsAble to communicate technical ideas clearly to both technical and non-technical stakeholdersUnlock your Altus Experience!If you’re looking to advance ...

Senior Engineering Manager - 9-10 month FTC

Location
Manchester, England, United Kingdom
practices, including test‐driven development and automated testing, helping teams build quality into the development process. Support strong operational ownership through CI/CD, observability, production support and the principle that teams own the systems they build and run. Help teams identify and address technical debt sustainably while maintaining appropriate … outcomes and translating these into clear engineering priorities. Strong understanding of modern software engineering practices, including automated testing and TDD principles, CI/CD, observability and production ownership. Experience facilitating technical decisions, bringing the right people and evidence together and constructively challenging thinking when needed. Comfortable balancing feature delivery ...

Data Architect

Hiring Organisation
PA Consulting
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
Company DescriptionWe believe in the power of ingenuity to build a positive human future. We challenge where it matters and own the outcome. As strategies, technologies, and innovation collide, we create opportunity from complexity. Our ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health ...

Vice President, Production Services Application Support

Location
Manchester, England, United Kingdom
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Deep understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Strong analytical and problem-solving skills with the ability to identify root causes ...

Lead AI Engineer

Location
Manchester, England, United Kingdom
traffic, within real latency budgets and reliability realities. Hold a high engineering bar on AWS and TypeScript — clean CI/CD, infrastructure as code, observability, testing, and LLMOps for running model‐backed systems in production. Shape delivery against the roadmap with the VP, product managers, data analysts and Staff Engineers … powered applications. Extensive experience building and operating cloud‐native applications on AWS, with strong knowledge of CI/CD, infrastructure as code, observability and modern engineering practices. Experience delivering high‐performance, real‐time systems that operate reliably at scale with demanding latency requirements. Demonstrated experience leading the successful transition ...

Head of Digital Technology

Location
Manchester, England, United Kingdom
We’re Pret: proud makers of freshly made food, organic coffee, and big ideas. Across 750+ shops and 20+ countries, our teams are shaping the future of Pret through innovation, inclusion, great customer service and ...

Data Architect

Location
Manchester, England, United Kingdom
We believe in the power of ingenuity to build a positive human future.We challenge where it matters and own the outcome.As strategies, technologies, and innovation collide, we create opportunity from complexity. Our teams of interdisciplinary ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
engineered into every service throughout its lifecycle. Working alongside the Director of Site Reliability Operations, this leader will define the engineering standards, automation, observability, production readiness, and resilience capabilities that enable world‐class operational performance. While Site Reliability Operations owns the day‐to‐day operation of production services, the Site … global Site Reliability Engineering organization. This role owns the engineering strategy, governance, architecture, and technical practices that improve service reliability through software engineering, observability, automation, resilience engineering, production engineering, and operational readiness. Rather than operating production systems on a day‐to‐day basis, the Site Reliability Engineering organization develops ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 05/08/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 30/09/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
urgency, recover priority incidents under pressure, and maintain core support coverage across on-site and offshore support hours. Use SQL scripting, automation, monitoring, and observability tools to improve operational resilience, service health, reliability, and incident response. To be successful in this role, were seeking the following: Excellent SQL scripting skills. … solutions for alert correlation, anomaly detection, predictive monitoring, and service optimisation. Strong understanding of Site Reliability Engineering (SRE) principles, including service health, reliability, availability, observability, incident reduction, and continuous service improvement. Experience with SRE practices such as monitoring and alert tuning, incident management, post-incident reviews, root cause analysis ...

DevOps and Machine Learning Operations Engineer

Location
Manchester, England, United Kingdom
runtime platform the project depends on, as well as the model-serving path. The role covers infrastructure as code, continuous integration and deployment, observability, cost control, and production support. Applicants must have experience running systems in production and being accountable for their reliability. Main Responsibilities Infrastructure and Environments: Define … gates that fail closed. Automate database migration and rollback for reversible releases. Support progressive delivery, including staged rollout and fast rollback, with deployment tracking. Observability and Operations: Instrument services with structured logging, metrics, and tracing, and define user-focused alerts. Establish service level objectives and report against them. Run incident ...

Senior Vice President, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise-scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. To be successful in this role, were seeking … platform design and optimisation. Proven track record of delivering production-grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands-on technical leader who can actively contribute to solution design and critical builds, while defining ...

Head of Data Engineering £98,000 OTE Manchester Data

Location
Manchester, England, United Kingdom
Location: Manchesterbased office – Hybrid working PHMG is continuing to develop and mature its data capability, focusing on modernising our data platform and enhancing the way we work. This creates further opportunities toleveragedata more effectively, allowing ...

Principal/Lead Consultant - Cloud & Engineering

Hiring Organisation
Zühlke Engineering
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
Founded in Switzerland in 1968, Zühlke is owned by its partners and located across Europe and Asia. We are a global transformation partner, with engineering and innovation in our DNA. We're trusted to help ...

Head of Observability & Security Engineering

Location
Manchester, England, United Kingdom
Head of Observability & Security Engineering Employer: Co-op Group Location: Pay: Meets national minimum wage Contract Type: Permanent Hours: Full time Disability Confident: No Closing Date: 20/08/2026 About this job Head of Observability & Security Engineering Up to £100.000 plus private healthcare, company car or car allowance … working so we can deliver even better services for our Co-op, our colleagues, members and customers. We're looking for a Head of Observability and Security Engineering to join our Technology Operations and Assurance team and help us build and run the platforms that keep our systems secure, resilient ...