51 to 75 of 239 Observability Jobs in the North West

ML/AI Engineer

Location
Manchester, England, United Kingdom
leverage TensorRT where appropriate. Operate scalable serving frameworks (NVIDIA Triton, TorchServe) with attention to latency, efficiency, resilience, and cost. Implement end‐to‐end observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace. Establish actionable alerting ...

Lead DevOps Engineer

Hiring Organisation
Pathfinder Business Solutions Ltd
Location
Chester, Cheshire, North West, United Kingdom
Employment Type
Permanent
Salary
£80,000
teams on resilient, cloud-native services. Youll lead improvements across automation, GitOps, infrastructure as code and CI/CD, reducing manual effort while improving observability, SLOs, security, resilience and reliability. Youll also act as a senior escalation point for complex production issues and participate in the shared on-call rota. ...

Lead DevOps Engineer

Hiring Organisation
Shortlist Recruitment
Location
Chester, Cheshire, UK
Employment Type
Full-time
deployment solutions using IaC and GitOps practicesBuild and maintain CI/CD pipelines to enable development teams to deploy applications quickly and reliablyImprove monitoring, observability and incident response to enhance system reliability and resilienceContribute to the ongoing replatforming of services towards AWS and Kubernetes, identifying opportunities to automate, improve performance ...

Head of Engineering

Hiring Organisation
WRK DIGITAL LTD
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
platform engineering and cloud services. Partner with Product teams to align engineering delivery with strategic objectives and business outcomes. Champion engineering excellence through automation, observability and data-driven decision making. Drive improvements across deployment frequency, lead time, reliability and engineering performance metrics. Lead adoption of modern architectures including cloud-native ...

Senior AI Engineer

Location
Manchester, England, United Kingdom
against live traffic and within real latency budgets. Hold a high engineering bar on AWS and TypeScript: clean CI/CD, infrastructure as code, observability, and testing. Work closely with R&D engineers, product managers, data analysts and Staff Engineers, bridging fast experimental work and production discipline. And what will ...

Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Knutsford, Cheshire, North West, United Kingdom
Employment Type
Permanent
Practical experience with microservices and API design, with a clear understanding of service boundaries, integration contracts, and non-functional requirements such as resilience, scalability, observability, and failure handling. Experience working with event-streaming or messaging platforms (e.g. Kafka or equivalent), including concepts such as topics, partitions, schemas, consumer groups ...

Senior Software Engineer (Vue.js/TypeScript) - iGaming

Location
Manchester, England, United Kingdom
tools such as Jenkins or GitLab CI Performance optimisation (bundle size, lazy loading, rendering performance) Cloud platforms and static hosting (AWS, Azure or GCP) Observability tooling including logging, metrics and frontend monitoring Package management with npm, yarn or pnpm What Success Looks Like You’ll demonstrate: Technical Excellence & Craft Mastery ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud … Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner ...

Senior Backend Developer

Hiring Organisation
Protein Works
Location
Liverpool, Merseyside, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
connect e-commerce, ERP, WMS, finance, and manufacturing systems. Operations, Security & AI: Own end-to-end delivery including containers, CI/CD, IaC, full observability (metrics, traces, logs), security/privacy compliance (OWASP, GDPR, pen testing), and day-to-day use of AI-assisted/agentic tools. Cross-Functional Leadership … microservices and monoliths. Data & Async Systems: Solid background in relational databases (PostgreSQL, MySQL, SQL Server) and asynchronous messaging (queues, events, retries, idempotency). DevOps & Observability: Practical experience with Docker, Linux/Windows, Git/GitHub, Jira/Agile, monitoring/alerting, and diagnostic/profiling tools in high-volume ...

AI-Driven Cloud & Platform Engineer for Digital Factory

Location
Manchester, England, United Kingdom
implement CI/CD automation, infrastructure-as-code, and self-service tooling, leveraging Docker, Kubernetes, and multiple cloud providers. You’ll define observability and reliability practices, collaborate with architects, and mentor teammates. You should bring hands-on experience in cloud, DevOps, and SRE disciplines, with a willingness to work across ...

Lead Platform Engineer

Location
Manchester, England, United Kingdom
using Terraform Supporting container platforms using Kubernetes and Docker Developing CI/CD pipelines to automate delivery and deployment Improving platform monitoring, reliability and observability Providing technical leadership and mentoring engineers across delivery teams LEAD PLATFORM ENGINEER ESSENTIAL SKILLS Infrastructure as Code experience using Terraform Container technologies such as Kubernetes ...

Senior DevOps Engineer

Location
Manchester, England, United Kingdom
troubleshooting across containers, networking, databases, and more, identifying root causes and improving platform reliability. Improve monitoring, logging, alerting and incident response processes, introducing proactive observability, operational runbooks and stronger disaster recovery procedures. Drive infrastructure automation through Infrastructure-as-Code and CI/CD, reducing manual changes and improving deployment consistency ...

Head of Cyber, Platforms & IT

Location
Chester, England, United Kingdom
enforce secure‐by‐design principles, including cybersecurity standards, cloud architecture guardrails and operational controls Lead DevOps and Site Reliability Engineering (SRE) maturity, embedding monitoring, observability, automated testing and structured incident response Drive adoption of automation and AI‐enabled tooling to improve anomaly detection, incident management, vulnerability management and operational efficiency ...

Back End Engineer

Hiring Organisation
Pontoon
Location
Manchester, United Kingdom
Employment Type
Contract
Conversant with GCP Implement OAuth, SAML and SSO Build cloud-native applications using Kubernetes/OpenShift and containers Develop CI/CD, monitoring and observability capabilities Work closely with AI, Data, Architecture and Security teams What we're looking for Strong backend engineering experience with Java, Python, Node.js or similar ...

Platform Engineer

Location
Stockport, England, United Kingdom
technical problems Comfortable operating in a senior engineering environment and making technical recommendations Highly Desirable Helm AWS experience CI/CD tooling Monitoring and observability Cloud networking and VPC design Infrastructure architecture Governance and security best practices Experience simplifying or modernising complex infrastructure environments The Ideal Candidate You will likely ...

IT Infrastructure Engineer

Location
Manchester, England, United Kingdom
build, integration, and deployment of products across all environments, including the use of fragments/includes for pipeline efficiency. Proven ability to utilize observability tools (logs, metrics, and tracing) and automated testing to monitor system health, troubleshoot issues, and confirm the ongoing viability of products in production. Experience of working ...

Databricks Architect - Lead/Principal Consultant

Location
Greater Manchester, England, United Kingdom
Engineering Fundamentals: Bring rigour to how we model, test, deploy and operate data platforms: dimensional and medallion modelling, CI/CD, infrastructure as code, observability, cost management and data quality by design. 1. Architectural Vision & Technical Ownership Lead from the Front: Build and lead high-performing delivery teams, setting direction ...

Principal Developer - AI

Location
Salford, England, United Kingdom
lifecycle management on Kubernetes across AWS, GCP, and Azure. Establish and evangelize engineering best practices across services, including standards for scalability, fault tolerance, observability, and operational excellence. Mentor senior engineers and partner closely with data scientists to elevate technical quality and shorten the path from research to production. Partner with ...

Associate Director, Data Science/Gen AI Lead - ER&I

Hiring Organisation
Deloitte
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
/GenAI governance & ethics (bias detection, explainability). GenAI Platform & Infrastructure Architecture (Cloud, Lakehouse). GenAI ModelOps & Performance Monitoring. AI-driven business intelligence & reporting. Observability & FinOps for AI/GenAI. Cloud Infrastructure, Networking, & Security for AI.Aligning GenAI Architectures Across Organizations: Experience aligning GenAI architecture blueprints across business units and geographies ...

Senior Platform/Dev Ops Engineer

Location
Manchester, England, United Kingdom
platform reliability across multi-cloud environments. The ideal candidate has strong experience building enterprise-scale CI/CD platforms, Kubernetes ecosystems, Infrastructure as Code, observability solutions, and developer enablement capabilities. Key Responsibilities Architect, build, and evolve scalable platform engineering and DevOps solutions for enterprise engineering teams. Lead the design … manage scalable Infrastructure as Code frameworks using Terraform and Terraform Enterprise. Lead Kubernetes platform architecture, Helm-based deployments, and container orchestration best practices. Establish observability standards and operational excellence using Grafana, Loki, Prometheus, OpenTelemetry, and related tooling. Collaborate with engineering, security, cloud, and architecture teams to improve platform capabilities ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Manchester, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Manchester, England, United Kingdom
national security. Join agile, multi-disciplinary teams focused on CI/CD, Infrastructure as Code and live service support, with emphasis on security, observability and incident readiness to ensure #J-18808-Ljbffr ...

MongoDB Site Reliability Engineer

Location
Knutsford, England, United Kingdom
paced environment, your role will be essential to ensuring our infrastructure remains resilient, secure, and scalable. You’ll work on automating operations, enhancing system observability, and driving continuous improvements that reduce downtime and improve efficiency. If you’re motivated by solving, multi-layered problems and building systems that perform reliably ...

Senior DevOps Engineer

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£50,000
technical leadership and subject matter expertise across Microsoft Azure * Build, maintain and optimise Infrastructure as Code solutions and CI/CD pipelines * Improve monitoring, observability and operational performance across digital services * Ensure the security, stability and availability of production and non-production environments Essential skills for the Senior DevOps Engineer ...

Director, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
peer review, automated testing, release governance, production validation, and support readiness. Champion engineering best practices in full stack architecture, reusable design patterns, API strategy, observability, and secure coding. Strengthen platform stability and operational resilience by reducing manual touchpoints, improving exception handling, and enhancing recovery and support processes. Support release planning ...