226 to 250 of 338 Observability Jobs in the North of England

Senior AI Engineer - Real-Time Personalization & AI Systems

Location
Manchester, England, United Kingdom
players. As a senior engineer, you’ll own the design and delivery of real-time AI systems, focusing on low latency, reliability, and observability, while collaborating with product managers and data analysts. #J-18808-Ljbffr ...

Operations Team Lead: Scale Reliability & Incident Mastery

Location
Stockport, England, United Kingdom
move from reactive firefighting to proactive reliability engineering. You will lead the Operations team, set standards, manage incidents, on-call rotations, and drive observability improvements. A strong SRE/DevOps background and calm, clear communication are essential. #J-18808-Ljbffr ...

Database Reliability Engineer - Hybrid, Cross-Cloud RDS

Location
Manchester, England, United Kingdom
data-focused engineering role within our hybrid data environment. You will help modernize and scale the RDS fleet, architect cross-cloud portability, and advance observability across a global database fleet. Strong Postgres, Kubernetes, and cloud-native experience are essential. Join a fast-moving fintech team with a strong focus ...

Lead Software Engineer (£80k + benefits)

Location
Wigan, England, United Kingdom
production. As a Lead Software Engineer you’d play a key role in system design, helping to modernise their existing microservices and improve observability and testability by using modern approaches like hexagonal architecture and agentic behaviour driven design. Skills: The money is good too – up to £80k plus benefits including ...

Senior & mid-level Data Platform Engineers

Hiring Organisation
ITSS Recruitment
Location
Manchester, United Kingdom
Employment Type
Permanent
Salary
£70000 - £100000/annum Bonus + Excellent benefits
consistency and reusability across environments. * Build and optimise CI/CD pipelines using Azure DevOps and GitHub Actions to support rapid, reliable deployments. * Implement observability practices including logging, metrics, and alerting using observability tools. * Collaborate with the Lead Engineer and Architects to align implementation with platform standards and patterns. * Provide … Fabric. * Proven experience with infrastructure-as-code using Terraform and building CI/CD pipelines via Azure DevOps and GitHub Actions. * Strong grasp of observability practices, including logging, metrics, alerting, and performance optimisation. * Deep understanding of cloud security, with experience applying secure-by-design principles in Azure and/ ...

Production Support Analyst

Location
Manchester, England, United Kingdom
thousands of clients and colleagues worldwide. This is more than a traditional production support role. As our team evolves towards a proactive resilience and observability model, you'll have the opportunity to help drive that transformation by identifying risks before they become incidents, improving operational excellence, and influencing ...

Principal Software Engineer - Full Stack - AI

Location
York and North Yorkshire, England, United Kingdom
full stack, guiding the development of responsive frontend applications (React/TypeScript) and robust, scalable backend services (Python, Java, Kotlin, or Node.js). LLM Observability & Reliability: Establish robust LLM observability, evaluations, and caching, implementing latency optimisations and comprehensive monitoring (logging, usage tracking, agent behaviour). Operational Excellence: Champion high availability … performance optimisation, and observability across frontends and backend microservices, focusing on practices that maintain platform reliability and optimise MTTD and MTTR. Mentorship & Collaboration: Elevate the engineering organisation by mentoring senior and junior engineers, conducting rigorous code and system design reviews, and partnering with product managers to translate product visions into ...

SC Cleared Tester (Performance)

Hiring Organisation
VIQU IT
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£450 - £475/day Inside IR35
role of Performance Tester you will enhance and shape the scalability and reliability - creating strategies and solutions in cloud and container based environments, leveraging observability tools and diagnostics to optimise application performance. Essential Criteria • Ability to design, execute, and analyse performance, load, stress, and volume tests. • Familiarity with monitoring … observability tools (Grafana, Splunk) and diagnostics platforms (New Relic, Dynatrace). • Understanding of containerisation technologies (Docker, Kubernetes) and big data platforms (Databricks). • Ability to analyse complex performance issues and provide actionable recommendations. • Strong scripting skills (e.g., Python, JavaScript) for test automation and data analysis. • Knowledge of Performance Test Strategy ...

Senior .NET Backend Developer

Location
York and North Yorkshire, England, United Kingdom
design, implementation, testing, review, refactoring, documentation and migration, while retaining clear ownership of the outcome. Set a high standard for maintainability, automated testing, security, observability and pull-request review. Investigate complex technical and production issues and drive them through to resolution. What we are looking for Strong experience designing … clinically important data. PostgreSQL, Redis, Elasticsearch or other data and caching technologies. GraphQL, including schema design and gateway patterns. Grafana, OpenTelemetry, Prometheus or equivalent observability tooling. CI/CD, production services, microservices and message-driven systems such as RabbitMQ. How we work We value engineers who own outcomes and remain ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Salford, Greater Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Cloud Engineer

Location
Manchester, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Lead Cloud Engineer

Location
Leeds, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

AI Platform & Site Reliability Engineering Managing Consultant

Location
Manchester, England, United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production‐grade, enterprise‐scale AI services. You will work with technology, engineering, operations … operational requirements. AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. Reliability Engineering & SRE: Establish SRE practices including ...

Lead, Production Reliability & Operations

Location
South Shields, England, United Kingdom
Operations Team Lead to own production and build scalable systems. You will lead operational excellence across all live customer-facing platforms, aiming for reliability, observability, and continuous improvement. This hands-on role requires shaping processes, incident management, and building a high-performing team. You will drive a culture of ownership ...

SRE Lead: Scalable, Reliable Betting Platform (Hybrid)

Location
Leeds, England, United Kingdom
evoke is seeking a Site Reliability Engineer to join our betting and gaming platforms, focusing on reliability, scalability and performance. You will work with observability, automation and engineering to deliver robust customer experiences across services. You will lead incident response, perform postmortems and drive resilience improvements. The role requires hands ...

Senior Ruby Backend Architect for E-commerce

Location
Manchester, England, United Kingdom
party integrations, and ensuring data parity across systems. You'll lead Ruby-focused architectures, mentor engineers, and drive best practices in testing, reliability, and observability while collaborating with product, design, and data teams to deliver exceptional user experiences. #J-18808-Ljbffr ...

Full-Stack Software Engineer — Hybrid, Health Benefits

Location
Sheffield, England, United Kingdom
work with a collaborative group of engineers, define scope with the Technical Lead, and contribute to design decisions while delivering reliable, observability-driven software through frequent, #J-18808-Ljbffr ...

IOS Engineer

Hiring Organisation
Randstad Technologies Recruitment
Location
Manchester, United Kingdom
Employment Type
Contract
Contract Rate
£43 - £84/hour Negotiable
work with backend-driven UI architectures. Write clean, testable code, maintain unit tests, and debug production issues. Monitor app stability using crash reporting and observability tools. Key Requirements Good years of professional iOS development experience. Strong proficiency in Swift, SwiftUI, and UIKit interoperability . Hands-on experience with modern architecture ...

iOS Engineer - £600 per day

Hiring Organisation
Ventula Consulting
Location
Manchester, Lancashire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 600 Daily
issues and improve application stability and performance. Participate in code reviews and contribute to engineering best practices. Monitor application quality through crash reporting and observability tools. Be familiar with Back End-driven UI solutions Must Have 3+ years of professional iOS development experience. Strong knowledge of Swift. Experience with SwiftUI ...

Founding Engineer

Location
Manchester, England, United Kingdom
each customer The systems that keep long-running agents dependable in production, including durable memory, scheduled work, sandbox lifecycle, tool execution, recovery and observability The trust layer between an agent and a founder, so work moves from draft to delivered with the right approvals, clear status and safe ways ...

3rd Line Network Engineer

Location
Salford, England, United Kingdom
enterprise networking platforms such as Palo Alto, Cisco and Juniper across complex customer or multi-site environments. Rich working experience with network monitoring and observability tools such as LogicMonitor, SolarWinds and NetFlow, including alert tuning, threshold review, root cause investigation and proactive service improvement. Strong documentation and network diagramming skills ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Manchester, England, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Newcastle upon Tyne, England, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

Site Reliability Engineer

Location
Manchester, Lancashire, United Kingdom
reliability and performance. Automating repetitive operational tasks and reducing manual intervention wherever possible. Monitoring and troubleshooting systems across the application and infrastructure stack. Improving observability and instrumentation to identify issues and measure system performance. Working alongside development and product teams to build scalable and resilient services. Responding to production incidents … including Bash or PowerShell. Cloud platforms such as AWS, Azure or OpenStack. Infrastructure automation and configuration management. CI/CD and deployment tooling. Monitoring, observability and troubleshooting of production systems. Docker, containers and/or microservices. Diagnosing issues across different levels of the technology stack. Working within Agile engineering teams. ...