1,226 to 1,250 of 1,791 Permanent Observability Jobs

Site Reliability Engineer (DV Security Clearance)

Hiring Organisation
CGI
Location
Manchester, United Kingdom
Employment Type
Full Time
Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production environments. The successful candidate will work closely with engineering, platform … services across cloud and containerised environments. - Manage and support Kubernetes clusters and Helm-based deployments across multiple environments. - Enhance monitoring, alerting, logging, and observability solutions to improve operational visibility and system reliability. - Investigate incidents, analyse logs, identify root causes, and drive timely resolution of production issues. - Participate in incident response ...

Platform Engineer

Hiring Organisation
CGI
Location
Newry Mourne and Down, United Kingdom
Employment Type
Full Time
take ownership of designing and evolving resilient digital platforms that accelerate delivery and power mission-critical services. Working across cloud infrastructure, automation, security, and observability, you will help set engineering standards and drive measurable improvements in reliability, performance, and cost efficiency. Within a collaborative, multidisciplinary environment, you will have … Code standards and reusable modules Build & Optimise CI/CD pipelines for efficient, secure delivery Enable & Support containerised workloads and orchestration platforms Implement & Improve observability, logging, metrics, and alerting standards Secure & Govern platforms in line with enterprise security frameworks Troubleshoot & Resolve critical incidents, strengthening resilience and diagnostics Optimise & Control cloud ...

Technical Release Manager

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
functional teams throughout the development lifecycle, maintaining the release calendar, championing engineering and DevOps best practice, and continuously raising the bar on deployment quality, observability and release velocity across both AWS and Azure. A little about you... Proven experience managing software releases within a SaaS environment, combined with strong project … services and configuration management. Familiarity with containers and orchestration and modern deployment patterns to support zero‐down‐time releases. Understanding of monitoring, logging and observability tooling and their role in validating release health. Awareness of change management, release governance and security/compliance considerations within a SaaS delivery context. Proactive ...

Head of Infrastructure & Security

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automation strategy, partnering with Product and Engineering to design delivery pipelines that drive growth and feature velocity. Measure system performance using clear KPIs and observability dashboards. Engineering Enablement & Reliability (approx. 15%): Partner with tech leaders to streamline deployment processes across the development lifecycle. Drive incident response planning and Service Level …/GCP ecosystems and Kubernetes orchestration. You must be an expert systems administrator with hands‐on experience in Infrastructure as Code (Terraform) and modern observability tools (Datadog, Prometheus). Problem‐Solving: You are a data‐driven strategist who can devise high‐level strategy and is equally comfortable rolling up your ...

DevOps Engineer

Hiring Organisation
Big Red Recruitment Midlands Limited
Location
Coalville, Stanton under Bardon, Leicestershire, United Kingdom
Employment Type
Permanent
Salary
£59999 - £65000/annum £60,000 - £65,000
/CD pipelines to enable faster, safer software delivery. Embedding DevSecOps principles and security tooling throughout the development lifecycle. Improving platform resilience, monitoring, observability and operational performance. Driving automation and adopting AI-assisted engineering to improve efficiency. Working closely with software engineers, infrastructure specialists and support teams to continuously improve … advantageous) CI/CD pipelines and deployment automation Docker, containers and Linux DevSecOps, SAST/DAST and cloud security best practices Monitoring, logging and observability platforms Agile software delivery environments Why join? Build technology that has a genuine positive impact on society. Join a business investing heavily in cloud, automation ...

Systems Engineer, Production

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
scalable solutions to thousands of customers every day. Our mission is to make deploying and operating software effortless and safe. We focus on automation, observability, and reliability, ensuring every engineering team at the company can move faster and with confidence. You’ll be part of a globally distributed team, collaborating … Code using Terraform, ensuring reproducibility and compliance. Collaborate with developers to improve CI/CD pipelines, deployment strategies, and overall developer experience. Enhance observability and reliability, refining alerting, monitoring, and incident response. Collaborate on cloud optimization projects, improving performance, cost efficiency, and security posture. Mentor and guide team members, fostering ...

Senior Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
health: tech debt, refactoring, security, and governance Take ownership of features from start to finish using agile methodologies, delivering with feature flags, tests, and observability Champion high‐quality technical communications: proposals, specs, testing reports, and release planning Drive AI‐first ways of working within the squad — embedding AI tooling into … relates to the Manage domain Contribute to CI/CD pipeline improvements and progressive delivery practices across squads Drive reliability monitoring and observability within Manage (Prometheus, Grafana, Sentry) Contribute to security posture improvements: vulnerability scanning, pen testing coordination, and enforcement of standards Cross‐Squad & Leadership Collaboration Work closely with ...

DV Cleared DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Malvern, England, United Kingdom
environment provisioning, configuration management, and deployment processes. Collaborate with cross-disciplinary teams to ensure system reliability and security. Support incident investigations and contribute to observability strategies involving metrics, logs, and tracing. Lead or contribute to architectural discussions around system scalability, security, and best practices. Essential Skills & Qualifications Proven expertise with … including IaC, environment parity, and automation strategies. Proficiency in scripting languages such as Python for automation and tooling. Demonstrable experience with system monitoring and observability platforms. Solid understanding of security automation practices and experience working with modern infrastructure patterns. Active SC clearance (Security Check) is mandatory for this role. Excellent ...

AWS API Engineer / Data Architect

Hiring Organisation
Capgemini
Location
City and Borough of Birmingham, United Kingdom
Employment Type
Full Time
will design RESTful APIs and service boundaries, define integration standards, and collaborate with product, delivery and cybersecurity stakeholders to ensure secure access patterns, strong observability, and smooth deployments and cutovers. Where required, you will also contribute to data architecture decisions (e.g., PostgreSQL and document stores) to enable robust … will be able to design secure, scalable and well-governed RESTful APIs and microservices on AWS, applying consistent standards for identity, access control, observability, auditability and operational resilience. You will be comfortable collaborating with stakeholders to translate outcomes into pragmatic architectures and delivery plans, and you will bring strong engineering ...

Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
/CD pipelines in BitBucket Pipelines to test and deploy infrastructure changes automatically and safely. Apply SRE principles by codifying reliability, self‐healing, and observability directly into automated solutions. Set coding and automation standards for the team and mentor colleagues while remaining hands‐on; identify manual, repetitive tasks and replace … automate platform migration and modernisation work across Azure, AWS, VMware, Citrix, storage services, and Office 365. Instrument systems and build automated monitoring and observability using tools such as Splunk, Grafana, and Opsgenie. Participate in on‐call rotations and incident response, automating detection and remediation to minimise downtime. Embed security ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

AI Consulting –Sr. AI Architect & Client Partner

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
cost‐efficient infrastructure that empowers engineering teams to deliver quickly without sacrificing reliability or compliance. You will be the driving force behind our DevOps, observability, and compliance readiness , ensuring our systems are audit‐ready, highly available, and optimized for both performance and cost. Key Responsibilities Architect, implement, and maintain cloud … incredible journey and learning a lot along the way. Requirements Technical stack : Azure (also AWS is a plus), Terraform, AKS (Kubernetes), Docker, GitHub Actions. Observability : Experience implementing logging, metrics, and tracing frameworks. Security : Familiarity with best practices, secrets management, and security scanning tools. Networking : Solid understanding of VPCs, private networking ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
17918
Location
United Kingdom
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Principal C# Engineer / Snr Dev Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
.NET practices, SOLID principles and appropriate design patterns. Produce and review solution designs covering service boundaries, contracts, data models, integration patterns, idempotency, resilience, security, observability and operational supportability. Contribute directly to code delivery while reviewing implementation approaches, pull requests and engineering standards across the team. Lead delivery streams across internal … supportability, auditability, security and controlled change are important. Experience applying engineering best practice including loose coupling, contract‐first design, idempotency, automated testing, secure coding, observability and operational readiness. Experience leading people to deliver, including mentoring, work allocation, delivery tracking, quality review and constructive challenge, while remaining substantially hands‐on. Experience ...

Platform Chapter Lead - Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Lead Java Developer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...