676 to 700 of 4,640 Permanent Observability Jobs

Lead Ai Engineer (Contract)

Location
Newcastle upon Tyne, England, United Kingdom
Responsible AI principles (ATRS alignment, NCSC "Secure by Design, " meaningful human control) are enforced in the running system, not just on paper. Reliability, Security & Observability Build AI services to be secure-by-default and observable in production: logging, monitoring, alerting, and rollback paths appropriate for sensitive public-sector data ...

Software Engineer

Location
Manchester, England, United Kingdom
understanding of what you're building and why. Champion quality and operational excellence Apply DevOps practices and contribute to CI/CD, automation, and observability as a natural part of how you work. Write and maintain tests (unit, integration) and apply security best practices in day‐to‐day development. Participate ...

Senior SRE

Location
Horsell, England, United Kingdom
Istio). Skilled in designing and operating within a complex multi-cloud/hybrid ecosystem, with a solid understanding of distributed systems. Proficient in observability stacks such as Prometheus, Grafana, Loki, and Tempo. Hands on experience with EDB Postgres for enterprise-grade database solutions. Ability to develop and maintain Infrastructure ...

Cloud Security Engineer Gloucester- National Security West

Hiring Organisation
Hackajob Ltd
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
cutting-edge cloud activities focused on Security Operations. Experience using Infrastructure as Code, for example Terraform, CloudFormation or AWS CDK. Experience with monitoring and observability tools such as CloudWatch, Grafana or Prometheus. Experience with container orchestration platforms such as Kubernetes, Amazon ECS or OpenShift, including container security best practices. Understanding ...

Azure Platform Engineer

Location
Greater London, England, United Kingdom
integrated with Azure platform services. Establish curated Azure Deployment Environments and self-service infrastructure blueprints for Development, Applications and wider IT teams. Configure monitoring, observability, log aggregation and operational metrics using Azure Monitor, Log Analytics and Application Insights. Work with Infrastructure, Applications and Security Architects to ensure systems follow Business ...

Consulting Engineer, Infrastructure & AI Architecture

Location
City of Edinburgh, Scotland, United Kingdom
Cloud Architect, TOGAF, or equivalent).. Familiarity with private and hybrid cloud platforms (e.g., VMware, OpenShift, Nutanix, Azure Stack, AWS Outposts). Expertise with observability and FinOps tooling across hybrid estates (e.g., Prometheus, Grafana, ELK/OpenSearch, Datadog, cloud cost management platforms). Experience architecting under data sovereignty, regulated industry ...

Platform Engineering Manager (Cloud Foundations) London, UK

Location
Greater London, England, United Kingdom
recover predictably. The platform is reliable, incident detection and resolution are smooth due to proactive monitoring, well‐maintained alerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategy is defined and operational, working across the business with validated requirements and outcomes. Incident management processes ...

Platform Engineering Manager (Cloud Foundations)

Location
Greater London, England, United Kingdom
motivations. Platform Engineering, Operations&Reliability Cloudplatform and keyservices arereliable,incident detection and resolutionaresmooth due to proactive monitoring, well‐maintainedalerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategyisdefined and operational,working across the business with validated requirements and outcomes. Backup,restoreand resilience mechanisms are regularly ...

Manager, Senior Platform Software Engineer

Location
Greater London, England, United Kingdom
that enable them to work more effectively and better serve our clients. You will help drive engineering best practices, ensuring that quality, reliability and observability are built into our products from inception. You will design and implement robust CI/CD automation flows, continually evolving our processes to improve ...

Senior Infrastructure Engineer (GCP) - Engine by Starling

Location
Cardiff, Wales, United Kingdom
Dedicated/Partner Interconnect Experience with Workload Identity and Workload Identity Federation for keyless authentication of workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity ...

Staff Infrastructure Engineer (GCP) - Engine by Starling

Location
Manchester, England, United Kingdom
Dedicated/Partner Interconnect Experience with Workload Identity and Workload Identity Federation for keyless authentication of workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity ...

Staff Infrastructure Engineer (GCP) - Engine by Starling

Location
Greater London, England, United Kingdom
Dedicated/Partner Interconnect Experience with Workload Identity and Workload Identity Federation for keyless authentication of workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity ...

Technical Site Reliability Engineer

Hiring Organisation
Anduril Industries
Location
London, United Kingdom
Salary
£ 70 K
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative ...

Architect – Security Solutions

Location
United Kingdom
Job Details: Architect – Security Solutions Full details of the job. Vacancy Name Vacancy Name Architect – Security Solutions Vacancy No Vacancy No VN666 Employment Type Employment Type Full-Time Location Location Home Type of Vacancy Type ...

Cyber Platform Engineer

Hiring Organisation
Hays Specialist Recruitment Limited
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£550.00 - £650.00 per day
Role Title: Cyber Platform Engineer (Python, Cloud, DevOps, Terraform, Cyber Security) Location: SheffieldRate - £550 - £650 - Inside IR35 Role Description: We're looking for a strong Software Engineer to join the Cyber Security Strategic Engineering team. ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
million Series C funding round – the largest fundraise ever completed by a quantum computing company in Europe. The Purpose As a Monitoring and Observability Engineer, you'll help keep OQC's live quantum computing systems running at their best. By improving system visibility, developing intelligent monitoring solutions and driving operational … Live Services team, you'll monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
secure supply chain configurations; Establish IaC and GitOps standards with automated testing for every infrastructure change; Prototype agentic infrastructure components, including deployment and observability platforms in service meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/… Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay current with CNCF and AI ecosystem innovations, including eBPF observability and agent‐aware orchestration; Lead and mentor a team of DevOps engineers while remaining hands‐on. Требования: Experience leading or mentoring engineering teams, setting direction ...

SC Cleared Tester (Performance)

Hiring Organisation
VIQU IT Recruitment
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £475.00 per day
role of Performance Tester you will enhance and shape the scalability and reliability - creating strategies and solutions in cloud and container based environments, leveraging observability tools and diagnostics to optimise application performance. Essential Criteria • Ability to design, execute, and analyse performance, load, stress, and volume tests. • Familiarity with monitoring … observability tools (Grafana, Splunk) and diagnostics platforms (New Relic, Dynatrace). • Understanding of containerisation technologies (Docker, Kubernetes) and big data platforms (Databricks). • Ability to analyse complex performance issues and provide actionable recommendations. • Strong scripting skills (e.g., Python, JavaScript) for test automation and data analysis. • Knowledge of Performance Test Strategy ...

Senior Forward Deployed ML Engineer, Agents

Location
Greater London, England, United Kingdom
agent development, MLOps pipeline implementation, and production optimization. You understand what makes agents perform well in production and how to systematically improve quality through observability and evaluation. Experience with voice AI platforms, RAG systems, and LLM orchestration frameworks is highly desirable. You bring exceptional communication skills, customer empathy … validate datasets for fine-tuning, evaluation, and synthetic data generation Work with other MLEs, MLOps, SREs to carry out model deployment and productionization Observability, Evaluation & Production Operations Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure Instrument applications ...

Senior ML Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
maintain cloud-native infrastructure using Kubernetes and Infrastructure as Code technologies such as Terraform or Bicep. Develop robust CI/CD processes and observability frameworks to ensure reliable and secure ML operations. Collaborate closely with Data Scientists, software engineers, and client-facing teams to deliver scalable AI solutions. Influence technical … containerised workloads. Expertise in Infrastructure as Code using Terraform, Bicep, Pulumi, or comparable technologies. Experience building CI/CD pipelines and implementing monitoring and observability practices. Experience working with cloud platforms, ideally Azure, although other cloud backgrounds will be considered. Exposure to LLMs, NLP, text analytics, or generative AI applications. ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Computershare
Location
Bristol, Gloucestershire, United Kingdom
Salary
£ 70 K
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … services while fostering a culture of shared ownership and continuous learning.Other key responsibilities:Drive adoption of SRE principles (SLOs, error budgets, toil reduction).Establish observability and monitoring standards.Lead automation-first operations.Improve incident and problem management maturity.Partner with software and infrastructure engineering teams to embed reliability into the product lifecycle.Establish ...

Head of Site Reliability Engineering (SRE)

Location
West of England, England, United Kingdom
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … fostering a culture of shared ownership and continuous learning. Other key responsibilities: Drive adoption of SRE principles (SLOs, error budgets, toil reduction). Establish observability and monitoring standards. Lead automation-first operations. Improve incident and problem management maturity. Partner with software and infrastructure engineering teams to embed reliability into ...

Staff Platform Engineer – Platform Engineering (UK & Ireland)

Location
United Kingdom
reliability, scalability, and developer productivity. What You’ll Do Help drive technical design and delivery of core platform capabilities spanning cloud infrastructure, Kubernetes platforms, Observability, CI/CD systems, and developer tooling Architect, build, and operate cloud-native platforms on AWS with a focus on reliability, security, scalability, and cost … developer needs Participate in and/or lead complex, cross-team initiatives and provide technical clarity in ambiguous problem spaces Improve operational excellence through observability, automation, and incident response Mentor engineers and participate in on-call rotations What You Bring 8+ years of experience in platform, infrastructure, DevOps, or software ...

Infrastructure Software Engineer, Apps Platform

Location
Greater London, England, United Kingdom
Apps Platform, London As an Infrastructure Engineer on the Apps Platform Infrastructure team, you'll help build and evolve our platform’s deployment and observability layers across multiple cloud providers and on-premises, for both internal and customer-managed environments. This is a role for someone who cares about building … both on-premise, all major cloud platforms and beyond Ensure fast, secure, and reproducible deployments across all supported CSPs and on-prem Expand the observability of the platform so that forward deployed infrastructure teams can easily and efficiently operate customer deployments Partner closely with internal teams deploying the platform ...