901 to 925 of 6,285 Permanent Observability Jobs

Cyber Platform Engineer

Hiring Organisation
Hays Specialist Recruitment Limited
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£550.00 - £650.00 per day
Role Title: Cyber Platform Engineer (Python, Cloud, DevOps, Terraform, Cyber Security) Location: SheffieldRate - £550 - £650 - Inside IR35 Role Description: We're looking for a strong Software Engineer to join the Cyber Security Strategic Engineering team. ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
million Series C funding round – the largest fundraise ever completed by a quantum computing company in Europe. The Purpose As a Monitoring and Observability Engineer, you'll help keep OQC's live quantum computing systems running at their best. By improving system visibility, developing intelligent monitoring solutions and driving operational … Live Services team, you'll monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce ...

Platform Engineer: Scalable AWS, CI/CD & Observability

Location
Greater London, England, United Kingdom
availability and performance of our platform Evolve and optimise our CI/CD pipelines to enable fast, safe and consistent deployments Implement and enhance observability (monitoring, alerting, logging) using tools like DataDog to ensure high-availability APM Apply best practices in security, scalability and cost optimisation across our infrastructure Support ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
secure supply chain configurations; Establish IaC and GitOps standards with automated testing for every infrastructure change; Prototype agentic infrastructure components, including deployment and observability platforms in service meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/… Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay current with CNCF and AI ecosystem innovations, including eBPF observability and agent‐aware orchestration; Lead and mentor a team of DevOps engineers while remaining hands‐on. Требования: Experience leading or mentoring engineering teams, setting direction ...

Expert Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

SC Cleared Tester (Performance)

Hiring Organisation
VIQU IT Recruitment
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £475.00 per day
role of Performance Tester you will enhance and shape the scalability and reliability - creating strategies and solutions in cloud and container based environments, leveraging observability tools and diagnostics to optimise application performance. Essential Criteria • Ability to design, execute, and analyse performance, load, stress, and volume tests. • Familiarity with monitoring … observability tools (Grafana, Splunk) and diagnostics platforms (New Relic, Dynatrace). • Understanding of containerisation technologies (Docker, Kubernetes) and big data platforms (Databricks). • Ability to analyse complex performance issues and provide actionable recommendations. • Strong scripting skills (e.g., Python, JavaScript) for test automation and data analysis. • Knowledge of Performance Test Strategy ...

Senior Forward Deployed ML Engineer, Agents

Location
Greater London, England, United Kingdom
agent development, MLOps pipeline implementation, and production optimization. You understand what makes agents perform well in production and how to systematically improve quality through observability and evaluation. Experience with voice AI platforms, RAG systems, and LLM orchestration frameworks is highly desirable. You bring exceptional communication skills, customer empathy … validate datasets for fine-tuning, evaluation, and synthetic data generation Work with other MLEs, MLOps, SREs to carry out model deployment and productionization Observability, Evaluation & Production Operations Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure Instrument applications ...

Senior ML Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
maintain cloud-native infrastructure using Kubernetes and Infrastructure as Code technologies such as Terraform or Bicep. Develop robust CI/CD processes and observability frameworks to ensure reliable and secure ML operations. Collaborate closely with Data Scientists, software engineers, and client-facing teams to deliver scalable AI solutions. Influence technical … containerised workloads. Expertise in Infrastructure as Code using Terraform, Bicep, Pulumi, or comparable technologies. Experience building CI/CD pipelines and implementing monitoring and observability practices. Experience working with cloud platforms, ideally Azure, although other cloud backgrounds will be considered. Exposure to LLMs, NLP, text analytics, or generative AI applications. ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Computershare
Location
Bristol, Gloucestershire, United Kingdom
Salary
£ 70 K
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … services while fostering a culture of shared ownership and continuous learning.Other key responsibilities:Drive adoption of SRE principles (SLOs, error budgets, toil reduction).Establish observability and monitoring standards.Lead automation-first operations.Improve incident and problem management maturity.Partner with software and infrastructure engineering teams to embed reliability into the product lifecycle.Establish ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Computershare
Location
Bristol, UK
Employment Type
Full-time
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … while fostering a culture of shared ownership and continuous learning. Other key responsibilities: Drive adoption of SRE principles (SLOs, error budgets, toil reduction).Establish observability and monitoring standards. Lead automation-first operations. Improve incident and problem management maturity. Partner with software and infrastructure engineering teams to embed reliability into ...

Head of Site Reliability Engineering (SRE)

Location
West of England, England, United Kingdom
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … fostering a culture of shared ownership and continuous learning. Other key responsibilities: Drive adoption of SRE principles (SLOs, error budgets, toil reduction). Establish observability and monitoring standards. Lead automation-first operations. Improve incident and problem management maturity. Partner with software and infrastructure engineering teams to embed reliability into ...

Staff Platform Engineer – Platform Engineering (UK & Ireland)

Location
United Kingdom
reliability, scalability, and developer productivity. What You’ll Do Help drive technical design and delivery of core platform capabilities spanning cloud infrastructure, Kubernetes platforms, Observability, CI/CD systems, and developer tooling Architect, build, and operate cloud-native platforms on AWS with a focus on reliability, security, scalability, and cost … developer needs Participate in and/or lead complex, cross-team initiatives and provide technical clarity in ambiguous problem spaces Improve operational excellence through observability, automation, and incident response Mentor engineers and participate in on-call rotations What You Bring 8+ years of experience in platform, infrastructure, DevOps, or software ...

Infrastructure Software Engineer, Apps Platform

Location
Greater London, England, United Kingdom
Apps Platform, London As an Infrastructure Engineer on the Apps Platform Infrastructure team, you'll help build and evolve our platform’s deployment and observability layers across multiple cloud providers and on-premises, for both internal and customer-managed environments. This is a role for someone who cares about building … both on-premise, all major cloud platforms and beyond Ensure fast, secure, and reproducible deployments across all supported CSPs and on-prem Expand the observability of the platform so that forward deployed infrastructure teams can easily and efficiently operate customer deployments Partner closely with internal teams deploying the platform ...

SC Cleared DevOps Engineer

Location
United Kingdom
Develop and maintain Infrastructure as Code solutions. Create and support CI/CD pipelines to enable secure, repeatable deployments. Implement monitoring, logging, alerting and observability capabilities. Ensure platform resilience, availability and operational readiness. Manage configuration, secrets and environment controls. Embed security best practice throughout the delivery lifecycle. Support incident response … Code, ideally Terraform, CloudFormation or AWS CDK. Strong CI/CD pipeline experience. Experience with containerisation technologies, ideally Docker and Kubernetes. Strong understanding of observability, monitoring and operational support. Knowledge of cloud security, IAM, secrets management and governance controls. Experience delivering and supporting production cloud environments within secure or regulated ...

Intelligent Automation Engineering Manager

Location
Greater London, England, United Kingdom
deployment and operational support of AI agents and AI-powered solutions. Establish engineering standards and best practices for AI architecture, orchestration, retrieval, tool invocation, observability, governance, privacy, security and cost management. Review technical designs and architecture documentation to ensure solutions align with engineering, security and governance standards. Translate emerging … with MCP (Model Context Protocol), MCP Servers, MCP Clients or enterprise AI integration frameworks. Knowledge of LLMOps, AI evaluation frameworks, model routing and AI observability tooling. Experience with enterprise integration technologies including APIs, Middleware, ESB or iPaaS platforms. Hands‐on experience integrating internal and third-party systems. Azure cloud experience ...

Senior Site Reliability Engineer (Azure)

Hiring Organisation
GTN Technical Staffing & Consulting
Location
Plano, Texas, United States
Employment Type
Any
Salary
USD Annual
Site Reliability Engineer (SRE) to lead reliability engineering initiatives across our Azure estate and Command Center operations. This role focuses on scripting, automation, and observability to ensure uptime, performance, and rapid incident response. The Senior SRE will design and implement monitoring-as-code, optimize alerting, and build self-healing automation … professional manner at all times. Team Members are required to observe the companys standards, work requirements and rules of conduct. Essential Duties & Responsibilities: Observability & Monitoring o Architect end-to-end monitoring using Azure Monitor, Log Analytics, Application Insights, and ITRS Geneos. o Implement monitoring-as-code with Terraform/Bicep ...

DevOps Manager

Hiring Organisation
Stott & May Professional Search Limited
Location
United Kingdom
Employment Type
Permanent
live service. * Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. * Manage monitoring and observability, using tools such as Dynatrace or similar platforms to improve visibility, alerting and platform performance. * Oversee cloud infrastructure, environments and costs, identifying opportunities for optimisation … networking, including DNS, routing, firewalls and load balancing. * Experience managing enterprise-scale web applications, microservices and high-traffic digital platforms. * Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. * Experience managing incidents, problem management, root cause analysis and live production environments. * Understanding of application and cloud security ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Software Development Engineer in Test

Location
Greater London, England, United Kingdom
accuracy, reliability, latency, and cost. Work with Engineering, Product, Data Science, QA, and Operations teams to deliver production-ready solutions. Design for scalability, reliability, observability, and operational excellence across the platform. Troubleshoot complex production issues across application services, data pipelines, AI workflows, and telemetry. Participate in architecture, technical design, code … C++. Strong experience working with large-scale data, telemetry, logs, events, and operational datasets. Strong understanding of cloud platforms, system scalability, reliability, performance, observability, and production operations. Strong analytical, problem-solving, communication, and collaboration skills, with the ability to work effectively across technical and cross-functional teams. Desirable skills ...

DevOps Manager

Hiring Organisation
Stott & May Professional Search Limited
Location
London, UK
live service. * Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. * Manage monitoring and observability, using tools such as Dynatrace or similar platforms to improve visibility, alerting and platform performance. * Oversee cloud infrastructure, environments and costs, identifying opportunities for optimisation … networking, including DNS, routing, firewalls and load balancing. * Experience managing enterprise-scale web applications, microservices and high-traffic digital platforms. * Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. * Experience managing incidents, problem management, root cause analysis and live production environments. * Understanding of application and cloud security ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
environments Define and implement infrastructure-as-code using Terraform, CDK, or equivalent Monitor platform health, define SLOs/SLAs, and build alerting and observability tooling Respond to and lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security … Actions, Jenkins, or equivalent) Solid understanding of networking, security groups, load balancing, and DNS Experience with container orchestration (Docker, Kubernetes, or ECS) Familiarity with observability tooling (Datadog, Grafana, CloudWatch, or equivalent) Understanding of HIPAA infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site ...

Senior Java Developer

Hiring Organisation
Clarify Consultancy Ltd
Location
Leyland, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £65,000 per annum
such as Kafka. Develop and maintain automated unit and integration tests. Work with Docker and Kubernetes within modern cloud environments. Contribute to monitoring, logging, observability and application performance. Participate in Agile ceremonies including sprint planning, refinement, reviews and retrospectives. Work closely with Product, Architecture and other technical teams to understand … performance within distributed systems. Experience with CI/CD pipelines and modern development practices is essential, as is familiarity with application monitoring, logging and observability tools. The position sits within an Agile/Scrum environment, so strong problem-solving skills, analytical thinking and excellent communication abilities are key to working ...

Corporate KYC Principle Software Engineer - Executive Director

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
infrastructure Troubleshoot and resolve infrastructure issues and fix root causes of issues at the source Implement monitoring, logging and alerting solutions to ensure system observability Collaborate with engineers and architects to improve platform documentation, standards and adoption Requirements Strong hands-on experience with cloud platforms (AWS & Azure ideally) Hands …/VNets, load balancers, DNS, security groups/NSGs) Experience with secrets management and identity/access control (IAM, OIDC, Azure AD) Familiarity with observability tooling (Prometheus, Grafana, CloudWatch, Azure Monitor) Risk Benefit Statement Learn more about the LexisNexis Risk team and how we work here We know your well ...