276 to 300 of 388 Remote/Hybrid Observability Jobs

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Platform Chapter Lead - Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Lead Java Developer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc. Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

Staff or Principal Data Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
create unnecessary complexity, risk, duplicated capability or long‐term support burden. Raise the quality bar for data products through clear ownership, robust testing, reconciliation, observability, lineage, documentation, performance and supportability. Collaborate with cross‐functional teams to address security, GDPR, PII handling, role‐based access, auditability and data governance are designed … services across batch, streaming and event‐driven patterns. Deep understanding of engineering practice: clean design, testing strategy, CI/CD, infrastructure as code, observability, performance, security, incident response and DevSecOps. Experience with cloud data services and modern data stacks. Relevant technologies may include Snowflake, Azure/AWS/GCP data ...

Head of Technology Operations

Hiring Organisation
Jobleads-UK
Location
Halifax, England, United Kingdom
adoption of infrastructure as code (IaC), CI/CD pipelines, and automated testing within platform operations. Champion site reliability engineering (SRE) practices, embedding monitoring, observability, and incident response, and continuously improving performance metrics including application load times, throughput, and error rates. Partner with the Head of Development and the Director … technologies. Deep expertise in Kubernetes, cloud networking, CI/CD pipelines, infrastructure as code, and platform security. Experience leading site reliability engineering (SRE), monitoring, observability, and performance optimisation (load times, application speed, latency management). Experience providing senior‐level escalation support for complex infrastructure, firewall, networking, and systems issues. ITIL ...

Global Head of SRE & Reliability – Hybrid Role

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
office and collaboration across Engineering, Infrastructure Operations and Security to boost reliability and performance of critical platforms. You will define reliability strategy, drive automation, observability, and incident maturity, and scale the organization while embedding reliability into the #J-18808-Ljbffr ...

DevOps Engineer

Hiring Organisation
Fruition Group
Location
Leeds, Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
Contract: Inside IR35 We're seeking an experienced Senior DevOps Engineer to join a small, highly skilled engineering team delivering a large-scale enterprise observability platform as they move away from Splunk This is an opportunity to work on a critical cloud platform supporting the migratio click apply for full ...

Data Platform Lead — AI-Ready, Scalable Pipelines (Hybrid)

Hiring Organisation
Jobleads-UK
Location
Cardiff, Wales, United Kingdom
models, ensuring reliability and AI readiness while coaching engineers and setting technical standards. You will balance immediate delivery with long-term platform evolution, improve observability and data quality, and drive governance across datasets to empower self-service analytics. #J-18808-Ljbffr ...

Lead Platform Engineer - Cloud & Security (Remote)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Platform Engineer to own the infrastructure and security foundation that our engineering teams rely on. You’ll manage cloud infrastructure end-to-end, establish observability practices, and drive security posture across the platform in a hybrid environment with remote options. Based in London or remote in Europe, you’ll play ...

Senior Backend Engineer - Remote or Hybrid, Warehouse Data

Hiring Organisation
Jobleads-UK
Location
Cambridge, England, United Kingdom
contracts that power the product — ingestion pipelines transforming warehouse data into a model, robust APIs, durable storage, and reliable background work. You will ensure observability that lets a small team operate them confidently at 3 am. Two engines power WareBee: Physical AI and Process AI. You will build systems ...

Principal AI Infrastructure Architect - Remote

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
squads to set standards and deliver scalable systems. This role reports to the Engineering Director, focusing on real-time data pipelines, LLM-driven agents, observability, and privacy-driven data governance, enabling the next generation of Grip's event platform. #J-18808-Ljbffr ...

AI Security Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
system design Embedding security within the SDLC Exposure to AI and agentic systems (practical or conceptual), including: Security and governance for AI Solutions (Identity, observability, risks) Agentic Architecture and patterns (MCP, RAG, Tools, Memory, Harnesses) Experience working in complex, enterprise‐scale environments Ability to translate security requirements into engineering‐aligned ...

AI Security Engineer

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Contract
Embedding security within the SDLC * Exposure to AI and agentic systems (practical or conceptual), including: o Security and governance for AI Solutions (Identity, observability, risks) o Agentic Architecture and patterns (MCP, RAG, Tools, Memory, Harnesses) * Experience working in complex, enterprise-scale environments * Ability to translate security requirements into engineering-aligned ...

Senior C++ Engineer

Hiring Organisation
Randstad Digital
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £700 per day
shared library across iOS and Android layers. Evolve Protocols: Shape client-server contracts, handle network state detection, and optimize real-time messaging systems. Enhance Observability: Build deep telemetry and metrics to diagnose latency and performance anomalies across the full client-to-backend path. Kill Legacy Code: Gracefully sunset legacy networking ...

Lead Data Architect - News Engineering

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
quality across the platform. Partner closely with engineering teams to improve signal discoverability through normalizing metadata and implementing industry standards. Ensure data architectures support observability, lineage, and a simple, straightforward experience for customers. Encourage pragmatic experimentation with new tools, libraries, and approaches to improve reliability, performance, or developer experience. Qualifications ...

Lead News Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
quality across the platform. Partner closely with engineering teams improving signal discoverability through normalizing metadata and implementing industry standards. Ensure data architectures support observability and lineage while providing customers with a simple, straight‐forward experience for working with News from LSEG. We encourage pragmatic experimentation with new tools, libraries ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
driven business valueTranslate business requirements into AI-enabled workflow solutionsSolution Design & ArchitectureDesign and support implementation of:AI Control Tower (AI lifecycle management, governance, observability)Agentic AI workflows enabling autonomous executionNow Assist/GenAI use cases across workflowsDefine data, integration, and workflow architectures for AI-enabled ServiceNow solutionsConversational AI & Employee ExperienceSupport ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Head of SRE to define, lead and evolve our global reliability strategy. This senior leadership role is responsible for driving operational excellence, service reliability, observability, automation and continuous improvement across our technology landscape. Key Responsibilities Drive adoption of SRE principles (SLOs, error budgets, toil reduction). Establish observability and monitoring … with strong scripting and development capabilities using technologies such as Python, PowerShell, Bash, Terraform and Ansible Automation Platform. Other key skills: Robust knowledge of observability and monitoring practices, and experience implementing and managing platforms such as Dynatrace, Prometheus, Grafana, and Splunk. Good understanding of CI/CD tooling and modern ...

Associate/Vice President, AI Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
aligned with firmwide risk and compliance standards. Implement and maintain infrastructure as code and automation to ensure repeatable, auditable platform provisioning. Build and operate observability, monitoring, and alerting solutions for AI platforms, ensuring availability, performance, and cost transparency. Collaborate with Security and Risk partners to integrate identity, access controls, data … cloud compute, networking, storage, and security services. Understanding of ML platform operations and governance concepts, including model deployment strategies, lifecycle management, monitoring/observability, and Disaster Recovery. Experience supporting LLMs, generative AI platforms, or model serving infrastructure. Experience supporting AI and machine learning workloads, with exposure to managed compute ...

Azure Platform Engineer

Hiring Organisation
EMBS Engineering
Location
Newbury, West Berkshire, Berkshire, United Kingdom
Employment Type
Permanent
Salary
£65000 - £75000/annum + Benefits
Azure platform templates and engineering patterns to support consistent platform adoption. Build and improve CI/CD pipelines, deployment automation and release processes. Implement observability, monitoring, logging, resilience and operational readiness across the platform. Embed FinOps principles, improving cloud cost visibility and optimisation. Work closely with engineering teams throughout sprint … engineering patterns. Strong understanding of DevOps and DevSecOps practices including CI/CD, source control, automated testing and release management. Experience with Azure observability, monitoring, alerting, logging and platform reliability. Practical knowledge of cloud security, governance and enterprise engineering standards. Experience applying FinOps principles including tagging strategies, right-sizing ...

Platform Engineer (SC Cleared or Eligible)

Hiring Organisation
Sanderson Recruitment
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£60,000
cost efficiency. Ensure platform security, governance and compliance standards are maintained. Act as an escalation point for platform-related issues. Champion DevOps culture, observability and operational excellence. Contribute to the adoption of modern cloud-native and serverless architectures. Skills & Experience Essential Experience working in Platform Engineering, DevOps, Infrastructure … other CI/CD tooling experience. Scripting experience with PowerShell, Python, Bash or Go. Experience supporting and operating production environments. Knowledge of monitoring, observability and incident management. Understanding of security and compliance best practices. Excellent communication skills and a collaborative approach. Passion for continuous improvement and knowledge sharing. Desirable Microsoft ...

AI Platform engineer

Hiring Organisation
Nextech Group Limited
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
strategies, embedding generation, and vector store management Architect event-driven microservices using Kafka/SQS for async processing of high-volume inference requests Implement observability and cost-tracking for token usage across multiple LLM providers (Anthropic, OpenAI, open-source models via vLLM) Own database performance for both relational (Postgres … pgvector Infra: AWS (ECS, Lambda, SQS/SNS), Docker, Kubernetes, Terraform AI/ML tooling: LangChain/LlamaIndex, vLLM, Anthropic & OpenAI APIs, embedding models Observability: Datadog, Grafana, OpenTelemetry CI/CD: GitHub Actions, ArgoCD Requirements: 4+ years backend development experience, ideally with at least 1 year working with LLM/ ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

DevSecOps Engineering Lead CGEMJP00346044

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
management, including allowlist processes and risk acceptance where required Secrets management and identity/access management Policy enforcement for workloads, container images and infrastructure Observability, monitoring, logging and audit controls Partner with developers to embed secure-by-design engineering and ensure compliance with CLIENT security standards. Enable and govern Infrastructure … compliance tooling (e.g. Trivy scanning and vulnerability management, HashiCorp Vault, cert-manager) Containers and orchestration (e.g. Docker, AWS EKS) Infrastructure as Code (e.g. Terraform) Observability (e.g. Grafana, Loki) Scripting and automation (e.g. Python, Bash) Cloud and networking fundamentals (e.g. AWS IAM, S3, network policies) Experience delivering within the UK Government ...

DevSecOps Engineering Lead CGEMJP00346044

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Contract, Work From Home
management, including allowlist processes and risk acceptance where required Secrets management and identity/access management Policy enforcement for workloads, container images and infrastructure Observability, monitoring, logging and audit controls Partner with developers to embed secure-by-design engineering and ensure compliance with CLIENT security standards. Enable and govern Infrastructure … compliance tooling (e.g. Trivy scanning and vulnerability management, HashiCorp Vault, cert-manager) Containers and orchestration (e.g. Docker, AWS EKS) Infrastructure as Code (e.g. Terraform) Observability (e.g. Grafana, Loki) Scripting and automation (e.g. Python, Bash) Cloud and networking fundamentals (e.g. AWS IAM, S3, network policies) Experience delivering within the UK Government ...