2,451 to 2,475 of 2,830 Remote/Hybrid Observability Jobs

Senior Platform Engineer

Hiring Organisation
WeDo Technology Solutions Limited
Location
Croydon, Surrey, United Kingdom
Employment Type
Full-Time
Salary
£85,000 - £95,000 per annum
platform changes through Terraform • Improve deployment processes and reduce the risk of downtime • Troubleshoot platform, networking and workload issues in production • Improve monitoring, observability and operational processes • Work closely with Software Engineering, Security and Architecture teams Required Skills Kubernetes/AKS is essential. You should also have strong experience across … Microsoft Azure • Terraform/Infrastructure as Code • Azure Front Door • Azure API Management • ArgoCD/GitOps • Renovate • CI/CD and deployment automation • Observability and production troubleshooting • Highly available, business-critical production environments Experience supporting high-volume SaaS platforms handling millions of requests would be particularly valuable. Why should ...

Senior Software Engineer - k6 Core | United Kingdom | Remote

Location
United Kingdom
Senior Software Engineer - k6 Core | United Kingdom | Remote United Kingdom (Remote) Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana … optimization Distributed systems or cloud-based services Backend systems for web or mobile applications Tools and platforms such as Docker, AWS, microservices architectures, and observability tools like Grafana or APM systems Compensation & Rewards: In the United Kingdom, the Base compensation range for this role ...

platform engineer for AI data workflows

Location
Greater London, England, United Kingdom
agents, including MCP connectors and workflow skills Design AI-powered data workflows for data engineers, analysts, business teams, and AI tools Implement data observability, quality frameworks, monitoring, and alerting Own infrastructure-as-code, CI/CD for data workflows, access management, security configurations, audit trails, and spend governance Build golden … data infrastructure Experience with Terraform/Terragrunt, CI/CD for data workflows, and Docker/Kubernetes Production-quality Python proficiency Reliability mindset, including observability, meaningful alerts, production ownership, and incident-driven improvements Platform-product mindset focused on developer experience and reducing friction for platform users Communicates effectively across different ...

Senior Architect Private Cloud

Hiring Organisation
Randstad Technologies Recruitment
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £550/day Inside IR35 via Umbrella
patterns, and platform guardrails. Kubernetes Platform Leadership: Own the end-to-end Kubernetes ecosystem, including cluster topology, multi-tenancy, ingress, service mesh, secrets management, observability, and seamless workload onboarding. Event-Driven Architecture: Direct the adoption and governance of Apache Kafka, covering partitioning strategies, schema management, resilience patterns, capacity planning … configuration management. Desirable Skills: Familiarity with Service Mesh (e.g., Istio), API Gateways, Policy as Code (e.g., OPA), developer portals (Platform-as-a-Product), and Observability stacks (OpenTelemetry, Prometheus, ELK). Randstad Technologies is acting as an Employment Business in relation to this vacancy. ...

Lead Data Engineer

Hiring Organisation
Develop
Location
Solihull, West Midlands, United Kingdom
Employment Type
Permanent
Salary
£95,000
data solutions. Key Responsibilities Lead the technical design and architecture of data pipelines, models and platform components. Establish standards across code quality, testing, security, observability and documentation. Drive technical discovery, design sessions, code reviews and resolution of technical debt. Design and deliver scalable batch and streaming data solutions. Mentor … similar SQL-based transformation tooling. Understanding of CI/CD, Git, automated testing and infrastructure as code, particularly Terraform. Experience implementing data quality, monitoring, observability and alerting. Ability to lead architecture decisions and communicate technical trade-offs clearly. Desirable: Spark/DataProc, Looker or similar BI tools, analytics engineering, data ...

Senior Infrastructure Platform Engineer – Veeam / VMware / Hyper-V

Hiring Organisation
100 Percent
Location
Bristol, City of Bristol, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum Bonus + Benefits
major incident recovery, particularly around backup, restore and platform availability. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. … Azure DevOps-focused role. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability to coordinate ...

Senior Infrastructure Platform Engineer Veeam / VMware / Hyper-V

Hiring Organisation
100% IT Recruitment Ltd
Location
United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
major incident recovery, particularly around backup, restore and platform availability. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. … Azure DevOps-focused role. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability to coordinate ...

Engineering & Release Lead

Location
Melbourne, England, United Kingdom
role built around three critical pillars: people leadership of the software engineering team, ownership of release and build management, and accountability for production platform observability and operations. On the people side, you'll provide direct line management for Movember's software engineers, running regular one-to-ones, setting clear expectations … model, making operational accountability a natural part of how the team works rather than something bolted on after the fact. On the observability and operations side, you'll own platform monitoring across the production environment, using insight to drive continuous improvement in reliability, performance, and incident response. You'll lead ...

Staff Backend Engineer - Data Platform

Location
Greater London, England, United Kingdom
drive the technical vision and implementation for our foundational data platform — from experimentation, event ingestion pipelines to our data lake, governance frameworks, and data observability and real-time analytics capabilities. You will work closely with data scientists, machine learning engineers, backend teams, and product leaders, providing deep technical expertise … analytics. Champion engineering excellence, setting high technical standards and advocating for best practices in system design, maintainability, performance, and privacy. Lead efforts in data observability, governance, and privacy-by-design principles, ensuring their robust implementation across the organization. Mentor and coach engineers, elevating the technical capabilities of the team ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Director, Generative AI Experience Engineering (EMEA) - London

Location
Greater London, England, United Kingdom
contextual generation, retrieval and adaptive orchestration. Work across modern LLM ecosystems including foundation models, retrieval pipelines, vector storage, embeddings, multi-agent systems and AI observability tooling. Design systems that support conversational rendering, dynamic content assembly, streaming UX and adaptive interfaces. Actively experiment with AI tooling, workflows and engineering methodologies, bringing … models RAG pipelines, vector databases, embeddings and chunking Tool calling, agentic systems, orchestration and memory frameworks Prompt engineering, prompt management and context engineering AI observability, evaluations, governance and guardrails AI-Augmented Engineering - AI-assisted development tooling including Claude Code, Cursor and Codex Prompt-driven engineering and AI pair-programming workflows ...

Senior Backend Software Engineer

Hiring Organisation
Spire
Location
Boulder, Colorado, United States
Employment Type
Permanent
Salary
USD Annual
next generation of Spire's File Transfer architecture (Rust) Ensure that Spire's file transfer services meet internal and mission-driven security standards Enhance observability and fault-tolerance of Spire's file transfer services Take responsibility for the performance of critical file transfer services that contribute to the movement … customers to ensure alignment and optimal integration with interfacing hardware and software services Create and maintain comprehensive documentation for APIs and system architectures Implement observability solutions for space and ground-based services to provide a clear and actionable view of service performance Basic Skills/Qualifications: 5+ years' experience ...

AI Platform Engineer

Location
Greater London, England, United Kingdom
patterns that support the safe and scalable adoption of AI-assisted development.Improve developer experience through streamlined workflows, tooling integration and self-service capabilities. Develop observability and measurement capabilities to provide insights into engineering productivity, quality and platform adoption. Collaborate with Technical Leads and AI Software Engineers to identify recurring engineering … cost optimisation, including token monitoring, caching strategies and model selection considerations. Knowledge of approaches for managing and reducing token consumption costs.Deep understanding of observability, automated testing and software delivery tooling. Knowledge of platform security, governance and operational controls. Strong programming, automation and problem-solving capabilities. Passionate about improving developer experience ...

Senior Software Engineer (Application Operations)

Location
Greater London, England, United Kingdom
Application Operations function. You’ll apply these practices across our AWS environment (EKS), PHP services, and MySQL databases, using Datadog as our observability platform, Kibana for log exploration, and Heap to help quantify and understand customer impact. Key Responsibilities L2.5 Operations Delivery Provide high-quality, timely L2.5 support … playbooks for known issue patterns across workloads, services, and operational scenarios. Identify and automate repetitive remediation tasks to reduce manual toil and improve MTTR. Observability and Service Readiness Collaborate with peers to ensure the right monitoring signals, dashboards, and alerts exist in Datadog. Tune app‐level alerts and dashboards ...

Director of Software Engineering

Hiring Organisation
Spire Global
Location
Glasgow, UK
Employment Type
Full-time
infrastructureStay hands-on: review code, prototype solutions, and get into the details when it mattersEstablish engineering standards across code quality, system design, testing, and observability, and hold the team to themBe the person engineers come to when the problem is genuinely hardTeam Building & Culture Recruit, develop, and retain a team … monitoring problemsExperience writing performance software in RustBackground in space systems, aerospace, or highly constrained real-time environmentsExperience building data lakes, telemetry platforms, or observability infrastructure at scaleA history of leading teams through technical transformations and not just maintaining the status quoSpire operates a hybrid work model, and this position will ...

ML Data & Platform Engineer

Location
United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

ML Data & Platform Engineer

Location
Cambridge, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

ML Data & Platform Engineer

Location
Greater London, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

Software Engineer - Affiliate Operations

Location
Greater London, England, United Kingdom
remain highly reliable while evolving to support new markets and acquisition channels. Reliability is fundamental to everything we build. We invest heavily in automation, observability, and operational excellence to reduce manual effort. We are also exploring how AI can transform our engineering productivity and marketing platforms. We operate … improve platform effectiveness, launch new capabilities, and reduce manual operational effort across the business. Improve You'll continuously improve our systems through better observability, automation, and thoughtful refactoring. You'll help evolve our architecture and engineering practices to ensure our platforms remain resilient as they scale. Own You'll take ...

Technical Lead, Lending & Savings

Hiring Organisation
Blockchain
Location
London, UK
Employment Type
Full-time
performance, security, and maintainability. Remain hands-on, contributing production-quality code and reviewing critical changes. Drive engineering best practices across system design, testing, deployment, observability, and operational excellence. Mentor engineers through code reviews, technical coaching, and day-to-day leadership. Engineering DeliveryOwn the delivery of technical initiatives from design through … design. Experience working with Redis or other NoSQL technologies. Deep understanding of microservices architecture, APIs, distributed systems, and cloud-native applications. Experience with monitoring, observability, incident response, and production operations. Strong debugging and performance optimisation skills. Demonstrated experience shipping reliable production systems that process financial transactions. LeadershipExperience leading engineering teams ...

Senior Full Stack Engineer (Realtime & Voice) Customer Experience Platform

Location
Greater London, England, United Kingdom
Build the safety and compliance plumbing enterprise partners audit, including guardrails, content filtering, and PII redaction integration points Keep revenue-critical deployments healthy through observability, alerting, incident response, and SLA performance Build the platform capabilities forward-deployed engineers configure for partner telephony integrations and go-lives Raise the engineering … standard part of their workflow, with the judgment to review, correct, and own everything that ships An operable-systems mindset, covering SLAs, observability, on-call rotations, and rollback plans The ability to break down complex problems, make pragmatic tradeoffs, and ship iteratively, backed by strong communication across product, design ...

Senior Applied AI Engineer (Defence Contractor)

Location
United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Senior Applied AI Engineer (Defence Contractor)

Location
Greater London, England, United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Director - Platform Engineering & Architecture Lead

Location
Greater London, England, United Kingdom
Engineer the future of global finance. At Citi, our Tech team doesn’t just support finance – we are helping to redefine it. Every day, $5 trillion crosses through our network. We do business in 180+ ...

Principal Consultant - Cloud & Engineering

Location
Greater London, England, United Kingdom
## Principal Consultant - Cloud & EngineeringApplylocations: Londontime type: Full timeposted on: Posted Todayjob requisition id: JR100890Founded in Switzerland in 1968, Zühlke is owned by its partners and located across Europe and Asia. We are a global ...