1,651 to 1,675 of 2,464 Remote/Hybrid Observability Jobs

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational ExcellenceDesign and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry).Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skillsFamiliarity with infrastructure ...

Intermediate Data Engineer

Location
Cheadle, England, United Kingdom
maintain automated deployment, testing and operational workflows using GitLab pipelines, version control, Terraform/OpenTofu and scripting. Contribute to data quality, lineage, governance and observability practices to ensure data is trusted, traceable, secure and supportable throughout its lifecycle. Investigate data quality, platform reliability, pipeline performance and operational issues, identifying practical … willingness to learn, AI-assisted engineering tools such as Kiro, GitHub Copilot or similar technologies to improve engineering effectiveness and delivery outcomes. Familiarity with observability and troubleshooting tooling such as Dynatrace, OpenSearch/Elasticsearch or CloudWatch logs. Ability to communicate clearly, analyse problems, work collaboratively and explain technical concepts ...

Engineering Manager: Platform & Growth Lead (Hybrid)

Location
Bristol, England, United Kingdom
client journeys on React and React Native. You'll coach engineers, own delivery, guide architectural decisions, partner with product, and drive observability and high-velocity delivery in a regulated financial services environment. #J-18808-Ljbffr ...

Lead Software Engineer - Hybrid, Share Options, £80k+

Location
Wigan, England, United Kingdom
production. You will lead a small cross-functional team, delivering features in a fast-paced SaaS environment. You will help modernise microservices, improve observability, and enhance testability using modern architectural approaches. Hybrid working with two days in the office is available. #J-18808-Ljbffr ...

Senior Value Engineer – Remote: Growth & ROI Storytelling

Location
Pathhead, Scotland, United Kingdom
persuasive value stories. The role emphasizes storytelling with data, enabling GTM teams, and partnerships to articulate Grafana's value propositions, with a focus on observability, open source roots, and a global remote #J-18808-Ljbffr ...

Senior Data Platform SRE — Hybrid, Obs & Reliability Leader

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer to embed within the Data Engineering team. You will own reliability, performance, and operability of the data platform, building observability from the ground up and leading incident response for data outages. You'll balance feature velocity with system stability, applying SRE practices and capacity planning ...

Senior SET: Platform Quality Engineer (Hybrid, London)

Location
Greater London, England, United Kingdom
architecture level, automate validation across environments, and manage delivery risk in distributed systems. You will define standards for testing distributed systems, enable observability-driven debugging, automate validation of availability and latency, and contribute to security posture. #J-18808-Ljbffr ...

Senior Data Analyst: AI/ML UX & Hybrid Data Pipelines

Location
City Of London, England, United Kingdom
three on-site days in City of London. Responsibilities include building scalable data pipelines (BigQuery, Dataflow/Apache Beam, Airflow), ensuring data quality and observability (Looker, Monte Carlo), and collaborating with product engineering and data science teams to plan data tracking and ingestion tasks. #J-18808-Ljbffr ...

Rust Sat OS Engineer — Space Systems, Hybrid Role

Location
East Hagbourne, England, United Kingdom
work with embedded teams on fault-tolerant designs. The role features a hybrid work model with three days per week in the office, supporting observability, testing, and documentation for new satellite software services. #J-18808-Ljbffr ...

Senior AI Engineer – GenAI, RAG & Agents (Hybrid)

Location
Greater London, England, United Kingdom
including Generative AI, RAG, and AI agents, collaborating with architecture, security, and infrastructure teams. You will lead the adoption of best practices, model evaluation, observability, and governance across SDLC while translating #J-18808-Ljbffr ...

Platform SRE Lead: On-Call, Reliability & Cloud Ops

Location
Greater London, England, United Kingdom
platform. Open to mid-level or senior SREs, the role is ops-heavy with focus on health of live systems, Kubernetes, cloud infra, deployments, observability, and collaboration with product teams. Hybrid work with 3 days in the office. #J-18808-Ljbffr ...

Senior AI Engineer — Production LLMs (Remote Europe)

Location
United Kingdom
Senior AI Engineer to build AI agents, agentic workflows, and production-grade AI capabilities. You will own end-to-end implementation, evaluation, and observability, interfacing with product and platform teams to deliver value to 3M+ traders. Location: Europe (EU/UK); fully remote with overlap to Central European hours. Start ...

Senior Java Backend Engineer - Hybrid London

Location
City Of London, England, United Kingdom
this hybrid role, you will collaborate with business analysts and stakeholders, develop RESTful APIs, optimize data access, and contribute to system stability, performance, and observability #J-18808-Ljbffr ...

Senior DevOps Engineer — Remote, Global Cloud Ops

Location
Belfast City District, Northern Ireland, United Kingdom
users, and data, across multiple regions. You will influence architecture, reliability, automation, and cloud efficiency, collaborating with engineers worldwide. The position emphasizes cost optimization, observability, and scalable deployment practices, with occasional travel to Belfast or remote collaboration. #J-18808-Ljbffr ...

Investment Tech Engineer — AI & Data Pipelines (Hybrid)

Location
United Kingdom
platform. You will work on cloud-based tools (Streamlit/ReactJS), apply AI-assisted tooling, and help drive CI/CD, data quality, and observability across the lifecycle. Collaborating with Cloud Teams, Enterprise Architecture and Security, you will translate business requirements into robust technical solutions, build data pipelines for analytics ...

MLOps Engineer — Real-Time AI in Production (Hybrid London)

Location
Greater London, England, United Kingdom
production systems, collaborating with research and engineering teams. The role emphasizes designing end-to-end ML pipelines, deploying and monitoring models in production, building observability and CI/CD, and enabling reproducible experimentation. #J-18808-Ljbffr ...

Senior ML Engineer - Conversational AI & GenAI | Hybrid

Location
Greater London, England, United Kingdom
full ML lifecycle from design to production, deploying models and building robust, scalable systems in collaboration with product managers. Strong focus on reliability, observability, and GenAI-enabled solutions. #J-18808-Ljbffr ...

Scala Engineer — Core Streaming Platform (Hybrid)

Location
Greater London, England, United Kingdom
contribute to design discussions, participate in reviews, and help juniors grow. Expect ownership of features end-to-end, with emphasis on reliability, observability, and production readiness. #J-18808-Ljbffr ...

Senior AI/ML Engineer — Agentic LLM Systems (Hybrid, London)

Location
Greater London, England, United Kingdom
services that support AI-driven workflows. You will work with researchers and engineers to diagnose failures, optimize latency and cost, and contribute to observability and monitoring of production AI systems. #J-18808-Ljbffr ...

Hybrid Backend Engineer — Commercial Planning

Location
Greater London, England, United Kingdom
planning platforms for Fashion, Home & Beauty, serving millions of customers and thousands of colleagues. Join us to craft robust, scalable backend services with modern observability and secure design. You will lead design and delivery of cloud-native solutions, mentor engineers, and collaborate with architecture and product teams in a hybrid ...

Senior AI Software Engineer - Hybrid, High-Impact Delivery

Location
Greater London, England, United Kingdom
secure production code, and mentor engineers and AI coding agents. The role blends deep software engineering with AI-native practices, overseeing code quality, testing, observability, deployment readiness, and governance while ensuring outputs are explainable and compliant with standards. #J-18808-Ljbffr ...

Senior GTM Automation Architect - AI-Powered Pipelines

Location
Wedmore, England, United Kingdom
enrichment, and routing, while shaping prompting architectures used by sales teams. You will influence architecture, vendor relations, and production readiness, with a focus on observability and data quality. This is a hands-on, production‐oriented role in a remote-first company. #J-18808-Ljbffr ...

Database Reliability Engineer - Hybrid, Cross-Cloud RDS

Location
Manchester, England, United Kingdom
data-focused engineering role within our hybrid data environment. You will help modernize and scale the RDS fleet, architect cross-cloud portability, and advance observability across a global database fleet. Strong Postgres, Kubernetes, and cloud-native experience are essential. Join a fast-moving fintech team with a strong focus ...

Remote Site Reliability Engineer – AI & Voice/UC Focus (UK)

Location
Greater London, England, United Kingdom
Intermedia is hiring a Site Reliability Engineer to ensure 24/7 service availability across global deployments. You will enhance reliability, observability, and scalability of cloud-based Voice/UC platforms with AI-enabled capabilities. This role collaborates with cross-functional teams and participates in incident response and post-incident ...