2,901 to 2,925 of 2,960 Remote/Hybrid Observability Jobs

Senior Specialist Engineer - Networks

Hiring Organisation
UK Health Security Agency
Location
Birmingham, Leeds, Liverpool or London (Canary Wharf), E14 4PU, United Kingdom
Salary
£56185.00 to £70566.00
mitigation strategies to reduce risk, improve service resilience and support secure cloud and on-premises environments. Service Management, Automation & Technical Leadership Lead the monitoring, observability, performance and capacity management of network services, integrating operational and security monitoring with ITSM and SIEM platforms. Drive network automation and Infrastructure as Code using … Experience with Zero Trust technologies such as Zscaler ZIA/ZPA, Illumio, or equivalent ZTNA platforms in an enterprise deployment context. Familiarity with network observability tools such as ThousandEyes, NetBrain, or Datadog Network Performance Monitoring for advanced path analysis and synthetic monitoring. Experience with ITSM platforms (e.g. ServiceNow or equivalent ...

Azure Cloud SRE — Remote 6-Month Contract, Automation

Location
Greater London, England, United Kingdom
Korn Ferry is hiring a Cloud SRE to support our client’s Azure infrastructure estate, focusing on discovery, planning, and delivery across VMs, networking, storage, and databases. You will monitor service health, drive changes, patching ...

Senior Technical Program Manager (InfoSec), London

Hiring Organisation
Isomorphic Labs
Location
London, UK
Employment Type
Full-time
Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease. The future is coming. A future enabled and enriched by ...

Senior Developer Relations Engineer

Hiring Organisation
Cloudsmith
Location
Belfast, UK
Employment Type
Full-time
TL;DR: We're seeking a technical, community-first Developer Relations Engineer to become the face of Cloudsmith in the open source world, with impact lasting from today until IPO and beyond. About CloudsmithCloudsmith is ...

Staff Observability Engineer: Scalable Profiling & Telemetry

Location
Greater London, England, United Kingdom
Anthropic is seeking Software Engineers to join our Observability team within the Infrastructure group. You will design scalable telemetry pipelines, build instrumentation libraries, and drive AI-assisted diagnostics across multi-cluster systems. The role focuses on end-to-end signals, kernel-level visibility, and cross-team collaboration to improve reliability ...

UK Observability Solutions Engineer - Demos & POCs

Location
Greater London, England, United Kingdom
Dash0 is hiring a Commercial Solutions Engineer in the UK to bridge product and customer outcomes, designing, demonstrating, and validating observability capabilities with OpenTelemetry and cloud-native tech for real-world deployments. You'll work with Sales and Product to craft compelling POCs, deliver demos for engineers and executives ...

Staff Backend Engineer - Grafana App Platform | UK | Remote

Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … fully multi-tenant and scalable, as well as a solid platform for our opinionated Cloud apps. We are turning Grafana into a proper observability app platform where OSS and proprietary apps can directly tap into dashboards, alerts, incidents, and telemetry and deliver even more integrated experiences. To get there ...

Applied AI Engineer New

Location
Greater London, England, United Kingdom
built quickly and as separate systems. The next challenge is to bring these approaches together: build reusable agentic infrastructure, establish a robust evaluation and observability layer, and create systems that allow us to automate new workflows quickly and reliably as Dwelly scales. This is not an AI research role. … loops. Move us from one-off AI solutions toward reusable infrastructure where new workflows can be introduced quickly and with predictable reliability. 2. Evaluation & observability Build the evaluation framework that allows us to understand how our agents perform and why they succeed or fail. Make testing, tracing, debugging, and evaluating ...

AI Engineer III

Hiring Organisation
Victra - Verizon Wireless Premium Retailer
Location
Durham, North Carolina, United States
Employment Type
Permanent
Salary
USD Annual
build the shared platform for agents, including business context, permitted actions, identity and permissions, runtime environments, and auditability. Provide guidance on agent frameworks and observability to support reliable, measurable system performance. Take agents from answering questions to taking actions, defining safe stages for recommendations, actions requiring approval, and autonomous execution. … cost management. Current hands-on coding ability (Python a plus). Fluency with at least one open-source agent framework and one LLM observability stack. Bachelor's degree in computer science, information systems, or equivalent experience. PREFERRED QUALIFICATIONS Experience with AWS Bedrock or another second model platform, and a view ...

Senior AI-Native Product Engineer

Location
Greater London, England, United Kingdom
design and build production systems with Elixir, Phoenix and LiveView, working across domain modelling, data, interfaces, integrations, background processing, testing, deployment and observability as each project requires. You’ll decide what to do yourself, what to delegate to agents and what evidence is needed before the result can be trusted. … initial environment centres on Elixir, Phoenix, LiveView, OTP, Ecto, Ash and PostgreSQL , alongside the usual testing, Git, CI/CD, cloud deployment and observability tooling. Experience in other languages is welcome but not required. Python is particularly useful , whether you’ve used it for services, data, automation, scientific computing ...

Senior Engineer - Agentic Systems

Location
Greater London, England, United Kingdom
changes Ensure work is meaningfully tested and production-ready Design for unreliable APIs, asynchronous processing, partial failure, retries and changing external systems Build appropriate observability and operational controls into the work rather than adding them afterwards Keep changes focused, understandable and maintainable Surface uncertainty early rather than allowing unclear work … tool calling, agent workflows, retrieval systems or AI-enabled software Ability to design systems that handle sensitive business data with appropriate access control, auditability, observability and human oversight Strong product engineering judgement, including the ability to turn broad operational problems into simple, reliable software A pragmatic approach to architecture, testing ...

Staff Site Reliability Engineer

Location
Greater London, England, United Kingdom
multi-trillion-row, petabyte scale — sharding and replication, materialised views, merge and query optimisation, tenant isolation and cost/performance trade-offs. Owning SLOs, observability, capacity planning and incident response for data-intensive systems and pipelines, alongside the orchestration and job execution that power them. Shaping Data-as-a-Service … level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data as a product ...

Staff Site Reliability Engineer

Hiring Organisation
Genomics
Location
Oxford, Oxfordshire, UK
Employment Type
Full-time
multi-trillion-row, petabyte scale — sharding and replication, materialised views, merge and query optimisation, tenant isolation and cost/performance trade-offs. Owning SLOs, observability, capacity planning and incident response for data-intensive systems and pipelines, alongside the orchestration and job execution that power them. Shaping Data-as-a-Service … level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data as a product ...

IT Infrastructure Solutions Architect

Location
Cambridge, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...

IT Infrastructure Solutions Architect

Location
Greater London, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...

Real-Time Banking Platform Engineer

Location
Greater London, England, United Kingdom
GlobalLogic in London is seeking a Software Engineer to join a permanent feature team, delivering software across the lifecycle of a real-time banking platform and collaborating with principal engineers and architects. You will design ...

Data Scientist Technology Data Science London

Location
Greater London, England, United Kingdom
Company Description Checkout.com powers the payments behind many digital experiences. We are where the world checks out, enabling over 10 billion transactions yearly for more than a billion global shoppers. Whether you book a holiday ...

Senior Java Lead: Real-Time Risk & Cloud Solutions (Hybrid)

Location
Greater London, England, United Kingdom
Citi is hiring a Lead Java Developer to advance Real‐Time and On‐Demand risk capabilities within the Credit Business. You will own end‐to‐end delivery from architecture through production support, collaborating with London ...

Senior Analytics Engineer (Hybrid, London, UK)

Location
Greater London, England, United Kingdom
We've signed up to an ambitious journey. Join us! As Arrive, we guide customers and communities towards brighter futures and more livable cities, it isn't a challenge just anyone could take on. Luckily ...

AI Engineering Lead

Location
Greater London, England, United Kingdom
Open roles across robotics, biology, and infrastructure. London-based;some roles remote-friendly. AI Engineering Lead Location London Employment Type Full time Location Type Hybrid Department The opportunity Substrate is building a network of fully ...

Data Scientist

Hiring Organisation
Checkout.com
Location
London, UK
Employment Type
Full-time
Company DescriptionWe'reCheckout.com. You might not know our name, but companies like eBay, Spotify, Klarna, Uber, and Sony do, because we're behind many of the digital experiences you use every day. We are where ...

Hybrid Network Splunk Developer - Telemetry & Observability

Location
Chester, England, United Kingdom
McGregor Recruitment is seeking a Network Splunk Developer to support telemetry, monitoring and observability across a complex enterprise network. The role focuses on collecting telemetry and logs into Splunk, building dashboards and alerts for Network Operations, and collaborating with routing, switching, wireless and SD-WAN teams to ensure uptime. ...

Remote Enterprise Observability Architect

Location
United Kingdom
trusted technical advisor for our largest customers. Based anywhere in the UK, you will design, demonstrate and validate Dash0’s observability platform, bridging product and customer outcomes across pre- and post-sales. Your work will span Proofs of Concept, architecture discussions, and high-impact technical delivery. You’ll collaborate with ...

FP&A Lead

Location
Greater London, England, United Kingdom
ITRS, we make society's critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries - trusted by 90% of Tier 1 capital markets ...

Site Reliability Engineer / Production Support

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
person when incidents occur, escalating only to Head Of when required. Run on-call and incident response; ensure fast detection, triage, and restoration. Maintain observability standards (logs, metrics, traces) and alert quality (low noise, high signal). Understand at a working level all key system flows, the services, partners … first - every manual task is a candidate for automation. You actively hunt for toil and eliminate it. Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source. WHAT YOU BRING Strong SRE or production support experience with accountability for incident response ...