2,826 to 2,850 of 2,902 Remote Observability Jobs

Software Engineer - Platform Productivity | United Kingdom | Remote

Location
United Kingdom
Software Engineer - Platform Productivity | United Kingdom | Remote United Kingdom (Remote) Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud … tools. At the end of the day, we’re the Platform Team for the teams that are building some of the most cherished observability tools– from Grafana, Mimir and Loki, to Tempo. The squad is responsible for setting its own roadmap, and as a part of the team ...

Gen AI Architect

Location
Greater London, England, United Kingdom
production-grade AI systems using Amazon Bedrock, retrieval-augmented generation (RAG), agentic workflows, and cloud-native AWS services. Drive architecture standards, model orchestration, governance, observability, and operational excellence across the GenAI lifecycle while collaborating with engineering, security, compliance, and business stakeholders**Hybrid working:**The places that you work from … customization, prompt orchestration, retrieval pipelines, and agentic workflows* Design agentic AI systems incorporating tool use, workflow orchestration, memory management, and autonomous decision flows* Implement observability for prompts, model responses, vector retrieval quality, and agent execution workflows* Integrate GenAI capabilities into enterprise applications, APIs, workflow platforms, and data ecosystems* Work with ...

Staff ML Engineer | Agentic AI & Applied ML | London (Hybrid) | Contract | Inside IR35

Hiring Organisation
WeDo Technology Solutions Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£800.00 - £1,200.00 per day
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high-risk and high-complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit-learn Experience operating in complex, regulated or high-risk environments would ...

Staff ML Engineer | Agentic AI & Applied ML | London (Hybrid) | Contract | Inside IR35

Location
City Of London, England, United Kingdom
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high-risk and high-complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit-learn Experience operating in complex, regulated or high-risk environments would ...

Software Engineering Lead

Location
Greater London, England, United Kingdom
leadership for one of our teams. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … modern services and legacy systems Shape event-driven and integration patterns so teams can build independently without breaking cross-squad journeys Champion automated testing, observability, and production readiness for backend services - PHPUnit, Go tests, Datadog, and safe release practices Oversee database and schema evolution practices - migrations, data integrity, and performance ...

Engineering Manager

Location
Greater London, England, United Kingdom
designs and code, shape architecture and make decisions when the team needs direction. You’ll also own the team’s production services, including reliability, observability, incidents and on-call, while maintaining a high bar for engineering quality, security and testing. A big part of the role is building and developing … having an interest in AI Experience working in regulated environments where changes require appropriate evidence and controls Strong focus on engineering quality, testing, security, observability and reliability A leadership style built around ownership, urgency and continuous improvement Experience within payments, FX, banking, ledgers, reconciliation, payment compliance, bank integrations or financial ...

Principal Product Engineer

Location
Greater London, England, United Kingdom
typed APIs, clear service and data boundaries, robust processing workflows and platform capabilities that can handle millions of records and events without compromising correctness, observability or operability. A key part of the role is separating the data layer from the application layer. You will help ensure an action taken … propagation separately, ensuring that customer-facing actions can be reversed safely without creating hidden inconsistency underneath. Reliability Engineering : Improve idempotency, retry behaviour, failure isolation, observability, alerting and recovery across critical workflows and integrations. Technical Leadership : Lead design reviews, mentor through code review and pairing, make technical standards explicit, and help ...

Senior Connectivity Engineer

Location
United Kingdom
WireGuard/Tailscale or equivalent): access-as-code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet-wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end-to-end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel-level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling bonus points for containers/IoT OS, ACL-as-code patterns, and carrier/router API integrations. Benefits Starting from the interview process and continuing ...

Senior Connectivity Engineer

Hiring Organisation
Hackajob Ltd
Location
Wallingford, Oxfordshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
WireGuard/Tailscale or equivalent): access-as-code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet-wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end-to-end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel-level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL-as-code patterns, and carrier/router API integrations. Benefits Starting from the interview process and continuing ...

Principal Software Engineer (Platforms)

Location
West of England, England, United Kingdom
business: run evaluations, evidence the benefits, and train teams on effective use. Build the infrastructure that makes agents trustworthy: sandboxing, permissioning, audit logging, and observability so that what an agent did is efficient, secure, and fully transparent to the engineer who has to stand behind it. Create repeatable development environments … experience implementing agentic AI in real engineering workflows (beyond individual experimentation). Experience designing the guardrails around agentic AI: sandboxing, permissioning, audit logging, or observability for tools operating with elevated access. Deep, hands‐on experience of CI/CD, infrastructure as code, containers, and Kubernetes. Experience of building and deploying ...

Software Engineer, Simulation

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components. Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling. Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently. Work with internal users and adjacent engineering teams … such as camera, radar, lidar or GNSS, including modelling uncertainty or noise. Experience integrating machine-learning inference into production systems. Experience with performance profiling, observability or debugging distributed systems. Familiarity with large-scale batch processing, cloud infrastructure or GPU-based workloads. This is a full-time role based ...

Head of Cloud Platform Engineering

Location
Greater London, England, United Kingdom
provision, deploy, and operate services without reinventing infrastructure. Define and improve our service reliability through clear SLOs, recovery objectives, and operational practices. Make observability a core part of every platform capability, with logging, metrics, tracing, and operational runbooks built in from the start. Continuously reduce operational toil through automation, self … based deployment models, with a strong understanding of isolation, resilience, and operational economics. A strong Site Reliability Engineering mindset, including service levels, incident management, observability, and operational excellence. Practical FinOps experience, with a track record of improving cloud cost efficiency. Leading and growing distributed engineering teams across multiple locations, including ...

Senior Software Engineer, Billing & Revenue

Location
Greater London, England, United Kingdom
right. You'll build systems that turn usage and commercial agreements into accurate, auditable invoices, and you'll build the checks and observability that keep them trustworthy as we scale. You'll also work well beyond billing itself, building integrations with other backend systems like our ERP and Data Platform … understand that billing and revenue systems have little room for error, and you design for it — with clear data models, good tests, and observability built in. You enjoy working across teams. You can talk to Finance or Business Development, understand what they need, and turn a commercial model into something ...

Senior Platform Owner - Customer Engagement

Location
United Kingdom
enablement. You’ll guide Value Stream owners and lead cross‐functional technology teams to deliver fast, safe, high ‐ quality change , championing automation‐first practices, observability, decoupling and evidence‐led improvement. You’ll have a track record of improving customer, services, change and team metrics, including DORA metrics, and bring that … deliver fast, safe, high‐quality change, reducing lead times, increasing deployment frequency, and maintaining low change‐failure rates through automation‐first, standards‐led and observability‐driven delivery practices. You know how to turn Dynamics and Power Platform ecosystems into high‐flow, low‐friction environments. You will have deep experience ...

Senior Platform Owner - Customer Engagement

Location
Skipton, England, United Kingdom
enablement. You’ll guide Value Stream owners and lead cross‐functional technology teams to deliver fast, safe, high‐quality change, championing automation‐first practices, observability, decoupling and evidence‐led improvement. You’ll have a track record of improving customer, services, change and team metrics, including DORA metrics, and bring that … deliver fast, safe, high‐quality change, reducing lead times, increasing deployment frequency, and maintaining low change‐failure rates through automation‐first, standards‐led and observability‐driven delivery practices. You know how to turn Dynamics and Power Platform ecosystems into high‐flow, low‐friction environments. You will have deep experience ...

Principal Harness Engineer

Hiring Organisation
FICO
Location
United Kingdom, UK
Employment Type
Full-time
expensive checks post-integration, and continuous sensors that scan for drift outside the change lifecycle - keeping quality as far left as is economical. Improve observability into agent work and track the measures that matter - cost per merged PR, time-to-merge for agent-assisted PRs, review velocity relative … build engineering tooling across a modern stack - linters and static analysis, CI/CD pipelines, containerised build/test environments, and instrumentation/observability - plus familiarity with agent instruction conventions such as AGENTS.md. Experience with spec-driven development, context engineering, agent orchestration, fitness functions, and developer-platform work. A systems ...

Senior Specialist Engineer - Networks

Hiring Organisation
UK Health Security Agency
Location
Birmingham, Leeds, Liverpool or London (Canary Wharf), E14 4PU, United Kingdom
Salary
£56185.00 to £70566.00
mitigation strategies to reduce risk, improve service resilience and support secure cloud and on-premises environments. Service Management, Automation & Technical Leadership Lead the monitoring, observability, performance and capacity management of network services, integrating operational and security monitoring with ITSM and SIEM platforms. Drive network automation and Infrastructure as Code using … Experience with Zero Trust technologies such as Zscaler ZIA/ZPA, Illumio, or equivalent ZTNA platforms in an enterprise deployment context. Familiarity with network observability tools such as ThousandEyes, NetBrain, or Datadog Network Performance Monitoring for advanced path analysis and synthetic monitoring. Experience with ITSM platforms (e.g. ServiceNow or equivalent ...

Azure Cloud SRE — Remote 6-Month Contract, Automation

Location
Greater London, England, United Kingdom
Korn Ferry is hiring a Cloud SRE to support our client’s Azure infrastructure estate, focusing on discovery, planning, and delivery across VMs, networking, storage, and databases. You will monitor service health, drive changes, patching ...

Senior Technical Program Manager (InfoSec), London

Hiring Organisation
Isomorphic Labs
Location
London, UK
Employment Type
Full-time
Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease. The future is coming. A future enabled and enriched by ...

Senior Developer Relations Engineer

Hiring Organisation
Cloudsmith
Location
Belfast, UK
Employment Type
Full-time
TL;DR: We're seeking a technical, community-first Developer Relations Engineer to become the face of Cloudsmith in the open source world, with impact lasting from today until IPO and beyond. About CloudsmithCloudsmith is ...

Staff Observability Engineer: Scalable Profiling & Telemetry

Location
Greater London, England, United Kingdom
Anthropic is seeking Software Engineers to join our Observability team within the Infrastructure group. You will design scalable telemetry pipelines, build instrumentation libraries, and drive AI-assisted diagnostics across multi-cluster systems. The role focuses on end-to-end signals, kernel-level visibility, and cross-team collaboration to improve reliability ...

UK Observability Solutions Engineer - Demos & POCs

Location
Greater London, England, United Kingdom
Dash0 is hiring a Commercial Solutions Engineer in the UK to bridge product and customer outcomes, designing, demonstrating, and validating observability capabilities with OpenTelemetry and cloud-native tech for real-world deployments. You'll work with Sales and Product to craft compelling POCs, deliver demos for engineers and executives ...

Staff Backend Engineer - Grafana App Platform | UK | Remote

Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … fully multi-tenant and scalable, as well as a solid platform for our opinionated Cloud apps. We are turning Grafana into a proper observability app platform where OSS and proprietary apps can directly tap into dashboards, alerts, incidents, and telemetry and deliver even more integrated experiences. To get there ...

Applied AI Engineer New

Location
Greater London, England, United Kingdom
built quickly and as separate systems. The next challenge is to bring these approaches together: build reusable agentic infrastructure, establish a robust evaluation and observability layer, and create systems that allow us to automate new workflows quickly and reliably as Dwelly scales. This is not an AI research role. … loops. Move us from one-off AI solutions toward reusable infrastructure where new workflows can be introduced quickly and with predictable reliability. 2. Evaluation & observability Build the evaluation framework that allows us to understand how our agents perform and why they succeed or fail. Make testing, tracing, debugging, and evaluating ...

AI Engineer III

Hiring Organisation
Victra - Verizon Wireless Premium Retailer
Location
Durham, North Carolina, United States
Employment Type
Permanent
Salary
USD Annual
build the shared platform for agents, including business context, permitted actions, identity and permissions, runtime environments, and auditability. Provide guidance on agent frameworks and observability to support reliable, measurable system performance. Take agents from answering questions to taking actions, defining safe stages for recommendations, actions requiring approval, and autonomous execution. … cost management. Current hands-on coding ability (Python a plus). Fluency with at least one open-source agent framework and one LLM observability stack. Bachelor's degree in computer science, information systems, or equivalent experience. PREFERRED QUALIFICATIONS Experience with AWS Bedrock or another second model platform, and a view ...