2,876 to 2,900 of 2,960 Remote Observability Jobs

Network Engineer - Network Source of Truth, Automation & Network Reliability

Location
Cambridge, England, United Kingdom
network security fundamentals. Automation Practical capability in Python, Ansible and REST APIs. Experience with Terraform, PowerShell and Git/GitHub is desirable. Monitoring & Observability Experience with network monitoring or observability platforms such as netbox, SolarWinds, Auvik, IP Fabric, LogicMonitor, Dynatrace, Azure Monitor or Grafana. #J-18808-Ljbffr ...

AI Assurance - Senior Consultant

Location
Greater London, England, United Kingdom
Communication - Ability to translate complex technical and regulatory concepts into practical, risk-based recommendations and communicate effectively with technical, operational and executive stakeholdersAI Monitoring & Observability - Strong knowledge of AI monitoring, observability, performance validation, drift detection, incident management and post-deployment assurance practices, including relevant tooling and best practiceDesirable ExperienceWorking across ...

AI Assurance - Senior Consultant

Location
United Kingdom
Communication- Ability to translate complex technical and regulatory concepts into practical, risk-based recommendations and communicate effectively with technical,operationaland executive stakeholders AI Monitoring & Observability- Strong knowledge of AI monitoring, observability, performance validation, drift detection, incidentmanagementand post-deployment assurance practices, including relevant tooling and best practice Desirable Experience Workingacross diverse ...

Observability Software Engineer - Data & Dashboards

Location
Greater London, England, United Kingdom
Vercel is seeking a Software Engineer to join the Observability team in London. You will design, implement, and maintain observability features that help users monitor and understand their applications’ health and performance. The role offers in-office anchor days on Monday, Tuesday, and Friday for those within commuting distance; otherwise ...

Data Engineering Lead

Hiring Organisation
YLD
Location
London, UK
Employment Type
Full-time
pick the right approach for the problemData testing: know what to catch at build time (schema contracts, assertions, transformation logic) versus defer to observability, and can make that call for a teamData observability: treat data reliability like site reliability, with measurable indicators, alerting, incident response, and root cause analysisCI/ ...

platform engineer

Location
Greater London, England, United Kingdom
Design, build, and operate core GenAI platform components, including an LLM routing gateway, vector search and RAG infrastructure, tool registry and MCP gateway, AI observability and evaluation tooling, and infrastructure for long-running agentic workflows; Own production-quality delivery of platform features from design through rollout, monitoring, and follow … Contribute to resilient system design with sensible APIs, failure handling, rate limiting, retries, idempotency, and safe change management; Improve reliability and observability through metrics, dashboards, alerting, incident follow-ups, and operational improvements; Partner with Applied AI Engineers and product teams to understand platform needs and help them build AI-powered ...

Network Engineer Apprentice

Hiring Organisation
QA
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£26,000 per annum
networking, network security and operational support, while developing awareness of modern data centre networking concepts such as EVPN/VXLAN, software-defined networking, automation, observability and high-performance AI fabrics including NVIDIA InfiniBand. Working under the guidance of senior network, infrastructure and security engineers, the apprentice will contribute to improving … modern data centre connectivity models. Responsibilities: Network support and engineering Firewall, security and access control Monitoring, troubleshooting and operational reliability Modern network automation and observability Documentation and standards Collaboration across teams Required skills: Foundational understanding of networking concepts, including IP addressing, subnetting, routing, switching, DNS, DHCP and common network services ...

Senior Java Software Developer

Location
Greater London, England, United Kingdom
ITRS is an Enterprise SaaS provider with industry-leading solutions. Our mission is to make society’s critical technology work via automated & holistic IT observability solutions that safeguard critical applications and enable innovation. With our prestigious customer base includes 90% of the world's top investment banks. We are backed … form part of a wider global Engineering Team. The Core Platform layer is a collection of distributed services which ingest, transform and materialise observability data to make it available to several similarly distributed visualisation, integration, analytics and other domain specific applications to provide solutions to a range of observability problems. ...

Senior QA Engineer

Hiring Organisation
SRG
Location
Warrington, Cheshire, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £50,000 per annum
both fast and reliable. You'll help move quality earlier into the process (shift-left), while also using real production insights to improve decisions (observability-led quality). There's strong scope to influence how QA operates within the team, from testing strategy through to continuous improvement. What … considered early in design and development Leading exploratory testing to uncover issues real users might experience Using monitoring, metrics, and logs to drive observability-led quality and improve production outcomes Identifying risks early and helping the team make informed decisions Improving QA processes, standards, and ways of working across ...

Software Engineer - Platform Productivity | United Kingdom | Remote

Location
United Kingdom
Software Engineer - Platform Productivity | United Kingdom | Remote United Kingdom (Remote) Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud … tools. At the end of the day, we’re the Platform Team for the teams that are building some of the most cherished observability tools– from Grafana, Mimir and Loki, to Tempo. The squad is responsible for setting its own roadmap, and as a part of the team ...

Gen AI Architect

Location
Greater London, England, United Kingdom
production-grade AI systems using Amazon Bedrock, retrieval-augmented generation (RAG), agentic workflows, and cloud-native AWS services. Drive architecture standards, model orchestration, governance, observability, and operational excellence across the GenAI lifecycle while collaborating with engineering, security, compliance, and business stakeholders**Hybrid working:**The places that you work from … customization, prompt orchestration, retrieval pipelines, and agentic workflows* Design agentic AI systems incorporating tool use, workflow orchestration, memory management, and autonomous decision flows* Implement observability for prompts, model responses, vector retrieval quality, and agent execution workflows* Integrate GenAI capabilities into enterprise applications, APIs, workflow platforms, and data ecosystems* Work with ...

Staff ML Engineer | Agentic AI & Applied ML | London (Hybrid) | Contract | Inside IR35

Hiring Organisation
WeDo Technology Solutions Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£800.00 - £1,200.00 per day
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high-risk and high-complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit-learn Experience operating in complex, regulated or high-risk environments would ...

Staff ML Engineer | Agentic AI & Applied ML | London (Hybrid) | Contract | Inside IR35

Location
City Of London, England, United Kingdom
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high-risk and high-complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit-learn Experience operating in complex, regulated or high-risk environments would ...

Software Engineering Lead

Location
Greater London, England, United Kingdom
leadership for one of our teams. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … modern services and legacy systems Shape event-driven and integration patterns so teams can build independently without breaking cross-squad journeys Champion automated testing, observability, and production readiness for backend services - PHPUnit, Go tests, Datadog, and safe release practices Oversee database and schema evolution practices - migrations, data integrity, and performance ...

Engineering Manager

Location
Greater London, England, United Kingdom
designs and code, shape architecture and make decisions when the team needs direction. You’ll also own the team’s production services, including reliability, observability, incidents and on-call, while maintaining a high bar for engineering quality, security and testing. A big part of the role is building and developing … having an interest in AI Experience working in regulated environments where changes require appropriate evidence and controls Strong focus on engineering quality, testing, security, observability and reliability A leadership style built around ownership, urgency and continuous improvement Experience within payments, FX, banking, ledgers, reconciliation, payment compliance, bank integrations or financial ...

Principal Product Engineer

Location
Greater London, England, United Kingdom
typed APIs, clear service and data boundaries, robust processing workflows and platform capabilities that can handle millions of records and events without compromising correctness, observability or operability. A key part of the role is separating the data layer from the application layer. You will help ensure an action taken … propagation separately, ensuring that customer-facing actions can be reversed safely without creating hidden inconsistency underneath. Reliability Engineering : Improve idempotency, retry behaviour, failure isolation, observability, alerting and recovery across critical workflows and integrations. Technical Leadership : Lead design reviews, mentor through code review and pairing, make technical standards explicit, and help ...

Senior Connectivity Engineer

Location
United Kingdom
WireGuard/Tailscale or equivalent): access-as-code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet-wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end-to-end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel-level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling bonus points for containers/IoT OS, ACL-as-code patterns, and carrier/router API integrations. Benefits Starting from the interview process and continuing ...

Senior Connectivity Engineer

Hiring Organisation
Hackajob Ltd
Location
Wallingford, Oxfordshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
WireGuard/Tailscale or equivalent): access-as-code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet-wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end-to-end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel-level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL-as-code patterns, and carrier/router API integrations. Benefits Starting from the interview process and continuing ...

Principal Software Engineer (Platforms)

Location
West of England, England, United Kingdom
business: run evaluations, evidence the benefits, and train teams on effective use. Build the infrastructure that makes agents trustworthy: sandboxing, permissioning, audit logging, and observability so that what an agent did is efficient, secure, and fully transparent to the engineer who has to stand behind it. Create repeatable development environments … experience implementing agentic AI in real engineering workflows (beyond individual experimentation). Experience designing the guardrails around agentic AI: sandboxing, permissioning, audit logging, or observability for tools operating with elevated access. Deep, hands‐on experience of CI/CD, infrastructure as code, containers, and Kubernetes. Experience of building and deploying ...

Software Engineer, Simulation

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components. Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling. Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently. Work with internal users and adjacent engineering teams … such as camera, radar, lidar or GNSS, including modelling uncertainty or noise. Experience integrating machine-learning inference into production systems. Experience with performance profiling, observability or debugging distributed systems. Familiarity with large-scale batch processing, cloud infrastructure or GPU-based workloads. This is a full-time role based ...

Head of Cloud Platform Engineering

Location
Greater London, England, United Kingdom
provision, deploy, and operate services without reinventing infrastructure. Define and improve our service reliability through clear SLOs, recovery objectives, and operational practices. Make observability a core part of every platform capability, with logging, metrics, tracing, and operational runbooks built in from the start. Continuously reduce operational toil through automation, self … based deployment models, with a strong understanding of isolation, resilience, and operational economics. A strong Site Reliability Engineering mindset, including service levels, incident management, observability, and operational excellence. Practical FinOps experience, with a track record of improving cloud cost efficiency. Leading and growing distributed engineering teams across multiple locations, including ...

Senior Software Engineer, Billing & Revenue

Location
Greater London, England, United Kingdom
right. You'll build systems that turn usage and commercial agreements into accurate, auditable invoices, and you'll build the checks and observability that keep them trustworthy as we scale. You'll also work well beyond billing itself, building integrations with other backend systems like our ERP and Data Platform … understand that billing and revenue systems have little room for error, and you design for it — with clear data models, good tests, and observability built in. You enjoy working across teams. You can talk to Finance or Business Development, understand what they need, and turn a commercial model into something ...

Senior Platform Owner - Customer Engagement

Location
United Kingdom
enablement. You’ll guide Value Stream owners and lead cross‐functional technology teams to deliver fast, safe, high ‐ quality change , championing automation‐first practices, observability, decoupling and evidence‐led improvement. You’ll have a track record of improving customer, services, change and team metrics, including DORA metrics, and bring that … deliver fast, safe, high‐quality change, reducing lead times, increasing deployment frequency, and maintaining low change‐failure rates through automation‐first, standards‐led and observability‐driven delivery practices. You know how to turn Dynamics and Power Platform ecosystems into high‐flow, low‐friction environments. You will have deep experience ...

Senior Platform Owner - Customer Engagement

Location
Skipton, England, United Kingdom
enablement. You’ll guide Value Stream owners and lead cross‐functional technology teams to deliver fast, safe, high‐quality change, championing automation‐first practices, observability, decoupling and evidence‐led improvement. You’ll have a track record of improving customer, services, change and team metrics, including DORA metrics, and bring that … deliver fast, safe, high‐quality change, reducing lead times, increasing deployment frequency, and maintaining low change‐failure rates through automation‐first, standards‐led and observability‐driven delivery practices. You know how to turn Dynamics and Power Platform ecosystems into high‐flow, low‐friction environments. You will have deep experience ...

Principal Harness Engineer

Hiring Organisation
FICO
Location
United Kingdom, UK
Employment Type
Full-time
expensive checks post-integration, and continuous sensors that scan for drift outside the change lifecycle - keeping quality as far left as is economical. Improve observability into agent work and track the measures that matter - cost per merged PR, time-to-merge for agent-assisted PRs, review velocity relative … build engineering tooling across a modern stack - linters and static analysis, CI/CD pipelines, containerised build/test environments, and instrumentation/observability - plus familiarity with agent instruction conventions such as AGENTS.md. Experience with spec-driven development, context engineering, agent orchestration, fitness functions, and developer-platform work. A systems ...