1,476 to 1,500 of 2,124 Remote/Hybrid Observability Jobs

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Bedford, Bedfordshire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Halifax, West Yorkshire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Senior or Staff Software Engineer, SRE/ Platform Team

Location
Greater London, England, United Kingdom
code with Kubernetes and Terraform. You'll be at the forefront of shaping our foundational architecture, ensuring it’s both resilient and scalable. Drive Observability and Monitoring: Establish and maintain a state‐of‐the‐art observability and monitoring stack. Your insights will enable us to stay ahead of potential issues ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Durham, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
High Wycombe, Buckinghamshire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Wareham, Dorset, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Preston, Lancashire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Church End, Bedfordshire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Scarborough, North Yorkshire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Computershare
Location
Bristol, Gloucestershire, United Kingdom
Salary
£ 70 K
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … services while fostering a culture of shared ownership and continuous learning.Other key responsibilities:Drive adoption of SRE principles (SLOs, error budgets, toil reduction).Establish observability and monitoring standards.Lead automation-first operations.Improve incident and problem management maturity.Partner with software and infrastructure engineering teams to embed reliability into the product lifecycle.Establish ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Computershare
Location
Bristol, UK
Employment Type
Full-time
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … while fostering a culture of shared ownership and continuous learning. Other key responsibilities: Drive adoption of SRE principles (SLOs, error budgets, toil reduction).Establish observability and monitoring standards. Lead automation-first operations. Improve incident and problem management maturity. Partner with software and infrastructure engineering teams to embed reliability into ...

Head of Site Reliability Engineering (SRE)

Location
West of England, England, United Kingdom
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … fostering a culture of shared ownership and continuous learning. Other key responsibilities: Drive adoption of SRE principles (SLOs, error budgets, toil reduction). Establish observability and monitoring standards. Lead automation-first operations. Improve incident and problem management maturity. Partner with software and infrastructure engineering teams to embed reliability into ...

Intelligent Automation Engineering Manager

Location
Greater London, England, United Kingdom
deployment and operational support of AI agents and AI-powered solutions. Establish engineering standards and best practices for AI architecture, orchestration, retrieval, tool invocation, observability, governance, privacy, security and cost management. Review technical designs and architecture documentation to ensure solutions align with engineering, security and governance standards. Translate emerging … with MCP (Model Context Protocol), MCP Servers, MCP Clients or enterprise AI integration frameworks. Knowledge of LLMOps, AI evaluation frameworks, model routing and AI observability tooling. Experience with enterprise integration technologies including APIs, Middleware, ESB or iPaaS platforms. Hands‐on experience integrating internal and third-party systems. Azure cloud experience ...

Software Development Engineer in Test

Location
Greater London, England, United Kingdom
accuracy, reliability, latency, and cost. Work with Engineering, Product, Data Science, QA, and Operations teams to deliver production-ready solutions. Design for scalability, reliability, observability, and operational excellence across the platform. Troubleshoot complex production issues across application services, data pipelines, AI workflows, and telemetry. Participate in architecture, technical design, code … C++. Strong experience working with large-scale data, telemetry, logs, events, and operational datasets. Strong understanding of cloud platforms, system scalability, reliability, performance, observability, and production operations. Strong analytical, problem-solving, communication, and collaboration skills, with the ability to work effectively across technical and cross-functional teams. Desirable skills ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
environments Define and implement infrastructure-as-code using Terraform, CDK, or equivalent Monitor platform health, define SLOs/SLAs, and build alerting and observability tooling Respond to and lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security … Actions, Jenkins, or equivalent) Solid understanding of networking, security groups, load balancing, and DNS Experience with container orchestration (Docker, Kubernetes, or ECS) Familiarity with observability tooling (Datadog, Grafana, CloudWatch, or equivalent) Understanding of HIPAA infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site ...

Advanced Engineer, Investment Technology

Location
United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

AI Consulting –Sr. AI Architect & Client Partner

Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Platform Engineer

Location
Greater London, England, United Kingdom
scalable solutions to thousands of customers every day. Our mission is to make deploying and operating software effortless and safe. We focus on automation, observability, and reliability, ensuring every engineering team at Clio can move faster and with confidence. You’ll be part of a globally distributed team, collaborating closely … Code using Terraform, ensuring reproducibility and compliance. Collaborate with developers to improve CI/CD pipelines, deployment strategies, and overall developer experience. Enhance observability and reliability, refining alerting, monitoring, and incident response. Collaborate on cloud optimization projects, improving performance, cost efficiency, and security posture. Mentor and guide team members, fostering ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
accelerate development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing … encourage applications without these:Hands-on with Helm or KustomizeExperience with GitOps (e.g., Argo CD)Knowledge of secrets management (e.g., HashiCorp Vault)Experience with observability (metrics/logs/tracing)CollabHiringWhy Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Senior .NET Backend Developer

Location
York and North Yorkshire, England, United Kingdom
design, implementation, testing, review, refactoring, documentation and migration, while retaining clear ownership of the outcome. Set a high standard for maintainability, automated testing, security, observability and pull-request review. Investigate complex technical and production issues and drive them through to resolution. What we are looking for Strong experience designing … clinically important data. PostgreSQL, Redis, Elasticsearch or other data and caching technologies. GraphQL, including schema design and gateway patterns. Grafana, OpenTelemetry, Prometheus or equivalent observability tooling. CI/CD, production services, microservices and message-driven systems such as RabbitMQ. How we work We value engineers who own outcomes and remain ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Value Stream Engineering Lead

Location
Greater London, England, United Kingdom
fostering a culture of collaboration and continuous improvement Own and continuously improve engineering standards and practices across code quality, testing, CI/CD, observability, performance, DevOps, automation and secure, resilient system design Partner closely with Product and Architecture to optimise the flow of value from idea to delivery, identifying blockers … coaching engineers through technical challenges in real time Strong ownership of modern engineering standards and practices, including code quality, testing strategy, CI/CD, observability, performance optimisation, DevOps and secure, resilient system design Experience designing and delivering within modern cloud environments such as AWS, Azure or GCP, with a strong ...