576 to 600 of 850 Remote Observability Jobs

Platform Chapter Lead - Engineering

Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Senior Software Engineer

Hiring Organisation
Hargreaves Lansdown
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Part Time
applications using modern engineering practices. Develop integrations with internal and external systems. Enhance platform capabilities for managing and maintaining customer data. Improve platform resilience, observability, performance and scalability. Contribute across the full stack, including React-based applications where required. Take ownership of features throughout the software development lifecycle, from design … schema design. Experience migrating from REST to GraphQL architectures. Typesense, Elasticsearch or Contentful. OAuth and authentication flows. LaunchDarkly or feature management platforms. Grafana and observability tooling. Cross-account AWS architectures. TanStack Query, Zustand, Vite, Tailwind CSS. Micro-frontend architectures. AI-assisted and agentic development workflows. Financial Services experience. Interview process ...

Software Engineering Manager

Location
Corby, England, United Kingdom
Collaboration: Collaborate effectively with agile teams, domains, architects, the engineering practice, scrum masters, and product owners to drive integrated solutions. Operational Excellence: Coordinate the observability and operational support of production infrastructure, ensuring system reliability and performance. Quality Assurance: Review engineering output for quality, completeness, and adherence to best practices. Production …/CD pipelines using tools such as Terraform and GitLab CI. Knowledge of cloud platform tooling including HashiCorp Vault, Nomad, Kong API Gateway, and observability platforms such as Datadog. Solid understanding of application security, performance optimisation, caching technologies (e.g. Redis), and test automation. Demonstrated knowledge of solution architecture, emerging technologies ...

Lead Java Developer

Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc. Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

DevOps Engineer

Location
Reigate, England, United Kingdom
infrastructure. Automate environment provisioning across development and production. Manage backend state, pipelines, and state-change detection integrations. Platform Engineering & SRE Own and improve reliability, observability, and performance of the platform. Implement SLOs, alerting, dashboards, and auto remediation where possible. Troubleshoot cluster level, networking, and workload deployment issues. Lead root cause … endpoints, Certificate/Secret management etc Strong debugging and operational experience (SRE mindset). Solid experience of DevSecOps architecture, processes & tooling Solid understanding of Observability Process & Tooling Logging, metrics, traces, dashboards Other highly desirable, but not essential skills are: Experience with: GitOps - ArgoCD or GitOps workflows Zero downtime deployments (blue ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services.You will work with technology, engineering, operations … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices including ...

AI Platform & Site Reliability Engineering Managing Consultant

Location
United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services. You will work with technology, engineering, operations … operational requirements. AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. Reliability Engineering & SRE: Establish SRE practices including ...

Lead AI Software Engineer – Hybrid Delivery Leader

Location
Newbury, England, United Kingdom
secure, production-ready code while guiding engineers and AI agents. You''ll combine deep software engineering with AI-native practices, ensuring code quality, testing, observability, and compliance with standards. The role is hybrid, located in Newbury or Paddington, and offers the opportunity to shape #J-18808-Ljbffr ...

Backend Engineer, EMEA — Remote & AI-Driven Ownership

Location
United Kingdom
work. You will implement and ship backend features using Ruby/Rails, Go, or other modern languages, design APIs, and ensure maintainability, performance, and observability across services. #J-18808-Ljbffr ...

Remote Applied AI Engineer: Agentic Systems & Automation

Location
Greater London, England, United Kingdom
Applied AI Engineer to build agentic infrastructure and production-grade AI systems powering automation across its real estate platform. You will focus on reliability, observability, and scalable workflows to push Dwelly's AI capabilities into core business processes. As an early core member, you will influence architectural decisions, developer tooling ...

Remote Backend Engineer - Multi-Tenant App Platform

Location
United Kingdom
Grafana Labs, the leader in open observability, is hiring for a cloud/Grafana-as-a-service role. You will develop backend features in Golang, improve system reliability, and influence the product roadmap while collaborating across teams. The role supports a fully remote setup with compensation guided by UK market ...

Hybrid Site Reliability Engineer — Healthcare Infra & Automation

Location
Greater London, England, United Kingdom
Cranial Technologies is seeking a Site Reliability Engineer to help build and maintain scalable, secure healthcare technology systems. The role focuses on automation, observability, and reducing manual operations to keep clinical and business processes reliable. You will support production apps, APIs, and integrations, while collaborating with IT, security, and clinical ...

Senior Backend Engineer - Remote or Hybrid, Warehouse Data

Location
Cambridge, England, United Kingdom
contracts that power the product — ingestion pipelines transforming warehouse data into a model, robust APIs, durable storage, and reliable background work. You will ensure observability that lets a small team operate them confidently at 3 am. Two engines power WareBee: Physical AI and Process AI. You will build systems ...

Remote Backend Engineer - Django, Python, Energy Tech

Location
Greater London, England, United Kingdom
Django backends, collaborate with hardware and product teams, and ship features across multiple countries. You’ll own medium-sized features, write tests, and improve observability while integrating with Kraken platforms and external services. The role offers hybrid work from London and UK-wide opportunities, with a strong focus on green ...

Elasticsearch Consultant

Hiring Organisation
Identifi Global Resources Ltd
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Contract, Work From Home
their Elastic consulting practice. We have multiple opportunities for experienced Elastic Consultants/Senior Elastic Data Engineers to work on complex Search, Security and Observability programmes across Government and enterprise environments. Rather than working in a traditional engineering or operational support role, youll be part of an experienced consulting team ...

Principal Software Architect

Hiring Organisation
Spectrum IT Recruitment
Location
Uxbridge, London, United Kingdom
Employment Type
Permanent
Salary
£100000 - £120000/annum
service boundaries, ownership models and integration patterns Reviewing significant technical initiatives and providing architectural guidance across multiple engineering teams Embedding security, resilience, scalability and observability into platform design from the outset Identifying architectural risk, technical debt and platform constraints Partnering with Engineering and Product leadership on strategic technical decisions Influencing ...

Enterprise Account Executive - UKI

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
above may vary based on the country of your employment and the nature of your employment with Datadog. About Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Strategic Account Executive (UKI) - Public Sector

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
above may vary based on the country of your employment and the nature of your employment with Datadog.#LI-HybridAbout Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Field Marketing Manager (UKI)

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
above may vary based on the country of your employment and the nature of your employment with Datadog. About Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Manchester, England, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Glasgow, Scotland, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Greater London, England, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Newcastle upon Tyne, England, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

ServiceNow CRM Transformation Lead - Managing Consultant

Hiring Organisation
Capgemini
Location
Cheshire West and Chester, United Kingdom
Employment Type
Full Time
realisation Translate business requirements into workflow-enabled operating models Solution Design & Architecture Design and support implementation of: AI Control Tower (AI lifecycle management, governance, observability) Agentic AI workflows enabling autonomous execution Now Assist/GenAI use cases across workflows Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

Senior Manager, AI Architect

Location
Greater London, England, United Kingdom
responsible-AI and governance controls aligned toemergingregulation and client risk appetites. Define human-in-the-loop and fallback strategies for high-stakes use cases. Observability & operations Architect logging, monitoring,tracingand cost-observability across the AI stack, including model,agentand platform telemetry. Design for drift detection, performancemonitoringand continuous evaluation in production ...