1,626 to 1,650 of 2,378 Observability Jobs in London

Staff Frontend Engineer

Location
Greater London, England, United Kingdom
TypeScript and modern React (hooks, component composition, transitions), and of frameworks like Next.js that handle SSR, routing, and data fetching. Deep experience with observability, Core Web Vitals, accessibility auditing, and frontend performance tuning - and the ability to evaluate new frameworks and identify where we gain an edge in development, deployment ...

Enterprise Sales Engineer, UK

Hiring Organisation
Rubrik
Location
London, UK
Employment Type
Full-time
innovative technical programs and overseeing day-to-day account-level activities. You will be responsible for evangelizing, positioning, and architecting Rubrik's Cyber Resilience, Observability, and Remediation tools to a targeted list of new & existing customers throughout the UKI region. What You'll Do: Provides technical leadership and direction ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Greater London, England, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

Solution Engineer

Location
Greater London, England, United Kingdom
## Solution EngineerLondon, UK · Full-time · Senior#### About The PositionCoralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs, metrics, trace … security events with features such as APM, RUM, SIEM, Kubernetes monitoring and more, all enhancing operational efficiency and reducing observability spend by up to 70%.Solution Architects in Coralogix are key in meeting our customers’ expectations and helping them utilize their observability and security data. We are looking for hard ...

Data Observability Engineer

Hiring Organisation
Ashdown Group
Location
London, UK
Employment Type
Full-time
successful multinational technology business is looking for a Data Observability Engineer to join its growing data team in Central London. This role is hybrid – you'll be able to work from home 2 days per week. This is a high-impact role focused on improving data quality, reducing incidents … building scalable observability across a modern enterprise data platform. You'll help ensure data across the organisation is accurate, reliable, and trusted for critical business decision-making. You'll take ownership of data reliability end-to-end, designing and implementing frameworks that monitor data health, detect anomalies, and enforce standards ...

Senior Manager, AI Architect

Location
Greater London, England, United Kingdom
responsible-AI and governance controls aligned toemergingregulation and client risk appetites. Define human-in-the-loop and fallback strategies for high-stakes use cases. Observability & operations Architect logging, monitoring,tracingand cost-observability across the AI stack, including model,agentand platform telemetry. Design for drift detection, performancemonitoringand continuous evaluation in production ...

AI Native SW Eng

Location
Greater London, England, United Kingdom
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … Vertex AI, and open-source models - with fallback routing, token, cost, and latency management ImplementLLMOpsin production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions - workshops, proofs ...

AI Native SW Eng

Location
Greater London, England, United Kingdom
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … open-source models — with fallback routing, token, cost, and latency management Implement LLMOps in production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions — workshops, proofs of concept ...

AI Native SW Engineering

Location
Greater London, England, United Kingdom
production‐grade agentic systems at enterprise scale: multi‐agent orchestration across complex environments, RAG pipelines, policy‐based routing, memory management, and programme‐level lifecycle observability Define RAG pipeline standards across engagements: establish chunking and embedding strategies, set quality benchmarks, and ensure metric‐backed tradeoff decisions are documented and transferable … standard design practice across providers including OpenAI, Anthropic, Vertex AI, and open‐source models Own LLMOps at programme scale: eval strategy, prompt governance, observability tooling standards, safety monitoring and cost controls across multiple concurrent systems Lead client engineering engagements at senior level — facilitate architecture design sessions, lead proof‐of‐concept ...

AWS Engineering Lead

Hiring Organisation
Interact Consulting Limited
Location
South West London, London, United Kingdom
Employment Type
Permanent
architecture, reliability, security and FinOps. Set standards for Terraform, OpenTofu and Infrastructure as Code. Own GitHub orchestration, CI/CD and release practices. Improve observability, incident management and operational readiness. Select the right tools and methods for sustainable engineering. Mentor engineers and drive continuous improvement. Partner with engineering, security … Cloud or Platform Engineering. Strong AWS, CI/CD, security and production operations experience. Expertise in Terraform or OpenTofu. Knowledge of GitHub Actions, Docker, observability and monorepo tooling. A pragmatic approach to architecture, standards and technology selection. A passion for mentoring teams and building reliable, secure platforms. Please Apply Now. ...

Principal AI Platform Engineer

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
golden paths, and standard service templates to simplify service provisioning and operations. Contribute to cloud-native platform architecture, including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Integrate AI-assisted engineering workflows to support faster delivery, improved code quality, automation, and data-driven operational decisions. … Establish governance, security, and compliance guardrails through policy-as-code and auditable platform patterns. Improve platform reliability using SLOs, observability practices, resilience engineering, and insights from incidents. Collaborate with product, engineering, security, and architecture teams to align platform capabilities with business priorities and user needs. Drive efficiency and sustainability through ...

Principal AI Platform Engineer

Location
Greater London, England, United Kingdom
golden paths, and standard service templates to simplify service provisioning and operations. Contribute to cloud-native platform architecture, including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Integrate AI-assisted engineering workflows to support faster delivery, improved code quality, automation, and data-driven operational decisions. … Establish governance, security, and compliance guardrails through policy-as-code and auditable platform patterns. Improve platform reliability using SLOs, observability practices, resilience engineering, and insights from incidents. Collaborate with product, engineering, security, and architecture teams to align platform capabilities with business priorities and user needs. Drive efficiency and sustainability through ...

Core Platform Developer

Location
Greater London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment.**Key Responsibilities*** Design and build shared backend services, frameworks … developer tooling that support internal application and service development.* Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities.* Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams.* Help ...

AWS Cloud Architect

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Permanent
platform architecture and modernisation roadmap, including migration from a Java monolith to microservices on EKS. Define standards for containers, runtime environments, observability, tenancy, security, and infrastructure automation. Lead SRE practices including SLI/SLOs, incident management, DR/BCP planning, post-mortems, and operational resilience. Own platform security, secure SDLC … networking, KMS, RDS, and multi-account architecture. Hands-on Kubernetes, CI/CD, Terraform, and cloud security experience. Strong understanding of SRE, observability, incident response, and disaster recovery. Experience operating within regulated environments such as ISO 27001, SOC 2, or GxP. Comfortable balancing strategic leadership with hands-on operational delivery. ...

Enterprise Architect

Location
Greater London, England, United Kingdom
observable platforms. The role requires strong hands‐on architecture depth across application modernization, integration architecture, microservices, APIs, event streaming, container platforms, database modernization, DevSecOps, observability, resiliency, and governance. You will collaborate with business, application, data, infrastructure, security, DevOps, operations, and delivery teams to define transition roadmaps, manage architectural risks, establish … Application Architecture Microservices Architecture Event-Driven Architecture API & Integration Architecture Cloud-Agnostic Architecture Domain-Driven Design Strangler Pattern Migration DevSecOps CI/CD & IaC Observability DataBricks Data Governance & Lineage Cloud Portability Kubernetes Kafka .NET MSSQL/PostgreSQL Nice to Have Financial Services/Banking Domain experience Key Competencies Proven enterprise ...

Senior Python AI Developer

Hiring Organisation
EMBS Engineering
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 60,000 - 75,000 Annual
Designing CI/CD pipelines, automation and reusable microservices Supporting LLM benchmarking, performance monitoring and cost optimisation Integrating AI services with enterprise security and observability frameworks Collaborating with architects and platform engineers on scalable AI solutions What we're looking for You'll be an experienced Python Developer … AgentCore RESTful APIs, microservices and distributed systems CI/CD pipelines, DevOps and automation LLM integration, evaluation and performance optimisation Secure application development and observability Experience with Terraform, GitOps, Kong API Gateway, SQL/NoSQL databases and AI evaluation tools such as Promptfoo or Arize would also be useful. Previous ...

Senior Python AI Developer

Hiring Organisation
EMBS Engineering
Location
London, United Kingdom
Employment Type
Permanent, Contract
Salary
£60000 - £75000/annum + Benefits
Designing CI/CD pipelines, automation and reusable microservices Supporting LLM benchmarking, performance monitoring and cost optimisation Integrating AI services with enterprise security and observability frameworks Collaborating with architects and platform engineers on scalable AI solutions What we're looking for You'll be an experienced Python Developer … AgentCore RESTful APIs, microservices and distributed systems CI/CD pipelines, DevOps and automation LLM integration, evaluation and performance optimisation Secure application development and observability Experience with Terraform, GitOps, Kong API Gateway, SQL/NoSQL databases and AI evaluation tools such as Promptfoo or Arize would also be useful. Previous ...

AI / Machine Learning Engineer – Agentic LLM Systems (Contract)

Location
Greater London, England, United Kingdom
services and APIs to support AI-driven applications and workflowsBuild and run experiments to improve reliability, latency, cost, and success ratesContribute to evaluation frameworks, observability, and monitoring of LLM/agent performance What We’re Looking For: 4+ years’ experience in ML/AI engineering (LLMs, recommender systems, optimisation … codeComfortable debugging complex, distributed AI/ML systemsExperience running and analysing large-scale experiments and performance metrics (latency, accuracy, cost)Exposure to monitoring/observability tools for production systems Nice to Have: Experience with multi-agent systems or distributed AI architecturesExperience optimising LLM usage across multiple providers (cost/performance ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
pipelines Supporting the deployment of applications into production Working with Azure Kubernetes Service (AKS) and containerised applications Improving cloud networking, security, resilience and observability Working with Entra ID and identity/access management Partnering closely with Software Engineers to make deployments simpler, safer and more reliable Helping shape platform standards … Cloud networking Deploying and supporting applications in production Experience with Kubernetes/AKS would be particularly useful. Azure DevOps, Entra ID, cloud security, observability, cost optimisation or relevant Azure/Kubernetes certifications would all add value, but you don't need to tick every box. More important is that ...

Network Service Manager

Location
Greater London, England, United Kingdom
transition to a NetDevOps-focused operating model. You'll work at the intersection of service management and engineering, ensuring that automation, observability, and modern delivery practices are embedded into how we run and improve the network. You'll partner closely with the Principal Network Engineer, Principal Network Architect, and Delivery … Champion the adoption of NetDevOps practices within service operations — promoting automation, self-healing infrastructure, and reduced manual toil. Partner with network engineers to embed observability, telemetry, and automated alerting into the operational framework. Drive the use of infrastructure-as-code and CI/CD pipelines for network changes, reducing risk ...

Senior AI Software Engineer

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
Software Engineer will take ownership of software components and services across their operational lifecycle, contributing to engineering excellence through code review, automated testing, observability, documentation, technical collaboration and continuous improvement. What you will doBuild, test and maintain APIs, services, integrations, customer journeys and platform capabilities. Translate assigned requirements, specifications … capabilities. Familiar with cloud-native architectures, distributed systems, containerised deployment models and event-driven integration patterns. Experienced with CI/CD pipelines, automated testing, observability and production support. Able to apply AI coding tools and AI-assisted development approaches responsibly within software delivery. Skilled in reviewing both AI-generated ...

Senior Platform Engineer (Azure)

Location
Greater London, England, United Kingdom
improve CI/CD pipelines and deployment processes Work closely with engineering, product and QA teams to support reliable delivery Improve platform reliability, security, observability and performance Help drive DevSecOps, SRE and automation best practices across the platform What they're looking for: Strong cloud-native Azure experience Production experience … haves... AZ-305 or CKA certifications Helm Azure DevOps pipelines Experience within a regulated environment such as financial services, insurance or healthcare Experience with observability, SRE and DevSecOps principles This role would suit a Senior Platform Engineer who still enjoys being hands on technically and wants real ownership over ...

Senior Staff AI Engineering

Hiring Organisation
American Express
Location
London, UK
Employment Type
Full-time
systems with significant strategic and long-term business impactDefines architecture and engineering patterns for agentic AI, retrieval/grounding, model integration, inference platforms, observability, evaluation, and safetyLeads resolution of complex and ambiguous technical challenges, balancing performance, cost, scalability, security, safety, and regulatory constraintsShapes long-term AI engineering strategy across platforms … planning, reasoning, tool use, memory, multi-agent coordination, and autonomy controlsStrong foundation in distributed systems, cloud-native architecture, Kubernetes, APIs, microservices, event-driven design, observability, and platform reliabilityKnowledge of DevOps and CI/CD platforms such as Harness, including automated deployment, environment promotion, governance, and operational controlsKnowledge of enterprise ...

Senior Cloud Network & Security Engineer , Vice President

Location
Greater London, England, United Kingdom
platform engineering capabilities that improve consistency, reliability, and speed of delivery. Develop AI-enabled operational capabilities that support network investigation, incident response, change validation, observability, and workflow automation. Define architecture standards, engineering patterns, and technology roadmaps for enterprise-wide adoption. Partner with cloud, security, infrastructure, application, and operations teams globally. … platforms. Exposure to platform engineering, GitOps, event-driven automation, or self-service infrastructure models. Interest or experience in AI-enabled operations, AI agent workflows, observability platforms, knowledge systems, or intelligent automation. Experience working in a regulated, global enterprise environment. What We Value Engineering ownership and accountability. A builder’s mindset ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
have deep expertise in only one or more areas, you will be able to contribute across the entire ecosystem. Our platform domains include: Observability: Monitoring, tracing, logging, and alerting infrastructure. Scheduling & Orchestration: HPC schedulers and container orchestration for compute-intensive workloads. Development Tools: CI/CD pipelines, version control, artifact … troubleshooting (Linux, Bash, Containerization).Infrastructure automation and configuration management experience (e.g., Ansible, Terraform, Puppet). Nice to have: Deep knowledge of a modern observability stack (e.g., Prometheus, Grafana, Elastic Stack, Vector, AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g. ...