1,926 to 1,950 of 2,213 Observability Jobs in London

Infrastructure Engineer, Db2

Hiring Organisation
NatWest Group
Location
London, UK
Employment Type
Full-time
Join us as an Infrastructure Engineer, Db2You'll engineer Db2 related infrastructure technology complying with security, resilience, sustainability, and operational requirements with observability and guardrails built inYou'll also use automation to provide testing and a route to live for the product, working with customers to help them … design and implement robust housekeeping, backup, and recovery strategiesWorking knowledge of the development of CI or CD pipelines using modern toolingExperience of using observability tools and techniques and the ability to use data, information and user sentiment to continuously improve solutionsExperience of working with technology deployed to an on premise ...

Backend Software Engineer – Infrastructure, Foundations

Location
Greater London, England, United Kingdom
compute workloads and efficiently schedule hundreds of thousands of containers hourly Design architecture and opinionated APIs that guide application developers Implement tracing and performance observability in high-scale distributed microservice architectures Build reliable, performant, and scalable systems for storage, authentication, and asset serving Automate deployment, management, and operations of distributed … keywords Distributed Systems Development Java Programming C++ Programming Python Programming Cloud Infrastructure Management ATS Optimization Keywords Hard Skills Software Engineering Data Processing Systems Performance Observability Microservice Architecture API Design Container Scheduling Open-Source Contribution High-Scale Systems Automation Prototyping Soft Skills Strong Communication Skills Team Collaboration Feedback Incorporation Quality Maintenance ...

Data Engineer

Location
Greater London, England, United Kingdom
/CD and infrastructure as code. Create reusable components and maintain clear technical documentation. Quality & Governance (10%) : Implement robust data validation, testing, lineage and observability to ensure high-quality, trusted datasets. Support governance and privacy-conscious data handling. Collaboration & Enablement (10%) : Partner with Data Science, MLOps, Product and commercial teams … cloud environments (preferably AWS) Engineering Best Practice: Knowledge of CI/CD, testing, version control and infrastructure as code Data Quality & Governance: Understanding of observability, validation and maintaining reliable data systems Collaboration & Communication: Ability to translate business and data science needs into scalable solutions and communicate clearly with stakeholders Mindset ...

Integration Architect - SAP SAAS Products

Location
Greater London, England, United Kingdom
4HANA Cloud, SuccessFactors, Ariba, Concur, Datasphere SAC) and SAP BTP. Define patterns, govern APIs/events, ensure secure, resilient data flows, and drive standardization, observability, and compliance. Core Responsibilities Architecture & Standards • Define canonical integration patterns (API-led, event-driven, batch/EDI) and reference architectures on SAP BTP. • Establish guidelines … SuccessFactors, Ariba, Concur, Datasphere, SAC, and S/4HANA Cloud integrations. • Strong security fundamentals (OAuth2, SAML, JWT, SCIM) and compliance awareness. • Hands-on with observability (Cloud ALM), performance tuning, and reliability engineering. • Experience with agile delivery, CI/CD (Git-based pipelines), and test automation for integrations. • Preferred Qualifications ...

Lead Cloud Network Engineer AWS - Hedge Fund

Location
Greater London, England, United Kingdom
network security controls, monitoring and architecture. You'll also design and maintain secure hybrid connectivity between AWS, Azure and on‐premises datacentres, with strong observability, monitoring and capacity planning across the wider network estate. Infrastructure as Code and automation will be central to the role. You'll build reusable Terraform …/CD experience, provisioning, pipelines, network tooling You have a good knowledge of Routing and Switching (BGP), VPNs (IPSec, SSL), Network Monitoring and Observability You're collaborative and pragmatic, able to push back where necessary What's in it for you: Competitive salary, to £140k Pension Private medical care ...

Staff Software Engineer - Customer Data Platform

Location
City of Westminster, England, United Kingdom
outcomes. Work with adjacent teams when needed to align on shared components and dependencies. Actively participate in on-call support, contributing to operational stability, observability, and performance. Coach and support other engineers through pairing, reviews, and mentoring, helping raise the team’s overall capability. Contribute to team OKRs and actively … that balance speed, maintainability, and long-term scalability. You actively identify technical debt, risks, or inefficiencies and take action to address them. You use observability and metrics to validate behavior, debug issues, and improve system health. You regularly support and unblock teammates, helping them deliver more effectively. You contribute ...

Platform Engineer, AI Enablement

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
language models and AI agents. You'll help create a governed model-access layer and a secure production environment for agentic workflows, with safety, observability, and operational excellence built in from the start. Your work will span the full lifecycle—from defining problems and designing systems to implementation, deployment … cost-effective access to AI models and tools. Build the runtime, services, and developer tooling used to run agentic workflows in production. Create observability across cost, performance, reliability, and usage through metrics, tracing, and logging. Implement governance and safety controls that make AI use secure, compliant, and auditable. Develop reusable ...

Software Engineer

Location
Greater London, England, United Kingdom
system. AI harnesses in production. Deploying agentic AI into real-world operational settings — acting on real money, tenancies and legal exposure, with the guardrails, observability and correctness that demands. Non-deterministic LLM working within compliant, secure deterministic software. Our stack We build on NestJS + TypeScript on GCP/… build Small, well-factored services. TDD and DDD as defaults. Trunk-based CI/CD — you ship to production and own it, with tests, observability and clean rollbacks. Lean frameworks, readable code, and we move fast because the tests and boundaries let us. What we're looking for #J ...

Platform Engineer, AI Enablement London, United Kingdom

Location
Greater London, England, United Kingdom
language models and AI agents. You’ll help create a governed model-access layer and a secure production environment for agentic workflows, with safety, observability, and operational excellence built in from the start. Your work will span the full lifecycle—from defining problems and designing systems to implementation, deployment … cost-effective access to AI models and tools. Build the runtime, services, and developer tooling used to run agentic workflows in production. Create observability across cost, performance, reliability, and usage through metrics, tracing, and logging. Implement governance and safety controls that make AI use secure, compliant, and auditable. Develop reusable ...

AI Engineer - Agentic AI & Back End

Location
Greater London, England, United Kingdom
Pydantic AI, LangGraph/LangChain or Microsoft agent platformsDefine evaluation strategies for agentic systems using automated and human-in-the-loop methodsImplement telemetry and observability solutions to monitor performance and qualityLeverage AI-assisted tools such as Claude Code, OpenAI Codex or GitHub Copilot to accelerate developmentRequirementsBachelor’s or Master … based solutions with quality evaluation mechanismsUnderstanding of evaluation methodologies for AI systems such as golden datasets and automated testsProficiency in telemetry, logging and observability tools for production systemsAbility to work in cross-functional client engagement teams delivering enterprise solutionsNice to haveExperience using AI development tools such as Claude Code, OpenAI ...

Software Development Engineer II - BDP

Location
Greater London, England, United Kingdom
backend event-driven platform using Java and Spring Boot Pairing with more senior engineers to design, implement, test, and ship code Learning to use observability tools like New Relic and Splunk to monitor live systems Participating in planning sessions and team discussions to understand requirements and contribute ideas Writing automated … feedback, and share what you're learning Nice to have: Exposure to Spring Boot, NoSQL databases, or cloud services Curiosity about performance, scalability, and observability in large-scale systems Familiarity with Git, CI/CD pipelines, and containerisation Some experience working with platforms like Kafka You might know ...

Lead AI Engineer

Hiring Organisation
Capco
Location
London, UK
Employment Type
Full-time
experience deploying LLMs and multi-modal models at scaleStrong engineering background in Python with proven backend and API development skillsSolid understanding of scalable MLOps, observability, and cloud-native AI deploymentExcellent communication, problem-solving, and project management skills in agile environmentsBonus Points ForExperience with agentic frameworks (e.g., LangChain, LlamaIndex)Experience … deep learning frameworks and front-end developmentFamiliarity with Langfuse, Langsmith, or other LLM observability toolsUnderstanding of Model Context Protocol and bias/hallucination mitigation techniquesPrevious success in integrating GenAI solutions into enterprise-scale systemsWhy Join CapcoDeliver high-impact technology solutions for Tier 1 financial institutionsWork in a collaborative, flat ...

Head of Engineering

Location
Greater London, England, United Kingdom
down the business. Raise engineering quality Set clear standards for technical design, clean code, testing, code reviews, documentation, and release readiness. Improve automated testing, observability, production stability, and development workflows. Challenge weak technical decisions and help the team build better, more maintainable software. Introduce lightweight and predictable engineering processes without … perform deep code reviews, challenge architectural decisions, and solve complex production issues. Experience working with sensitive financial and personal data. Strong knowledge of testing, observability, incident management and software quality. Experience leading and developing engineering teams. Practical experience using AI across the software development lifecycle, beyond basic code generation. Strong ...

Software Engineer - Workflow Authoring & Deployment

Location
Greater London, England, United Kingdom
that let operators understand what robots are doing and why Build and maintain reliable, secure cloud infrastructure for running these services, including deployment pipelines, observability, and alerting Implement testing and validation pipelines to catch regressions across the UI, backend, and on-robot services - including simulation-based workflow validation before changes … Experience with cloud infrastructure and deployment: containerisation, CI/CD, infrastructure-as-code Comfort working across the stack from frontend to backend Experience with observability and alerting in systems where failures have real operational consequences Good instincts for where complexity should live and how to keep systems debuggable and understandable ...

Senior Director of Software Engineering

Location
Greater London, England, United Kingdom
delivery across the SDLC, ensuring predictable execution and front-office responsiveness. Establish delivery rhythms and drive a production-first culture focused on observability, performance, and operational controls. Lead the design and delivery of integrated workflows, ensuring consistency and traceability across the trade lifecycle. Coordinate cross-team delivery for clean, aligned … experience with front-office UI/tooling (e.g., React/TypeScript) and automation for workflow-critical systems. Experience modernizing legacy estates and improving resilience, observability, and incident management. #J-18808-Ljbffr ...

Senior Platform Architect

Location
Greater London, England, United Kingdom
ensure on-time, high-quality delivery Proactively manage technical risk, migrations, and architectural debt Ensure Long-Term Technical Health Set and validate reliability, observability, and operational standards Monitor health signals and intervene before the customer knows they're happening Lead technical reviews, readiness checkpoints, and maturity assessments Certify teams … Willingness to travel occasionally for customer engagements and team events Nice to Haves Prior experience with Temporal or similar workflow orchestration platforms Experience with observability tools and practices for distributed systems Background in professional services, technical consulting, or solutions engineering roles Experience with Kubernetes, cloud platforms, and infrastructure-as-code ...

IT Operations Director

Hiring Organisation
Bain & Company
Location
London, UK
Employment Type
Full-time
exposures are consistently triaged, remediated, and evidenced. A key part of your role will be advancing the maturity of our technology operations through observability, automation, disciplined incident response, and continuous improvement. You'll identify opportunities for automation and AI-enabled agent workflows to reduce routine operational effort and embed these … reduce unplanned work and strengthen service reliability. Operationalize event-management integrations across platforms such as Datadog, Wiz, CrowdStrike, and ServiceNow. Advance the adoption of observability and operational monitoring practices within day-to-day support activities. Help mature our 24/7 operating model by improving operational readiness and incident response ...

Network Engineer

Hiring Organisation
Cboe Exchange
Location
London, UK
Employment Type
Full-time
leadership, translating complex issues into actionable outcomes. The role also emphasizes building efficient, modern operations through scripting, automation, and API-driven workflows to improve observability, accelerate root cause analysis, and reduce mean time to repair (MTTR). While familiarity with AI-assisted tooling is beneficial, the primary focus … cloud, and infrastructure domains, collaborating with engineering teams and vendors to restore service and resolve complex issues Develop and maintain automation, monitoring enhancements, and observability tooling using scripting and API-driven integrations to improve efficiency, visibility, and troubleshooting speed Perform deep packet analysis and contribute to root cause investigations, translating ...

Senior Sales Engineer - Public Sector (UK)

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscapeWho You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customerSomeone with strong written and oral communication … above may vary based on the country of your employment and the nature of your employment with Datadog. About Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Principal Machine Learning Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
data engineering teams to implement scalable data lakehouse oriented feature architecture and enterprise‐grade ML governance. Champion engineering standards for model quality, documentation, observability, and platform resilience. Feature Engineering & Data ArchitectureArchitect highly scalable, production‐ready feature pipelines within Lakehouse environments. Set the technical direction for fallback and resilience strategies (e.g. … scoring metrics, latency, error analytics, and SLOs. Partner with platform teams to optimise cost, scale, and reliability of inference endpoints. Monitoring, Drift Detection & ObservabilityDefine observability standards for feature drift, concept drift, performance degradation, and data integrity. Lead the creation of dashboards, benchmarks, and automated alerting across the ML ecosystem. Ensure ...

Backend Software Engineer (AI Squad)

Hiring Organisation
Spendesk
Location
London, UK
Employment Type
Full-time
turn ambiguous ideas into concrete backend implementations with measurable impact. Bring pragmatism to delivery, balancing experimentation speed with long-term maintainability and trust. Reliability, observability & operational ownershipYou will: Instrument services with logs, tracing, and metrics to support production visibility and continuous improvement. Define and uphold standards around latency, resilience, failure … integrating predictive models, LLM APIs, or other AI capabilities into product backends. Familiarity with technologies such as Kafka, SQS, Step Functions, PostgreSQL, and modern observability practices. Leadership & collaborationYou are: Highly autonomous and comfortable owning backend systems from design to production. Product-minded, customer-focused, and motivated by building features that ...

Senior Security Engineer: Observability, Automation & Risk

Location
City Of London, England, United Kingdom
seeking a Staff Security Engineer to advance security observability, automate controls and strengthen governance across its global tech estate. This senior individual contributor role partners with engineering, data, AI and digital workplace teams to raise security maturity in a fast-moving environment. You will design scalable security observability, automate assurance ...

Senior/Principal Product Manager - Storage

Location
Greater London, England, United Kingdom
integration. Work with customers and internal teams to understand workload requirements around capacity, throughput, latency, resilience, and cost. Drive integrations with compute, Kubernetes, IAM, observability, and developer tooling. Own prioritisation, delivery, launch, adoption, and continuous iteration. Define success metrics across performance, reliability, utilisation, adoption, and unit economics. Qualifications: Experience owning … replication, caching Protocols & interfaces: S3, NFS, CSI, NVMe/NVMe-oF, APIs Infrastructure: Kubernetes, containers, networking, cloud platforms Reliability: durability, availability, consistency, failure recovery Observability & security: metrics, logging, IAM, encryption, auditability Familiarity with technologies such as Weka, Vast, DDN, MinIO, Lustre, Pure Storage, or similar. Able to reason about trade ...

Team Leader - Data Engineering & Integration - Commodities Data

Hiring Organisation
Bloomberg
Location
London, UK
Employment Type
Full-time
direction, balancing modernization, production stability, partner needs, and business-as-usual delivery. Build team capability in data pipelines, integration patterns, Python, SQL, orchestration, automation, observability, controls, and production support. Partner with Product, Engineering, Data Modelling, Data Quality, Content Acquisition, Enablement, and regional teams to deliver business-aligned outcomes. Contribute … move teams toward scalable, automated, supportable, and well-controlled operating models. Strong working knowledge of Python, SQL, orchestration tools, workflow platforms, automation frameworks, observability, and production support practices. Experience embedding data quality controls, reconciliation, completeness checks, timeliness checks, and exception workflows into production processes. Ability to work closely with data ...

Senior Product Manager, Payments

Location
Greater London, England, United Kingdom
touchpoints.The role will own and optimise Collinson's internal payment systems while managing key external partnerships with PSPs, acquirers, payment orchestration, fraud prevention and observability providers. In addition, the role will oversee payment risk and fraud management, ensuring regulatory compliance and enhancing payment security.Leading a high-performing product team … streamline global payment routing, retries and conversion optimisation.* Integrate with fraud prevention providers, implementing real-time risk assessment and fraud mitigation tools.* Work with observability partners to ensure real-time monitoring, reporting and payment analytics for proactive issue resolution.* Oversee payment security, fraud prevention and risk mitigation strategies across ...