2,076 to 2,100 of 2,372 Observability Jobs in London

Head of Engineering

Location
Greater London, England, United Kingdom
down the business. Raise engineering quality Set clear standards for technical design, clean code, testing, code reviews, documentation, and release readiness. Improve automated testing, observability, production stability, and development workflows. Challenge weak technical decisions and help the team build better, more maintainable software. Introduce lightweight and predictable engineering processes without … perform deep code reviews, challenge architectural decisions, and solve complex production issues. Experience working with sensitive financial and personal data. Strong knowledge of testing, observability, incident management and software quality. Experience leading and developing engineering teams. Practical experience using AI across the software development lifecycle, beyond basic code generation. Strong ...

Software Engineer - Workflow Authoring & Deployment

Location
Greater London, England, United Kingdom
that let operators understand what robots are doing and why Build and maintain reliable, secure cloud infrastructure for running these services, including deployment pipelines, observability, and alerting Implement testing and validation pipelines to catch regressions across the UI, backend, and on-robot services - including simulation-based workflow validation before changes … Experience with cloud infrastructure and deployment: containerisation, CI/CD, infrastructure-as-code Comfort working across the stack from frontend to backend Experience with observability and alerting in systems where failures have real operational consequences Good instincts for where complexity should live and how to keep systems debuggable and understandable ...

Senior Director of Software Engineering

Location
Greater London, England, United Kingdom
delivery across the SDLC, ensuring predictable execution and front-office responsiveness. Establish delivery rhythms and drive a production-first culture focused on observability, performance, and operational controls. Lead the design and delivery of integrated workflows, ensuring consistency and traceability across the trade lifecycle. Coordinate cross-team delivery for clean, aligned … experience with front-office UI/tooling (e.g., React/TypeScript) and automation for workflow-critical systems. Experience modernizing legacy estates and improving resilience, observability, and incident management. #J-18808-Ljbffr ...

Senior Platform Architect

Location
Greater London, England, United Kingdom
ensure on-time, high-quality delivery Proactively manage technical risk, migrations, and architectural debt Ensure Long-Term Technical Health Set and validate reliability, observability, and operational standards Monitor health signals and intervene before the customer knows they're happening Lead technical reviews, readiness checkpoints, and maturity assessments Certify teams … Willingness to travel occasionally for customer engagements and team events Nice to Haves Prior experience with Temporal or similar workflow orchestration platforms Experience with observability tools and practices for distributed systems Background in professional services, technical consulting, or solutions engineering roles Experience with Kubernetes, cloud platforms, and infrastructure-as-code ...

IT Operations Director

Hiring Organisation
Bain & Company
Location
London, UK
Employment Type
Full-time
exposures are consistently triaged, remediated, and evidenced. A key part of your role will be advancing the maturity of our technology operations through observability, automation, disciplined incident response, and continuous improvement. You'll identify opportunities for automation and AI-enabled agent workflows to reduce routine operational effort and embed these … reduce unplanned work and strengthen service reliability. Operationalize event-management integrations across platforms such as Datadog, Wiz, CrowdStrike, and ServiceNow. Advance the adoption of observability and operational monitoring practices within day-to-day support activities. Help mature our 24/7 operating model by improving operational readiness and incident response ...

Network Engineer

Hiring Organisation
Cboe Exchange
Location
London, UK
Employment Type
Full-time
leadership, translating complex issues into actionable outcomes. The role also emphasizes building efficient, modern operations through scripting, automation, and API-driven workflows to improve observability, accelerate root cause analysis, and reduce mean time to repair (MTTR). While familiarity with AI-assisted tooling is beneficial, the primary focus … cloud, and infrastructure domains, collaborating with engineering teams and vendors to restore service and resolve complex issues Develop and maintain automation, monitoring enhancements, and observability tooling using scripting and API-driven integrations to improve efficiency, visibility, and troubleshooting speed Perform deep packet analysis and contribute to root cause investigations, translating ...

Senior Sales Engineer - Public Sector (UK)

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscapeWho You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customerSomeone with strong written and oral communication … above may vary based on the country of your employment and the nature of your employment with Datadog. About Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Principal Machine Learning Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
data engineering teams to implement scalable data lakehouse oriented feature architecture and enterprise‐grade ML governance. Champion engineering standards for model quality, documentation, observability, and platform resilience. Feature Engineering & Data ArchitectureArchitect highly scalable, production‐ready feature pipelines within Lakehouse environments. Set the technical direction for fallback and resilience strategies (e.g. … scoring metrics, latency, error analytics, and SLOs. Partner with platform teams to optimise cost, scale, and reliability of inference endpoints. Monitoring, Drift Detection & ObservabilityDefine observability standards for feature drift, concept drift, performance degradation, and data integrity. Lead the creation of dashboards, benchmarks, and automated alerting across the ML ecosystem. Ensure ...

Backend Software Engineer (AI Squad)

Hiring Organisation
Spendesk
Location
London, UK
Employment Type
Full-time
turn ambiguous ideas into concrete backend implementations with measurable impact. Bring pragmatism to delivery, balancing experimentation speed with long-term maintainability and trust. Reliability, observability & operational ownershipYou will: Instrument services with logs, tracing, and metrics to support production visibility and continuous improvement. Define and uphold standards around latency, resilience, failure … integrating predictive models, LLM APIs, or other AI capabilities into product backends. Familiarity with technologies such as Kafka, SQS, Step Functions, PostgreSQL, and modern observability practices. Leadership & collaborationYou are: Highly autonomous and comfortable owning backend systems from design to production. Product-minded, customer-focused, and motivated by building features that ...

Senior Security Engineer: Observability, Automation & Risk

Location
City Of London, England, United Kingdom
seeking a Staff Security Engineer to advance security observability, automate controls and strengthen governance across its global tech estate. This senior individual contributor role partners with engineering, data, AI and digital workplace teams to raise security maturity in a fast-moving environment. You will design scalable security observability, automate assurance ...

Senior/Principal Product Manager - Storage

Location
Greater London, England, United Kingdom
integration. Work with customers and internal teams to understand workload requirements around capacity, throughput, latency, resilience, and cost. Drive integrations with compute, Kubernetes, IAM, observability, and developer tooling. Own prioritisation, delivery, launch, adoption, and continuous iteration. Define success metrics across performance, reliability, utilisation, adoption, and unit economics. Qualifications: Experience owning … replication, caching Protocols & interfaces: S3, NFS, CSI, NVMe/NVMe-oF, APIs Infrastructure: Kubernetes, containers, networking, cloud platforms Reliability: durability, availability, consistency, failure recovery Observability & security: metrics, logging, IAM, encryption, auditability Familiarity with technologies such as Weka, Vast, DDN, MinIO, Lustre, Pure Storage, or similar. Able to reason about trade ...

Senior Software Engineer - (Java/AI) Brokerage Services

Location
City Of London, England, United Kingdom
Services team to push hard on AI across Brokerage Services engineering: AI-assisted coding, testing and deployments, shift left on security, defect triage and observability, and the internal tooling that connects our AI tools safely to Jira, GitLab, and our CI/CD pipeline. This is a hands‐on role. … maintain internal tooling that embeds AI into day‐to‐day engineering work: coding assistants, AI‐assisted test generation, deployments, and defect triage. Contribute to observability in the AI‐assisted pipeline, so every stage shows what the tooling is doing, catching, and missing. Identify manual steps in coding, testing, and release ...

Team Leader - Data Engineering & Integration - Commodities Data

Hiring Organisation
Bloomberg
Location
London, UK
Employment Type
Full-time
direction, balancing modernization, production stability, partner needs, and business-as-usual delivery. Build team capability in data pipelines, integration patterns, Python, SQL, orchestration, automation, observability, controls, and production support. Partner with Product, Engineering, Data Modelling, Data Quality, Content Acquisition, Enablement, and regional teams to deliver business-aligned outcomes. Contribute … move teams toward scalable, automated, supportable, and well-controlled operating models. Strong working knowledge of Python, SQL, orchestration tools, workflow platforms, automation frameworks, observability, and production support practices. Experience embedding data quality controls, reconciliation, completeness checks, timeliness checks, and exception workflows into production processes. Ability to work closely with data ...

Team Leader - Data Engineering & Integration - Commodities Data

Hiring Organisation
Bloomberg
Location
London, United Kingdom
Salary
£ 70 K
technical direction, balancing modernization, production stability, partner needs, and business-as-usual delivery.Build team capability in data pipelines, integration patterns, Python, SQL, orchestration, automation, observability, controls, and production support.Partner with Product, Engineering, Data Modelling, Data Quality, Content Acquisition, Enablement, and regional teams to deliver business-aligned outcomes.Contribute to global Commodities … workflows and move teams toward scalable, automated, supportable, and well-controlled operating models.Strong working knowledge of Python, SQL, orchestration tools, workflow platforms, automation frameworks, observability, and production support practices.Experience embedding data quality controls, reconciliation, completeness checks, timeliness checks, and exception workflows into production processes.Ability to work closely with data modelling ...

Senior Product Manager, Payments

Location
Greater London, England, United Kingdom
touchpoints.The role will own and optimise Collinson's internal payment systems while managing key external partnerships with PSPs, acquirers, payment orchestration, fraud prevention and observability providers. In addition, the role will oversee payment risk and fraud management, ensuring regulatory compliance and enhancing payment security.Leading a high-performing product team … streamline global payment routing, retries and conversion optimisation.* Integrate with fraud prevention providers, implementing real-time risk assessment and fraud mitigation tools.* Work with observability partners to ensure real-time monitoring, reporting and payment analytics for proactive issue resolution.* Oversee payment security, fraud prevention and risk mitigation strategies across ...

Frontend Engineering Associate Manager (GenAI experience)

Location
Greater London, England, United Kingdom
including architecture, engineering standards, and governance Raise the bar for web quality and reliability by championing testing, accessibility, performance (Core Web Vitals) and production observability Identify, prototype and scale Generative AI use cases in products and delivery workflows, ensuring responsible adoption and measurable value Grow and retain high-performing teams …/component libraries (governance, documentation, and adoption) Frontend build tooling and developer experience (e.g., Vite/Webpack, monorepos, linting/formatting, dependency management) Production observability and debugging (frontend monitoring, error tracking, and analytics instrumentation) Frontend security fundamentals (e.g., XSS/CSRF, secure authentication flows, and secure handling of client-side ...

Senior Developer Experience Platform Engineer

Location
Greater London, England, United Kingdom
come in. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster with … metrics, monitoring, and alerting for web platforms, with a strong intuition for actionable vs noisy signals. Experience working with platform-level security controls and observability, including identifying and responding to abnormal or unsafe system behaviour. Experience designing and operating continuous delivery systems that support safe, repeatable, zero-downtime deployments. Experience ...

Senior Software Engineer - (Java/AI) Client Platform

Location
City Of London, England, United Kingdom
hard on AI across our React Native and UI engineering practice: AI-assisted coding, testing and deployments, shift left on security, defect triage and observability, and the internal tooling that connects our AI tools safely to Jira, GitLab, and our CI/CD pipeline. You'll bring strong React Native … maintain internal tooling that embeds AI into day-to-day engineering work: coding assistants, AI-assisted test generation, deployments, and defect triage. Contribute to observability in the AI-assisted pipeline, so every stage shows what the tooling is doing, catching, and missing. Identify manual steps in coding, testing, and release ...

Staff Software Engineer - Backend

Hiring Organisation
DataBricks
Location
London, UK
Employment Type
Full-time
London presence. Below are some example teams you can join: Lakebase Platform Reliability: Constantly evolving Lakebase Reliability. That includes cross-functional infrastructure, observability, SLI/SLOs definition and measurements, as well as achieving higher operational efficiency without the need for humans-in-the-loop (both internal and customer-facing … API. We route and manage large volumes of concurrent connections, solving load balancing, authentication, and data access at high throughput. We focus on performance, observability, and reliability to ensure the data path is fast and resilient. The impact you will have: Our backend teams span many domains across our essential ...

Senior AI Engineer

Location
Greater London, England, United Kingdom
cutting‐edge techniques with pragmatism to deliver measurable impact. Apply strong software engineering principles, such as modularity, testing, code reviews, CI/CD and observability, to ensure AI systems are reliable, maintainable, production‐ready and can be readily adapted to future developments. Choose the right approach for the problem … engineers, and PMs, to scope and ship AI features iteratively. Ability to reason about system behavior end‐to‐end, including model performance, latency, and observability, and how these impact user experience. Clear, structured communicator, comfortable documenting and defending architectural decisions and engaging in thoughtful technical debate. Not required ...

Senior Analytics Engineer

Hiring Organisation
Sony Interactive Entertainment America
Location
London, UK
Employment Type
Full-time
evolve the analytics layer architecture, ensuring it scales with data volume and business complexity. Ensure data quality and trust through testing, documentation, and observability within dbt. Business PartnershipPartner closely with business stakeholders, analysts, and product teams to understand requirements and translate them into well‐designed data models. … technically robust and intuitive for business users. Experience building greenfield analytics or data platforms, or significantly evolving existing ones. Experience with data quality and observability, using dbt or complementary tooling. Familiarity with orchestration tools such as Airflow, Dagster, or Prefect. Experience with semantic layers, metrics frameworks, or analytics consumption patterns. ...

Senior AI Engineer

Location
Greater London, England, United Kingdom
cutting‐edge techniques with pragmatism to deliver measurable impact. Apply strong software engineering principles, such as modularity, testing, code reviews, CI/CD and observability, to ensure AI systems are reliable, maintainable, production‐ready and can be readily adapted to future developments. Choose the right approach for the problem … engineers, and PMs, to scope and ship AI features iteratively Ability to reason about system behavior end‐to‐end, including model performance, latency, and observability, and how these impact user experience. Clear, structured communicator, comfortable documenting and defending architectural decisions and engaging in thoughtful technical debate. Not required ...

VP, Head of Data Science

Location
Greater London, England, United Kingdom
effective. Scale practical AI adoption Design and implement AI workflows as skills, models and MCP tools in OpenWebUI, with cost control, provider routing and observability built in. Partner with other departments to turn AI capabilities into everyday workflows, reusable tools and knowledge layers that improve internal productivity and client delivery. … open‐source AI software, including LiteLLM and OpenWebUI, to maximise the internal value of AI and build durable institutional capability. Reliability, cost control, observability, instruction‐following, usability and adoption matter most. The team designs reusable workflows, skills, models and MCP tools that support Penta’s move towards an AI‐native ...

Staff Engineer

Location
Greater London, England, United Kingdom
mobile technologies, build proof‐of‐concepts, and create roadmaps for adoption. Operational Excellence: Champion the "you build it, you run it" philosophy - driving pipeline observability, zero‐downtime deployments, robust monitoring, and alerting for front‐end applications. Resilience & Performance: Design solution‐wide UI resilience (graceful degradation, back‐pressure, retry strategies … understanding of SOLID principles, TDD, DDD, and software design patterns . Experience designing for resilience and performance at scale (graceful degradation, zero‐downtime deployments, observability, performance budgets). Experience with Azure development and cloud‐based infrastructure , including CI/CD and pipeline observability. Experience owning threat modelling and security hardening ...

Senior Infrastructure Operations Engineer (EMEA) New London, England, United Kingdom

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … scripting skills using Python, Go, Bash, Ansible, or similar tools. Experience with Kubernetes, Slurm, or other cluster and workload orchestration systems. Experience using monitoring, observability, and telemetry systems to diagnose and troubleshoot production infrastructure. Strong systems and networking fundamentals, with a track record of owning production issues through resolution. Comfortable ...