2,076 to 2,100 of 2,283 Observability Jobs in London

Principal Product Engineer

Location
Greater London, England, United Kingdom
typed APIs, clear service and data boundaries, robust processing workflows and platform capabilities that can handle millions of records and events without compromising correctness, observability or operability. A key part of the role is separating the data layer from the application layer. You will help ensure an action taken … propagation separately, ensuring that customer-facing actions can be reversed safely without creating hidden inconsistency underneath. Reliability Engineering : Improve idempotency, retry behaviour, failure isolation, observability, alerting and recovery across critical workflows and integrations. Technical Leadership : Lead design reviews, mentor through code review and pairing, make technical standards explicit, and help ...

Software Engineer - Fleet Management & Repair Automation

Location
Greater London, England, United Kingdom
between competing operations, rate limiting, graceful cancellation, and emergency stop. Autonomy: replacing human-driven operational decisions with automation and agentic workflows, backed by the observability and quality signals needed to trust them. Practically, this means we build large-scale workflow orchestration engines: systems that model operational intent, schedule it against … maintaining production systems in one or more of Python, C++, Java, Rust or Go Track record of operating critical systems: oncall ownership, incident leadership, observability and SLO design, and structural reliability improvement Ability to lead ambiguous, cross-organisational work to completion, and to communicate clearly in writing to both engineers ...

Product Engineer (Backend/AI @Briink)

Location
Greater London, England, United Kingdom
that support multiple user journeys and reporting workflows rather than solving each problem in isolation Improve the robustness of our systems, strengthening reliability, testing, observability, performance, and maintainability as we scale Work directly with users, product, and data, using qualitative feedback and product data to understand the real problem … combined with the ability and willingness to become productive in Python quickly) A solid software-quality mindset, including testing, code review, CI/CD, observability, and pragmatic approaches to reliability and security Experience in a small, high-ownership product team, ideally within a startup or scale-up engineering organisation ...

Platform Software Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
exception pathsModelling infrastructure state and making changes idempotent, attributable and auditableDefining stable contracts between storage platform and developer platform servicesBuilding tests, fixtures and observability that make failures safe to diagnoseWorking with users and partner teams to turn operational problems into bounded engineering workParticipating in incident response … infrastructure automation, workflows or reconciliation systemsUnderstanding of retries, partial failure, concurrency, idempotency and eventual consistencyStrong Linux and production troubleshooting skillsExperience with CI/CD, observability and safe deployment practicesA practical approach to security, permissions, audit and change controlA calm, methodical approach to incidentsDesirable experience includes orchestration, configuration management, enterprise identity ...

Engineering Director

Location
Greater London, England, United Kingdom
Structuring technical problem-solving frameworks and coordinating incident response to resolve complex, interdependent system failures and maintain system reliability Guiding the development of comprehensive observability, monitoring, and alerting strategies in order to ensure system health, visibility, and operational excellence Driving structured delivery governance, CI/CD practices, and automated pipelines … environment (such as healthcare, life sciences, or financial services), including information security, data protection, and quality management obligations Proven success developing and implementing comprehensive observability, monitoring, and delivery governance strategies aligned with business objectives In-depth knowledge of how to navigate ambiguous technical and organisational environments, structuring decision-making frameworks ...

Senior Product Manager - Storage & Networking

Location
Greater London, England, United Kingdom
underpinning Radiant’s GPU platform. You’ll work across bare-metal GPU clusters, Kubernetes, high-performance storage, data‐centre networking, infrastructure inventory, automation, and observability, partnering closely with engineering, SRE, infrastructure, networking, and operations. This is a highly technical product role focused on ensuring GPU workloads have reliable, high‐throughput … infrastructure through to customer and workload connectivity. Partner with engineering to turn requirements into scalable, reliable platform services. Drive improvements in automation, self‐service, observability, and operational efficiency. Define success metrics and use them to guide performance, reliability, capacity, and investment decisions. Qualifications: Experience owning technical infrastructure, platform, storage, networking ...

Model Release Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
with teams across Wayve to understand their requirements, agree interfaces and resolve technical or delivery conflicts across shared workflows. Improve the reliability, scalability and observability of the platform through effective monitoring, alerting and operational tooling. Provide clear visibility of model candidates, their progress, evaluation results, approvals and release status. … work through conflicting priorities. Experience operating cloud-based services using Kubernetes, with a good understanding of reliability, scalability and performance. Practical knowledge of observability, monitoring and alerting, including defining meaningful service-health metrics. Confidence using AI coding tools and agents to improve engineering productivity and automate repeatable work. A pragmatic ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
Anaplan’s platform and third‐party integrations Optimise model inference pipelines for performance, cost, and scalability in production environments Implement monitoring, logging, and observability for GenAI systems to track usage, errors, and model behaviour Collaborate with data scientists to productionise ML models and forecasting algorithms Your Skills Extensive hands … Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with model observability tools (LangSmith, W&B;, MLflow) Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive ...

Model Release Engineer

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
with teams across Wayve to understand their requirements, agree interfaces and resolve technical or delivery conflicts across shared workflows. Improve the reliability, scalability and observability of the platform through effective monitoring, alerting and operational tooling. Provide clear visibility of model candidates, their progress, evaluation results, approvals and release status. … work through conflicting priorities. Experience operating cloud-based services using Kubernetes, with a good understanding of reliability, scalability and performance. Practical knowledge of observability, monitoring and alerting, including defining meaningful service-health metrics. Confidence using AI coding tools and agents to improve engineering productivity and automate repeatable work. A pragmatic ...

Partner Sales Manager - EMEA

Location
Greater London, England, United Kingdom
March Capital, Lightspeed, Sorenson Ventures, Industry Ventures, and Emergent Ventures, we are a Series-C funded company headquartered in Silicon Valley. Our Enterprise Data Observability Platform - the first of its kind - helps enterprises build and operate world-class data products by ensuring data is reliable, trusted, and ready to power … strategy with hyperscaler priorities and customer modernization initiatives. Product and Industry Expertise and Demonstration Maintain a strong understanding of Acceldata's platform, including data observability use cases across modern and legacy data architectures. Confidently articulate Acceldata's value to partner sales, technical, and executive audiences. Support and, when needed, deliver ...

Solutions Architect

Location
Greater London, England, United Kingdom
Maintain comprehensive documentation of design decisions, patterns, standards, and trade-offs. Define non-functional platform architecture standards covering resilience, backup/restore, disaster recovery, observability, service levels, auditability and operational readiness. Define platform security and privacy architecture in partnership with Cyber Security and Data Governance, including PII handling, access recertification … authority in environments transitioning from outsourced to in-house ownership is beneficial. Experience defining non-functional requirements and architecture patterns for enterprise-grade resilience, observability, disaster recovery, data lifecycle management and operational readiness. Experience using architecture decision records, design authorities, exception management and measurable standards adoption to embed architectural governance ...

Software Development Engineer III

Location
Greater London, England, United Kingdom
across Expedia, Vrbo and Hotels.com run on what we help build, so the code you ship is felt quickly and good engineering habits (tests, observability, clear PRs) are valued as much as raw output. We lean heavily on coding agents (Cursor, Claude Code and friends) to accelerate the delivery cycle … meaningful scale. Demonstrated ability to lead architecture and implementation decisions within a service, technical domain, or cross-service initiative. Strong operational excellence, including automation, observability, resilience engineering, and data-driven improvements to reliability and performance. Experience applying AI/ML concepts, tools, or workflows to production software, including safely evaluating ...

Associate Director, Platform Engineering

Location
Greater London, England, United Kingdom
pipelines, internal services and developer tooling, treating the platform as a product and our internal engineers and computational scientists as its customers. Establish comprehensive observability and robust disaster recovery and failover strategies across on-premises and cloud environments. Own Relation’s platform security controls, including IAM, secrets management, network policy … understanding of the reliability, scalability and reproducibility requirements of scientific data pipelines. A track record of maturing CI/CD pipelines, developer tooling and observability across complex, multi-environment platforms. Strong security fundamentals, with hands-on experience designing and implementing appropriate security controls. Excellent written and verbal communication skills, with ...

Lead Infrastructure Engineer - Proxy/SSE Network Security

Location
Greater London, England, United Kingdom
resilience outcomes. Drive operational excellence at scale for perimeter, proxy, and SSE services in the US, including incident, change, and problem management rigor, observability and resiliency validation practices, automation to improve repeatability and evidence quality, reduction of client and partner impact, and execution of Technology Lifecycle Management (TLM) and modernization … design, exception frameworks, audit-ready traceability, and measurable risk reduction reporting. Experience with large-scale operations for externally facing or security enforcement services, including observability strategy, resilience testing, incident response alignment, and reduction of repeat incidents and client-impacting events. Experience designing and operating hybrid edge architectures and cloud interconnect ...

Frontend Engineer

Hiring Organisation
Zopa
Location
London, UK
Employment Type
Full-time
performance, reliability and accessibility, using data, experimentation and customer feedback to continuously improve our products. Participate in code reviews, testing, CI/CD and observability, helping maintain high engineering standards while learning from experienced teammates. Take ownership of your work, proactively identifying opportunities to improve our applications, developer experience … crawlability, structured data, indexation and semantic HTML), with experience optimising applications for both traditional search engines and AI-powered search. Experience with GraphQL and observability tools such as Honeycomb or Grafana. Working on customer-facing products within fintech, regulated industries or fast-growing technology businesses. Contributing to shared libraries, design ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
execution capabilities to agents Correctness, auditability and data integrity across financial workflows, so material actions remain traceable long after they happen Production ownership : observability, alerting, incident response, performance and scalability for the services you build Shared backend libraries, patterns and platform capabilities , plus technical direction, design and code review standards … data-critical systems across SQL and NoSQL, with a clear sense of APIs, service boundaries and data contracts Designing for correctness, resilience and observability , and building reliable integrations with external systems you do not control Delivery at scale of months. A record of scoping and delivering work measured in months ...

Software Engineer, Simulation

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
differences between on-road and simulated execution, identifying issues across data, inference and simulated components. Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling. Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently. Work with internal users and adjacent engineering teams … such as camera, radar, lidar or GNSS, including modelling uncertainty or noise. Experience integrating machine-learning inference into production systems. Experience with performance profiling, observability or debugging distributed systems. Familiarity with large-scale batch processing, cloud infrastructure or GPU-based workloads. This is a full-time role based ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
company's platform and third‐party integrations Optimise model inference pipelines for performance, cost, and scalability in production environments Implement monitoring, logging, and observability for GenAI systems to track usage, errors, and model behaviour Collaborate with data scientists to productionise ML models and forecasting algorithms Your Skills Extensive hands … Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with model observability tools (LangSmith, W& B;, MLflow) Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive ...

Software Engineer, Simulation London, United Kingdom

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components. Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling. Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently. Work with internal users and adjacent engineering teams … such as camera, radar, lidar or GNSS, including modelling uncertainty or noise. Experience integrating machine-learning inference into production systems. Experience with performance profiling, observability or debugging distributed systems. Familiarity with large-scale batch processing, cloud infrastructure or GPU-based workloads. This is a full-time role based ...

Technical Architect

Location
Greater London, England, United Kingdom
search layer used to ground agent responses in approved data sources. Delivering a production Internal MCP Gateway providing discovery, security, policy enforcement, observability, and lifecycle management for MCP tools and Skills. Designing and building MCP servers and Skills that expose internal and vendor systems safely to agents. Establishing evaluation, quality … Compliance stakeholders to validate and approve designs. Represent solutions in governance forums, clearly explaining architecture, risks, and controls. Ensure solutions meet security, resilience, scalability, observability, and compliance requirements. Align implementation with approved design, maintaining traceability and documentation integrity. Partner with ML, Quality, and SRE teams to ensure designs are deliverable ...

Technical Architect

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
search layer used to ground agent responses in approved data sources.* Delivering a production Internal MCP Gateway providing discovery, security, policy enforcement, observability, and lifecycle management for MCP tools and Skills.* Designing and building MCP servers and Skills that expose internal and vendor systems safely to agents.* Establishing evaluation, quality … Compliance stakeholders to validate and approve designs.* Represent solutions in governance forums, clearly explaining architecture, risks, and controls.* Ensure solutions meet security, resilience, scalability, observability, and compliance requirements.* Align implementation with approved design, maintaining traceability and documentation integrity.* Partner with ML, Quality, and SRE teams to ensure designs are deliverable ...

Director of Analytics

Location
Greater London, England, United Kingdom
drive strategic decision‐making Technical Direction and Innovation: Provide expert guidance on horizontal technical analytical areas, including experimentation, attribution, forecasting, predictive analytics, and observability Continuously identify opportunities for advanced analytics and data methodologies and enable implementation Identify and track appropriate metrics to measure success across all areas of the business … leadership skills with a track record of managing large, multi‐disciplinary teams Expertise in doing analytical work, including data modelling, experimentation, attribution, forecasting, and observability Experience leading product analytics and improving product experience based on analysis and insights Experience introducing and driving adoption of agentic analytics to an organisation Excellent ...

Software Engineer, Engineering Acceleration

Location
Greater London, England, United Kingdom
management. Own systems through delivery and operation. Take responsibility for technical design, implementation, deployment and ongoing operation. Build for reliability and maintainability, with appropriate observability, access controls and recovery mechanisms. What Can You Expect? Join the team early. Help shape and invent how engineers and agents design, build and verify … experience. You've built or maintained serverless infrastructure on AWS, deployment pipelines or development environments. You understand the trade-offs around access control, isolation, observability, reliability and cost. Hands-on experience with AI-assisted development. You've applied coding agents to software development and understand their capabilities and failure modes. ...

Management Consultant - Manager - AI Engineering

Location
Greater London, England, United Kingdom
production Integrate AI components with client and Moorhouse systems via APIs, event streams and workflow tools Implement robust engineering practices including version control, testing, observability and performance tuning Work with architects and data engineers to define scalable, secure and cost‐effective solution designs Collaborate with AI Delivery Leads to estimate … services) Experience integrating with enterprise systems (REST APIs, microservices, identity/permissions) Understanding of software engineering best practice (CI/CD, testing, code quality, observability) Confident communicating with both technical and non-technical audiences, with experience building effective relationships with senior client stakeholders and leadership teams. A passion for building ...

Product Engineer (Backend)

Location
Greater London, England, United Kingdom
V7 At V7, we’re building AI platforms that help humans do their best work, at incredible scale and speed. Our mission is to turn human knowledge into trustworthy AI, making complex tasks faster, smarter ...