2,001 to 2,025 of 2,213 Observability Jobs in London

Product Engineering Environment Lead

Hiring Organisation
Randstad Technologies Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £650/day
tested recovery capabilities with defensible RTO/RPO metrics. Automation & Self-Service: Drive Infrastructure/Environments as Code (IaC/EaC), automated provisioning, and observability to enable on-demand instantiation. Your Experience Professional Background: Proven track record leading enterprise-scale environment estates, spanning hybrid cloud and physical/network infrastructure. … Automation & Tooling: Hands-on experience with Infrastructure-as-Code and Environment-as-Code tools (e.g., Terraform, Ansible), CI/CD pipelines, and observability frameworks. SRE, SDLC & DR Depth: Strong understanding of SRE principles (SLIs/SLOs, toil reduction), release management, and High Availability/Disaster Recovery design and testing. ...

Data Scientist - Senior Associate - Applied AI & Machine Learning

Location
Greater London, England, United Kingdom
world’s largest financial institutions. Job Responsibilities: Design and implement scalable architecture for LLM-powered investment tools, ensuring integration, governance, observability, and access control Build and scale AI applications and automated workflows using models, retrieval, and tools, with orchestration for state management and human oversight Develop context-management and retrieval … applications Establish engineering standards for reusable tools and model integrations, including interfaces, permissions, testing, failure handling, and documentation Implement evaluation and end-to-end observability for AI systems, optimizing for quality, groundedness, task completion, latency, token usage, cost, reliability, and business impact Partner with portfolio managers and research teams ...

Staff ML Engineer | Agentic AI & Applied ML | London | |

Location
Greater London, England, United Kingdom
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high‐risk and high‐complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit‐learn Experience operating in complex, regulated or high‐risk environments would ...

Senior Systems Engineer SRE Golang - FinTech

Hiring Organisation
Client Server
Location
London, UK
Employment Type
Full-time
everything you do. You'll design systems with the assumption that failures will happen, building resilient, adaptable and highly available services, working across performance, observability and security, using logs, metrics and traces to understand system behaviour and correlate technical performance with business outcomes. You'll also work with Infrastructure … application and the infrastructure needed to support itYou have a strong knowledge of network, operating system and application level securityYou have experience with systems observability/SRE e.g. logs, metricsYou have experience with IaC, Terraform preferredYou have a good understanding of Kubernetes and how it worksYou have a good knowledge ...

Senior Systems Engineer SRE Golang - FinTech

Hiring Organisation
Client Server
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
everything you do. You'll design systems with the assumption that failures will happen, building resilient, adaptable and highly available services, working across performance, observability and security, using logs, metrics and traces to understand system behaviour and correlate technical performance with business outcomes. You'll also work with Infrastructure … infrastructure needed to support it You have a strong knowledge of network, operating system and application level security You have experience with systems observability/SRE e.g. logs, metrics You have experience with IaC, Terraform preferred You have a good understanding of Kubernetes and how it works You have ...

Senior Database Administrator (Open-Source/FinTech Stack)

Hiring Organisation
N.P.A
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 80,000 - 95,000 Annual
Time-Series: QuestDB (Used for massive market data ingestion & operational metrics) Version-Controlled SQL: Dolt (Git-for-data branching technology) Caching & Search: Redis & Elasticsearch Observability: Grafana & Linux-based monitoring tools Key Responsibilities Estate Ownership: Administer, monitor, and scale the database ecosystem to ensure continuous high availability and reliability. Developer Collaboration … Partner closely with engineering teams on schema design, query optimization, and database access patterns. Build Observability: Write complex SQL queries to surface vital business and operational metrics onto Grafana dashboards. Infrastructure Resilience: Plan and execute seamless upgrades, patch management, and routinely test point-in-time recovery (RPO/RTO) strategies. ...

Operations and SRE Manager

Location
Carshalton, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities Lead the implementation of the team’s strategic direction, translating priorities into clear operational plans, backlogs … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Senior Engineering Manager (GenAI & Agentic Platforms)

Location
Greater London, England, United Kingdom
Agentic Platform: approximately 8-12 engineers across two teams, while partnering closely with ML Platform leadership. GenAI Platform owns shared model access, routing, evaluation, observability, prompt and configuration lifecycle, retrieval, grounding, and guardrails. Agentic Platform builds on those foundations with durable execution, tools, state, context, memory, permissions, human controls, agent … coaching. Own hiring quality, team composition, evolving team boundaries, and operating models as the platforms and company change. Guide architecture across model access, evaluation, observability, retrieval, agent orchestration, tools, state, memory, permissions and control planes. Translate strategy into realistic plans, balancing foundational investment with near-term product needs while surfacing ...

Software Engineer II, AI Enablement Team

Location
Greater London, England, United Kingdom
Swap, we're building a culture that values clarity, creativity, and shared ownership as we redefine how global commerce works. Responsibilities Build cost observability and usage‐tracking tooling across LLM and agent workloads, so teams can see what AI actually costs and where we can accelerate productivity and impact. Design … problems rather than working on particular tech stack/languages Clear written and verbal communicator. Bonus: experience with LLM/agentic tooling, AI cost observability, or internal developer platforms. A generalist with an attitude for learning and growth. Diversity & Equal Opportunities: We embrace diversity and equality in a serious way. ...

Software Developer - Data Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
platforms, and automation that our technology and investment teams depend on. The team operates at the intersection of engineering and operations. We build the observability tooling that surfaces problems before they become incidents, the job orchestration platform that runs production workloads at scale, the self-serve systems that let teams … with a global team and external vendorsMindset: Proactive, detail-oriented, and self-driven with a strong sense of ownership and accountabilityNice to haveExperience with observability and monitoring tools such as Grafana, Kibana, or PrometheusExperience developing automation tooling and implementing configuration managementExperience with cloud platforms such as Google Cloud or AWSExperience ...

AI TEST ENGINEER

Location
Greater London, England, United Kingdom
Automation. The role is focused on building and evolving a next-generation testing framework for generative AI agents, including evaluation pipelines, synthetic data generation, observability layers, and quality metrics for LLM-based systems. You will work across AI engineering, data pipelines, and quality automation, contributing both to architecture design … Responsibilities: Design and develop Python-based frameworks for testing and evaluating AI agents and LLM-based systems Contribute to architecture design for evaluation pipelines, observability, and data‐driven testing systems Build and maintain tools for test data generation (including synthetic and adversarial datasets) Define and implement evaluation strategies and quality ...

Software Engineering Lead

Hiring Organisation
WTW
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
leadership for one of our teams. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. We're particularly interested in leaders who have experience working in fast-growing … modern services and legacy systems Shape event-driven and integration patterns so teams can build independently without breaking cross-squad journeys Champion automated testing, observability, and production readiness for backend services - PHPUnit, Go tests, Datadog, and safe release practices Oversee database and schema evolution practices - migrations, data integrity, and performance ...

Senior Software Development Engineer

Location
City Of London, England, United Kingdom
expectations and can be reused across multiple brands and platforms. Drive engineering excellence for the services you own by championing code quality, automated testing, observability, performance optimization, and simplification, taking technical responsibility for service health, scalability, resilience, and the ongoing reduction of technical debt and operational overhead. Provide technical mentorship … integrations, and communicating trade-offs to both technical and non-technical stakeholders. Track record of improving operational excellence at the team level through enhanced observability, automation, performance tuning, and data-driven analysis of incidents and customer impact. Hands-on experience integrating or consuming AI/ML-enabled services or platforms ...

Data Platform SRE

Location
Greater London, England, United Kingdom
wider business depends on for analytics, reporting, and product features. This is a hands‐on engineering role: you'll write infrastructure‐as‐code, build observability into data systems from the ground up, lead incident response for data platform outages, and work directly with data engineers to raise the reliability … maintain the infrastructure that underpins the data platform (compute, storage, orchestration, networking) using infrastructure-as-code primarily focussed in Microsoft Azure. Build and improve observability for data systems — metrics, logging, tracing, and data‐quality/freshness monitoring — so issues are caught before they reach downstream consumers. Lead incident response ...

Integration Engineering Manager

Location
Greater London, England, United Kingdom
Agent-ready API products, reusable integration assets, and governed API catalogs AI-enabled integration patterns across Salesforce, Data, ERP, customer, operational, and platform domains Observability, policy enforcement, auditability, and compliance for AI-driven integration flows The manager will oversee a broad portfolio of APIs, connectors, MCP-enabled tools, and integration … routing, policy enforcement, and agent-to-system orchestration. Ensure AI-enabled integrations follow strong standards for identity, consent, access control, data minimisation, rate limiting, observability, and auditability. Provide architectural guidance on complex use cases involving multi-system orchestration, agent workflows, human-in-the-loop controls, and event-driven automation. Ensure ...

Principal Product Engineer

Location
Greater London, England, United Kingdom
typed APIs, clear service and data boundaries, robust processing workflows and platform capabilities that can handle millions of records and events without compromising correctness, observability or operability. A key part of the role is separating the data layer from the application layer. You will help ensure an action taken … propagation separately, ensuring that customer-facing actions can be reversed safely without creating hidden inconsistency underneath. Reliability Engineering : Improve idempotency, retry behaviour, failure isolation, observability, alerting and recovery across critical workflows and integrations. Technical Leadership : Lead design reviews, mentor through code review and pairing, make technical standards explicit, and help ...

Software Engineer - Fleet Management & Repair Automation

Location
Greater London, England, United Kingdom
between competing operations, rate limiting, graceful cancellation, and emergency stop. Autonomy: replacing human-driven operational decisions with automation and agentic workflows, backed by the observability and quality signals needed to trust them. Practically, this means we build large-scale workflow orchestration engines: systems that model operational intent, schedule it against … maintaining production systems in one or more of Python, C++, Java, Rust or Go Track record of operating critical systems: oncall ownership, incident leadership, observability and SLO design, and structural reliability improvement Ability to lead ambiguous, cross-organisational work to completion, and to communicate clearly in writing to both engineers ...

Product Engineer (Backend/AI @Briink)

Location
Greater London, England, United Kingdom
that support multiple user journeys and reporting workflows rather than solving each problem in isolation Improve the robustness of our systems, strengthening reliability, testing, observability, performance, and maintainability as we scale Work directly with users, product, and data, using qualitative feedback and product data to understand the real problem … combined with the ability and willingness to become productive in Python quickly) A solid software-quality mindset, including testing, code review, CI/CD, observability, and pragmatic approaches to reliability and security Experience in a small, high-ownership product team, ideally within a startup or scale-up engineering organisation ...

Engineering Director

Location
Greater London, England, United Kingdom
Structuring technical problem-solving frameworks and coordinating incident response to resolve complex, interdependent system failures and maintain system reliability Guiding the development of comprehensive observability, monitoring, and alerting strategies in order to ensure system health, visibility, and operational excellence Driving structured delivery governance, CI/CD practices, and automated pipelines … environment (such as healthcare, life sciences, or financial services), including information security, data protection, and quality management obligations Proven success developing and implementing comprehensive observability, monitoring, and delivery governance strategies aligned with business objectives In-depth knowledge of how to navigate ambiguous technical and organisational environments, structuring decision-making frameworks ...

Senior Product Manager - Storage & Networking

Location
Greater London, England, United Kingdom
underpinning Radiant’s GPU platform. You’ll work across bare-metal GPU clusters, Kubernetes, high-performance storage, data‐centre networking, infrastructure inventory, automation, and observability, partnering closely with engineering, SRE, infrastructure, networking, and operations. This is a highly technical product role focused on ensuring GPU workloads have reliable, high‐throughput … infrastructure through to customer and workload connectivity. Partner with engineering to turn requirements into scalable, reliable platform services. Drive improvements in automation, self‐service, observability, and operational efficiency. Define success metrics and use them to guide performance, reliability, capacity, and investment decisions. Qualifications: Experience owning technical infrastructure, platform, storage, networking ...

Model Release Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
with teams across Wayve to understand their requirements, agree interfaces and resolve technical or delivery conflicts across shared workflows. Improve the reliability, scalability and observability of the platform through effective monitoring, alerting and operational tooling. Provide clear visibility of model candidates, their progress, evaluation results, approvals and release status. … work through conflicting priorities. Experience operating cloud-based services using Kubernetes, with a good understanding of reliability, scalability and performance. Practical knowledge of observability, monitoring and alerting, including defining meaningful service-health metrics. Confidence using AI coding tools and agents to improve engineering productivity and automate repeatable work. A pragmatic ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
Anaplan’s platform and third‐party integrations Optimise model inference pipelines for performance, cost, and scalability in production environments Implement monitoring, logging, and observability for GenAI systems to track usage, errors, and model behaviour Collaborate with data scientists to productionise ML models and forecasting algorithms Your Skills Extensive hands … Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with model observability tools (LangSmith, W&B;, MLflow) Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive ...

Model Release Engineer

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
with teams across Wayve to understand their requirements, agree interfaces and resolve technical or delivery conflicts across shared workflows. Improve the reliability, scalability and observability of the platform through effective monitoring, alerting and operational tooling. Provide clear visibility of model candidates, their progress, evaluation results, approvals and release status. … work through conflicting priorities. Experience operating cloud-based services using Kubernetes, with a good understanding of reliability, scalability and performance. Practical knowledge of observability, monitoring and alerting, including defining meaningful service-health metrics. Confidence using AI coding tools and agents to improve engineering productivity and automate repeatable work. A pragmatic ...

Partner Sales Manager - EMEA

Location
Greater London, England, United Kingdom
March Capital, Lightspeed, Sorenson Ventures, Industry Ventures, and Emergent Ventures, we are a Series-C funded company headquartered in Silicon Valley. Our Enterprise Data Observability Platform - the first of its kind - helps enterprises build and operate world-class data products by ensuring data is reliable, trusted, and ready to power … strategy with hyperscaler priorities and customer modernization initiatives. Product and Industry Expertise and Demonstration Maintain a strong understanding of Acceldata's platform, including data observability use cases across modern and legacy data architectures. Confidently articulate Acceldata's value to partner sales, technical, and executive audiences. Support and, when needed, deliver ...

Solutions Architect

Location
Greater London, England, United Kingdom
Maintain comprehensive documentation of design decisions, patterns, standards, and trade-offs. Define non-functional platform architecture standards covering resilience, backup/restore, disaster recovery, observability, service levels, auditability and operational readiness. Define platform security and privacy architecture in partnership with Cyber Security and Data Governance, including PII handling, access recertification … authority in environments transitioning from outsourced to in-house ownership is beneficial. Experience defining non-functional requirements and architecture patterns for enterprise-grade resilience, observability, disaster recovery, data lifecycle management and operational readiness. Experience using architecture decision records, design authorities, exception management and measurable standards adoption to embed architectural governance ...