2,051 to 2,075 of 2,284 Observability Jobs in London

Observability Software Engineer - Data & Dashboards

Location
Greater London, England, United Kingdom
Vercel is seeking a Software Engineer to join the Observability team in London. You will design, implement, and maintain observability features that help users monitor and understand their applications’ health and performance. The role offers in-office anchor days on Monday, Tuesday, and Friday for those within commuting distance; otherwise ...

Software Engineer II, AI Enablement Team

Location
Greater London, England, United Kingdom
Build cost observability and usage-tracking tooling across LLM and agent workloads Design and ship internal tooling and guardrails for safe, productive agentic coding workflows Evaluate AI tools, models, and vendors and produce build-vs-buy recommendations Write and maintain production-quality code and automated tests Review peers' code … particular technology stack or language Clear written and verbal communication Generalist attitude toward learning and growth Experience with LLM/agentic tooling, AI cost observability, or internal developer platforms is a bonus Core Competencies Demonstrates expertise in building cost observability and usage-tracking tooling for AI workloads, with a strong ...

Senior Product Manager (SaaS)

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
comfort with technical details are must-haves. You'll be on the technical side too, having experience with containerised platforms using Kubernetes, databases, and observability tools such as Prometheus and OpenTelemetry. This is a chance to shape the future of observability and security, build products people count ...

Lead Software Engineer - Agent Safety

Location
Greater London, England, United Kingdom
enforce pre- and post-generation guardrails, managing the overarching governance of AI models operating within internal tools and platforms. Drive AI security and observability: Build out dedicated auth/permissions for internal AI agents, establish deep monitoring/observability pipelines, and define incident response protocols for AI-specific anomalies. Verification … Inspect AI, Ragas, OpenAI Evals, NeMo Guardrails). Django experience and strong backend engineering patterns (security, performance, maintainability). Experience with Datadog for complex observability, tracing, and monitoring in AI environments. Familiarity with foundational AI engineering tooling like Pydantic AI, LiteLLM, or LangChain. Are you ready for a career with ...

Senior Infrastructure & Operations Engineer (Kubernetes / Platform Reliability)

Location
Greater London, England, United Kingdom
production systems at scale and who focuses on making infrastructure predictable and stable. You’ll work across Kubernetes, networking, CI/CD, Cloudflare, and observability to create a platform engineers can trust. What You’ll Do Design, deploy, and maintain production Kubernetes clusters. Own cluster reliability, upgrades, security, and performance. … Build and operate monitoring, logging, and alerting pipelines. Ensure full-stack observability across infrastructure and services. Design and maintain CI/CD pipelines that are fast, reproducible, and safe. Improve deployment strategies (rollouts, canaries, rollbacks). Automate infrastructure provisioning and configuration. Investigate and resolve production incidents. Improve system resilience, redundancy ...

Senior Java Software Developer London, UK

Location
Greater London, England, United Kingdom
ITRS is an Enterprise SaaS provider with industry-leading solutions. Our mission is to make society’s critical technology work via automated & holistic IT observability solutions that safeguard critical applications and enable innovation. With our prestigious customer base includes 90% of the world's top investment banks. We are backed … form part of a wider global Engineering Team. The Core Platform layer is a collection of distributed services which ingest, transform and materialise observability data to make it available to several similarly distributed visualisation, integration, analytics and other domain specific applications to provide solutions to a range of observability problems. ...

Senior Software Engineer - Login Platform

Location
Greater London, England, United Kingdom
teams to deliver a reliable and secure login experience. Contribute to architectural decisions and long-term technical strategy for the Login Platform. Improve tooling, observability, scalability, and performance across critical infrastructure services. Explore and champion new ideas to evolve our systems responsibly over time. Engineers on the team are involved … throughout the software lifecycle, including design, deployment, observability, incident response, and long-term platform evolution. Our technologies We work primarily in C++ and some Python, running mostly on Linux, with continued support for Solaris clients until full migration is complete. Because we sit deep in the infrastructure stack ...

Platform Engineer

Location
Greater London, England, United Kingdom
direction and build the systems, tooling and processes that the wider engineering team relies on. You’ll work across production infrastructure, Linux performance, observability, deployments, developer experience and internal tooling, with significant freedom to decide what needs improving and take ownership of delivering it. Responsibilities Own and improve infrastructure supporting … real-time, 24/7 production systems Build reliable deployment, rollback and operational workflows Improve observability across metrics, logging, dashboards, tracing and alerting Develop tooling and automation that allows engineers to ship faster and more safely Improve CI/CD, build processes, test environments and configuration workflows Work ...

Product Engineering Environment Lead

Hiring Organisation
Randstad Technologies Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £650/day
tested recovery capabilities with defensible RTO/RPO metrics. Automation & Self-Service: Drive Infrastructure/Environments as Code (IaC/EaC), automated provisioning, and observability to enable on-demand instantiation. Your Experience Professional Background: Proven track record leading enterprise-scale environment estates, spanning hybrid cloud and physical/network infrastructure. … Automation & Tooling: Hands-on experience with Infrastructure-as-Code and Environment-as-Code tools (e.g., Terraform, Ansible), CI/CD pipelines, and observability frameworks. SRE, SDLC & DR Depth: Strong understanding of SRE principles (SLIs/SLOs, toil reduction), release management, and High Availability/Disaster Recovery design and testing. ...

Data Scientist - Senior Associate - Applied AI & Machine Learning

Location
Greater London, England, United Kingdom
world’s largest financial institutions. Job Responsibilities: Design and implement scalable architecture for LLM-powered investment tools, ensuring integration, governance, observability, and access control Build and scale AI applications and automated workflows using models, retrieval, and tools, with orchestration for state management and human oversight Develop context-management and retrieval … applications Establish engineering standards for reusable tools and model integrations, including interfaces, permissions, testing, failure handling, and documentation Implement evaluation and end-to-end observability for AI systems, optimizing for quality, groundedness, task completion, latency, token usage, cost, reliability, and business impact Partner with portfolio managers and research teams ...

Linux Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
RHEL and Ubuntu) that underpin our compute environment. You'll work across the full Linux stack, from OS builds and provisioning systems through to observability, security and custom tooling. Mathematics and science are at our core; AI pushes them further to unlock bigger thinking and bigger impact. We expect every … using Ansible, to automate deploymentsLeading vulnerability response, including CVE triage and kernel and package remediationHelping design agentic frameworks for safe, autonomous detection and remediationImproving observability and telemetry to track OS health, provisioning and complianceWho are we looking for? The ideal candidate will have the following skills and experience: Strong hands ...

Staff ML Engineer | Agentic AI & Applied ML | London | |

Location
Greater London, England, United Kingdom
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high‐risk and high‐complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit‐learn Experience operating in complex, regulated or high‐risk environments would ...

Senior Systems Engineer SRE Golang - FinTech

Hiring Organisation
Client Server
Location
London, UK
Employment Type
Full-time
everything you do. You'll design systems with the assumption that failures will happen, building resilient, adaptable and highly available services, working across performance, observability and security, using logs, metrics and traces to understand system behaviour and correlate technical performance with business outcomes. You'll also work with Infrastructure … application and the infrastructure needed to support itYou have a strong knowledge of network, operating system and application level securityYou have experience with systems observability/SRE e.g. logs, metricsYou have experience with IaC, Terraform preferredYou have a good understanding of Kubernetes and how it worksYou have a good knowledge ...

Senior Systems Engineer SRE Golang - FinTech

Hiring Organisation
Client Server
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
everything you do. You'll design systems with the assumption that failures will happen, building resilient, adaptable and highly available services, working across performance, observability and security, using logs, metrics and traces to understand system behaviour and correlate technical performance with business outcomes. You'll also work with Infrastructure … infrastructure needed to support it You have a strong knowledge of network, operating system and application level security You have experience with systems observability/SRE e.g. logs, metrics You have experience with IaC, Terraform preferred You have a good understanding of Kubernetes and how it works You have ...

Senior Database Administrator (Open-Source/FinTech Stack)

Hiring Organisation
N.P.A
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 80,000 - 95,000 Annual
Time-Series: QuestDB (Used for massive market data ingestion & operational metrics) Version-Controlled SQL: Dolt (Git-for-data branching technology) Caching & Search: Redis & Elasticsearch Observability: Grafana & Linux-based monitoring tools Key Responsibilities Estate Ownership: Administer, monitor, and scale the database ecosystem to ensure continuous high availability and reliability. Developer Collaboration … Partner closely with engineering teams on schema design, query optimization, and database access patterns. Build Observability: Write complex SQL queries to surface vital business and operational metrics onto Grafana dashboards. Infrastructure Resilience: Plan and execute seamless upgrades, patch management, and routinely test point-in-time recovery (RPO/RTO) strategies. ...

Operations and SRE Manager

Location
Carshalton, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities Lead the implementation of the team’s strategic direction, translating priorities into clear operational plans, backlogs … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Senior Engineering Manager (GenAI & Agentic Platforms)

Location
Greater London, England, United Kingdom
Agentic Platform: approximately 8-12 engineers across two teams, while partnering closely with ML Platform leadership. GenAI Platform owns shared model access, routing, evaluation, observability, prompt and configuration lifecycle, retrieval, grounding, and guardrails. Agentic Platform builds on those foundations with durable execution, tools, state, context, memory, permissions, human controls, agent … coaching. Own hiring quality, team composition, evolving team boundaries, and operating models as the platforms and company change. Guide architecture across model access, evaluation, observability, retrieval, agent orchestration, tools, state, memory, permissions and control planes. Translate strategy into realistic plans, balancing foundational investment with near-term product needs while surfacing ...

Software Engineer II, AI Enablement Team

Location
Greater London, England, United Kingdom
Swap, we're building a culture that values clarity, creativity, and shared ownership as we redefine how global commerce works. Responsibilities Build cost observability and usage‐tracking tooling across LLM and agent workloads, so teams can see what AI actually costs and where we can accelerate productivity and impact. Design … problems rather than working on particular tech stack/languages Clear written and verbal communicator. Bonus: experience with LLM/agentic tooling, AI cost observability, or internal developer platforms. A generalist with an attitude for learning and growth. Diversity & Equal Opportunities: We embrace diversity and equality in a serious way. ...

Software Developer - Data Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
platforms, and automation that our technology and investment teams depend on. The team operates at the intersection of engineering and operations. We build the observability tooling that surfaces problems before they become incidents, the job orchestration platform that runs production workloads at scale, the self-serve systems that let teams … with a global team and external vendorsMindset: Proactive, detail-oriented, and self-driven with a strong sense of ownership and accountabilityNice to haveExperience with observability and monitoring tools such as Grafana, Kibana, or PrometheusExperience developing automation tooling and implementing configuration managementExperience with cloud platforms such as Google Cloud or AWSExperience ...

AI TEST ENGINEER

Location
Greater London, England, United Kingdom
Automation. The role is focused on building and evolving a next-generation testing framework for generative AI agents, including evaluation pipelines, synthetic data generation, observability layers, and quality metrics for LLM-based systems. You will work across AI engineering, data pipelines, and quality automation, contributing both to architecture design … Responsibilities: Design and develop Python-based frameworks for testing and evaluating AI agents and LLM-based systems Contribute to architecture design for evaluation pipelines, observability, and data‐driven testing systems Build and maintain tools for test data generation (including synthetic and adversarial datasets) Define and implement evaluation strategies and quality ...

Software Engineering Lead

Hiring Organisation
WTW
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
leadership for one of our teams. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. We're particularly interested in leaders who have experience working in fast-growing … modern services and legacy systems Shape event-driven and integration patterns so teams can build independently without breaking cross-squad journeys Champion automated testing, observability, and production readiness for backend services - PHPUnit, Go tests, Datadog, and safe release practices Oversee database and schema evolution practices - migrations, data integrity, and performance ...

Software Engineering Lead

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
leadership for one of our teams. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. We're particularly interested in leaders who have experience working in fast-growing … integration between modern services and legacy systemsShape event-driven and integration patterns so teams can build independently without breaking cross-squad journeysChampion automated testing, observability, and production readiness for backend services - PHPUnit, Go tests, Datadog, and safe release practicesOversee database and schema evolution practices - migrations, data integrity, and performance ...

Senior Software Development Engineer

Location
City Of London, England, United Kingdom
expectations and can be reused across multiple brands and platforms. Drive engineering excellence for the services you own by championing code quality, automated testing, observability, performance optimization, and simplification, taking technical responsibility for service health, scalability, resilience, and the ongoing reduction of technical debt and operational overhead. Provide technical mentorship … integrations, and communicating trade-offs to both technical and non-technical stakeholders. Track record of improving operational excellence at the team level through enhanced observability, automation, performance tuning, and data-driven analysis of incidents and customer impact. Hands-on experience integrating or consuming AI/ML-enabled services or platforms ...

Data Platform SRE

Location
Greater London, England, United Kingdom
wider business depends on for analytics, reporting, and product features. This is a hands‐on engineering role: you'll write infrastructure‐as‐code, build observability into data systems from the ground up, lead incident response for data platform outages, and work directly with data engineers to raise the reliability … maintain the infrastructure that underpins the data platform (compute, storage, orchestration, networking) using infrastructure-as-code primarily focussed in Microsoft Azure. Build and improve observability for data systems — metrics, logging, tracing, and data‐quality/freshness monitoring — so issues are caught before they reach downstream consumers. Lead incident response ...