576 to 600 of 2,278 Observability Jobs in London

Product Manager - Data

Location
Greater London, England, United Kingdom
connected vehicles, and compliance with automotive industry standards such as ISO 26262 and ISO 21434 Application Management - Open source solutions in the enterprise including Observability, IAM, App Stores and technologies such Grafana, GitOps, and Juju Charms We will route you to the most suitable team. Location: These roles are home ...

Backend Software Engineer - Infrastructure

Hiring Organisation
Palantir Technologies
Location
London, UK
Employment Type
Full-time
efficiently scheduling hundreds of thousands of containers every hourDesigning architecture and opinionated APIs to keep application developers on the happy pathTracing and performance observability in high scale distributed microservice architecturesBuilding reliant, performant, and scalable systems for storage, auth, or asset serving to enable other product teams to build robust applications ...

Director of Data & AI

Location
Greater London, England, United Kingdom
evolution of Hyve's enterprise data platform and data warehouse capabilities. Lead the Data Engineering team responsible for data ingestion, transformation, orchestration, modelling, testing, observability and operational support. Ensure appropriate use of Microsoft Fabric, Azure Data Factory, Snowflake and dbt across Hyve's Azure and AWS cloud environments. Establish engineering ...

Lead Platform Engineer

Location
City of Westminster, England, United Kingdom
testing, documentation and code organisation. Responsible for evolving those standards over time, protecting core principles while adapting to new tools and approaches. Ensure appropriate observability, logging and error-handling patterns are in place across applications. Responsible for ensuring documentation exists where it adds long-term value, and remains accurate. Problem ...

Staff Product Engineer

Location
Greater London, England, United Kingdom
quickly and decide whether it is worth pursuing before Parloa invests a full quarter in it. Own production quality and security. Evals, production readiness, observability and incident playbooks for enterprise‐scale traffic. Meet the data handling standards a regulated customer expects, including PII, access control and audit trails. Debug live ...

Senior Manager - Microsoft AI Technical Architect - TC - FS

Location
Greater London, England, United Kingdom
Foundry, Azure OpenAI Service, Microsoft Fabric and supporting Azure/AI services. Design agent orchestration, grounding, tool invocation, memory and context handling, fallback behaviour, observability and failure recovery. Define non‐functional requirements covering scalability, latency, resilience, availability, supportability, cost, performance and recoverability. Integration, data and enterprise architecture Design secure integration ...

Data Science & AI Delivery Lead

Hiring Organisation
Axis Capital
Location
London, UK
Employment Type
Full-time
applications, agentic systems, machine learning solutions, and AI platforms. Lead evaluation of emerging AI technologies and determine appropriate adoption. Establish standards for MLOps, AI observability, model lifecycle management, and responsible AI implementation. Collaboration: Work closely with AXIS's Business Technology Solutions for solution architecture and delivery, and with the Program ...

Senior Manager - Microsoft AI Technical Architect - TC - FS

Hiring Organisation
EY (Ernst & Young)
Location
London, UK
Employment Type
Full-time
Foundry, Azure OpenAI Service, Microsoft Fabric and supporting Azure/AI services. Design agent orchestration, grounding, tool invocation, memory and context handling, fallback behaviour, observability and failure recovery. Define non-functional requirements covering scalability, latency, resilience, availability, supportability, cost, performance and recoverability. Integration, data and enterprise architectureDesign secure integration with ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
APIs, and serverless services Infrastructure & tooling: AWS, Terraform, Docker, Kubernetes, Redis, CI/CD pipelines Practices: Automation-first, metrics-driven, incident write-ups and observability baked in Our 2026 Engineering Strategy We’re not just here to write code, we’re here to redefine how insurance works. Our 2026 engineering ...

AI Data/Graph Engineer

Location
Greater London, England, United Kingdom
Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, OpenTelemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyperscaler or on-prem. A tool-for-tool match is not expected ...

Principal Data Engineer (we have office locations in Cambridge, Leeds and London) London

Location
Greater London, England, United Kingdom
into platforms handling sensitive clinical and research data Strengthening engineering practices and platform resilience through DataOps, DevOps, CI/CD, infrastructure as code, automation, observability and security by design Evaluating emerging technologies and helping teams make thoughtful, appropriate and sustainable technology choices Mentoring experienced engineers, sharing knowledge and contributing ...

Rapid Application Engineer London ·

Location
Greater London, England, United Kingdom
your current level. Comfortable with production operations: You understand that code runs in production with real traders depending on it. You think about reliability, observability, and operational concerns, not just features. Growth mindset: You're early in your career relative to senior engineers, and you're excited about the opportunity ...

Senior QA Engineer Performance & Scalability

Location
Greater London, England, United Kingdom
tests, and interpreting results meaningfully. K6 or similar - You have real, hands‐on depth with k6 for building and running performance tests. Monitoring and observability know‐how. You're comfortable with tools like Datadog, Grafana, CloudWatch or similar, and you use them proactively to correlate test results with what ...

Senior Engineering Manager

Hiring Organisation
Multiverse Group
Location
London, UK
Employment Type
Full-time
market, and eligibility — closer to the customer than most engineering roles. Craft that's respected, not traded off. Strict typing, layered testing, ADRs, observability — and AI-native ways of working as the baseline, not an experiment. A company mid-transformation, told straight. Multiverse is rebuilding itself into a product-powered ...

Cloud Platform Engineer — AWS, Terraform & Observability

Location
Greater London, England, United Kingdom
Clio is the global leader in legal AI technology, empowering legal professionals and law firms of every size to work smarter, faster, and more securely. We are transforming the legal experience for all by bettering ...

Director of Platform Engineering - SaaS & Observability

Location
Greater London, England, United Kingdom
ITRS is seeking a Director of Platform Engineering to lead a hands-on, high-availability SaaS platform that serves global financial institutions. You will supervise the Analytics SaaS stack, guide architecture decisions, and drive automation ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams.It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling.**Responsibilities*** Lead the design and evolution of scalable, secure, and highly available trading infrastructure.* Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows.* Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering.* Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices.* Lead and participate ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
monitor, diagnose, and auto-remediate global SaaS infrastructure. What you'll do Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands‐on experience building LLM pipelines, AI Agents, Model ...

Software Engineer

Location
Greater London, England, United Kingdom
diagnose, and auto-remediate global SaaS infrastructure. What You'll Do Technical Design & Architecture: Design and implement high-resilience software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design, build, and maintain production-grade AI agents, MCP tool integrations, and deterministic evaluation … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated root cause analysis (RCA). Preferred Qualifications AI & Agentic Systems: E xperience building LLM pipelines ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands‐on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement. We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve … teams to productionize AI solutions and ensure services are ready to operate reliably at scale. Establish engineering standards and reusable patterns for infrastructure, security, observability, resilience, documentation, and operational readiness. Lead architectural decisions and evaluate trade-offs across reliability, security, scalability, performance, cost, and maintainability. Take ownership of operational risks ...

DevOps Engineer, Blockchain Infra (Fully Remote)

Hiring Organisation
Binance
Location
London, UK
Employment Type
Full-time
pipelines for application and infrastructure deployment. Deploy and operate middleware platforms such as Kafka, Redis and NGINX.Automate operational tasks using Golang and Python. Improve observability through monitoring, logging, alerting, and distributed tracing. Ensure platform reliability, scalability, security, and disaster recovery. Troubleshoot production incidents and perform root cause analysis. Work closely …/Redis/NGINXStrong scripting and programming skills in: Golang/PythonExperience with Git, GitOps workflows, and CI/CD platforms. Strong understanding of observability tools such as Prometheus, OpenTelemetry. Familiarity with container technologies including Docker and Kubernetes. Strong troubleshooting and problem-solving skills. Preferred QualificationsExperience operating blockchain infrastructure ...

Senior SRE

Hiring Organisation
Pigment
Location
London, UK
Employment Type
Full-time
platform usage. Secure high availability and redundancy of the Pigment platform. Ensure that the platform's performance and correctness are monitored accordingly, spread observability best practices across the engineering team. Participate in incident response. Accompany Pigment geographical expansion as the company grows and we sign clients overseas. Work with … experience in software developments with languages such as C#, Java, C++, Golang, Rust, JavaScript, Python, or Ruby (this list is not exhaustive).Experience with observability tools (e.g. Datadog, Prometheus, ELK, Jaeger...)Great team spirit with a problem-solving attitude. A good dose of humility and the willingness to grow ...

AI Platform Engineer- Senior Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
Implement MLOps and LLMOps pipelines (model deployment, monitoring, retraining, and fine-tuning where relevant) using Infrastructure-as-Code, GitOps, and CI/CD• Establish observability, security, and governance frameworks specific to AI systems, including cost attribution and lifecycle management• Work with clients and internal teams to develop new opportunities … equivalent), including familiarity with the Model Context Protocol (MCP)• Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)• Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust)**MLOps & LLMOps**• Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone ...