1,026 to 1,050 of 2,373 Observability Jobs in London

Engineering Manager, Connectivity - London

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
engineering, with the depth to guide architectural decisions and engage credibly in technical discussionsExperience operating production distributed systems, with an understanding of what reliability, observability, and incident response require at scaleExcellent communication skills, with the ability to build consensus across teams and time zonesDemonstrated success building a culture of belonging ...

Team Lead, Platform Engineering

Location
Greater London, England, United Kingdom
remediated effectively. Coordinate with the other platform engineering leads on shared architecture, joint initiatives, and cross-team and cross-timezone delivery. Own reliability, observability, security, and interface and versioning standards for the services your team ships. Stay hands‐on: take on complex or high‐risk engineering work directly alongside ...

Senior Data Analyst I

Location
Greater London, England, United Kingdom
Spark, SQL-based systems). Ensure high standards for data quality, metric consistency, and instrumentation reliability. Collaborate with engineering to improve logging, tracking, and observability of search systems. Cross-functional Impact & Mentorship Act as a key analytics partner to search data scientists, engineers, and product teams. Help elevate the team ...

AI Roku Developer

Location
Greater London, England, United Kingdom
agentic coding tools. Claude Code experience is strongly preferred. Experience with OpenAI Codex is desirable. Experience with MCP servers, agent frameworks, custom tool integrations, observability, or automated AI evaluations is advantageous. Experience with Roku test automation, for example Rooibos unit testing or ECP-based device automation, is advantageous. Knowledge ...

Engineering Manager, Edge SRE

Location
Greater London, England, United Kingdom
deliver with predictability Incident root cause analysis and follow-ups Comfortable managing teams/projections with deadlines and short release cycles Experience using observability tools such as Jaeger, OpenTracing, ELK, Prometheus, Thanos, Grafana, Clickhouse Experience running and maturing distributed systems Familiarity working with Proxies, DNS, Databases, Internet and Security Experience ...

Global IT Software Engineer Senior Director & Tech Area Lead - Legal

Hiring Organisation
The Boston Consulting Group
Location
London, UK
Employment Type
Full-time
offs vs. long-term capability build, platform and vendor selection, SDLC best-practices, and the establishment and regular review of technology and AI-quality observability and health metrics — while providing clarity and direction where priorities, ownership, or approaches are not well defined. From a people-management perspective, you will lead ...

Senior Azure DevOps Engineer

Location
Greater London, England, United Kingdom
position focuses on building and improve containerised environments using Docker and Kubernetes. A highly varied technical environment spanning Azure, Kubernetes, Terraform, CI/CD, observability and automation. What you'll bring Proficiency in at least one programming or scripting language such as Python, Go or Node.js Familiarity with … large-scale systems using tools such as GitLab CI/CD, GitHub Actions, Jenkins or similar You’ll need experience with monitoring, logging and observability technologies such as Prometheus, Grafana, Loki, Open You’ll need experience designing and implementing scalable systems capable of handling high workloads Role details Work model ...

AI Native DevOps Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Permanent
within an ambitious, product-led environment. You'll work closely with Product Engineering teams and Technical Leadership to deliver cloud platforms, infrastructure, deployment pipelines, observability and developer tooling, while introducing AI-assisted and agent-driven approaches to improve engineering productivity, reliability and operational efficiency. What You'll Do Design, build … governance and compliance throughout the software delivery lifecycle Optimise Azure environments for performance, scalability, reliability and cost efficiency Support model-serving infrastructure and AI observability capabilities Lead incident response and drive continuous improvement across production environments Work closely with engineering teams delivering microservices and distributed cloud-native applications Required Experience ...

Cloud UX Design Leader – AI, Observability & Scale

Location
Greater London, England, United Kingdom
Google in London seeks a Senior UX Design Manager for Cloud to guide product strategy, reduce complexity, and design ecosystem-level experiences. You will champion user-centered design and scale design across teams. You will ...

Observability Platform Engineer — Hybrid London

Location
Greater London, England, United Kingdom
N Consulting Limited is looking for a Platform Engineer to join their team in London, UK. This hybrid position requires hands-on experience with OTel collectors, telemetry pipelines, and strong knowledge of cloud platforms. The ...

Solutions Engineer (Hybrid) — AI-Driven Observability

Location
Greater London, England, United Kingdom
Riverbed is seeking a Solutions Engineer in London, operating in a hybrid role. You will act as the technical expert during the sales process, delivering demonstrations and proofs of concept while collaborating with the Sales ...

Site Reliability Engineer: Automation & Observability

Location
Greater London, England, United Kingdom
Apple Inc. is seeking a Site Reliability Engineer in London to join the Apple Services Engineering team. You will help sustain large-scale services powering the App Store, Apple Music, TV, Podcasts and Books for ...

AI-Powered Observability Tech Lead

Location
Greater London, England, United Kingdom
Cisco’s Collaboration Technology Group in London seeks a Technical Leader to drive architectural vision and implement an AI-powered Production Intelligence platform. You will blend Site Reliability Engineering with agentic AI to improve monitoring ...

Product Manager, XDR & Observability Intelligence

Location
Greater London, England, United Kingdom
Cato Networks is seeking a Product Manager to define and build the intelligence layer of our SASE platform. You will own how Cato converts unified network and security telemetry into high-confidence detections, actionable insights ...

DevOps Engineer

Location
Harrow, England, United Kingdom
automation, platforms, and tooling that allow development teams to deploy and manage applications reliably and securely. Key areas include IaC, CI/CD, Kubernetes, observability, AWS services, and developer self-service. The engineer will collaborate closely with global cloud, infrastructure, security, and application development teams. As our cloud and platform … pipelines using GitHub Actions and related tooling. Manage and improve Kubernetes platforms, GitOps workflows, Helm charts, and application onboarding. Develop monitoring, logging, alerting, and observability capabilities using Grafana, Prometheus, DataDog, Loki, and CloudWatch. Support cloud networking and connectivity across AWS accounts, regions, on-prem environments, and other cloud platforms ...

Lead AI Software Engineer - TRP Labs London

Location
Greater London, England, United Kingdom
Guide the development of reusable AI capabilities and shared platform components across areas such as agent orchestration, tool use, retrieval‐based systems, evaluation frameworks, observability, and guardrails. Champion engineering excellence through strong software design, code quality, automated testing, continuous integration, and continuous delivery practices. Ensure AI systems are built with … measurable quality, production readiness, operational observability, and appropriate safety controls. Oversee technical debt and drive continuous improvement across AI platforms, services, and development standards. Identify and pursue opportunities to apply AI in ways that accelerate workflows, improve decision‐making, and create scalable business impact across the firm. Present and demonstrate ...

Lead Software Engineer

Location
Greater London, England, United Kingdom
stack technical leadership role with responsibility for the platform end-to-end, including user interfaces, APIs, backend services, integrations, data processing, cloud infrastructure, security, observability, and operational support. The role combines software engineering, solution architecture, operational ownership, and team leadership. You will work closely with Product, Operations, Compliance, and Architecture … PRIIP Cloud platform in production. Support the team's responsibility for production incidents, customer issues, operational troubleshooting, and service restoration activities. Drive monitoring, observability, alerting, incident management, capacity planning, and service improvement initiatives. Ensure operational processes, runbooks, support procedures, and resilience capabilities are maintained and continuously improved. Work closely with ...

Senior Software Engineer in Test (SET) New London

Location
Greater London, England, United Kingdom
performing engineering teams. In this role, you’ll define and drive quality strategy forplatform and infrastructure-level products — from container orchestration and microservices to observability tooling and CI/CD pipelines.This is a hands‐on engineering position within cross‐functional teams where quality iseveryone’s responsibility but you’ll lead … platform behavesreliably under real-world conditions Be a Technical Leader in Quality Engineering Establish standards and practices for testing distributed, event-driven systems Enable observability-driven debugging by working closely with platform and service teams Automate validation of operational characteristics like availability, latency, throughput and recoverability Contribute to security posture ...

Senior Platform Engineer /Manager London

Location
Greater London, England, United Kingdom
platforms supporting analytical and risk workloads. Developing orchestration frameworks for batch and on-demand calculations. Managing containerised environments and distributed compute services. Implementing monitoring, observability and operational tooling. Building resilient services with fault tolerance, recovery and retry capabilities. Creating CI/CD pipelines and deployment automation. Ensuring security, governance … orchestration, workflow engines and message queues. Strong SQL and data integration experience. Experience building highly available, production-grade systems. Understanding of monitoring, observability and operational support. Strong stakeholder engagement and communication skills. Desirable Experience Risk systems, portfolio analytics or financial services platforms. Large-scale batch processing or calculation-heavy environments. ...

Platform Engineer

Location
Greater London, England, United Kingdom
overhead. Support and enhance CI/CD platforms and engineering workflows, including the safe promotion of infrastructure and platform changes. Improve platform reliability, security, observability, performance and operational excellence. Partner with software engineers, quantitative developers, research teams, Technology Operations and security colleagues to deliver secure, scalable and resilient platforms. Education … experience with AWS and cloud‐native infrastructure. Experience with Kubernetes, including Amazon EKS, and Docker or other container runtimes. Experience with monitoring and observability tools such as Grafana, Prometheus, Loki or Datadog. Experience with artifact‐management platforms such as JFrog Artifactory. Knowledge of AWS Batch, AWS Step Functions, AWS Identity ...

Senior TypeScript Engineer

Location
Greater London, England, United Kingdom
reliable solutions. Use AI tools to accelerate coding, debugging, testing, research, and documentation, while validating outputs carefully and applying sound judgment. Strengthen service reliability, observability, and engineering quality by improving monitoring, incident response, testing, and development practices. The candidate 5 to 8 years of professional software engineering experience delivering production … experience supporting or mentoring other engineers within a product engineering team. Familiarity with infrastructure as code, for example CDK or Terraform. Exposure to observability tooling, incident response, or production monitoring practices. Experience working with React when contributing to end‐to‐end product delivery. What’s in it for me? They ...

Senior MLOps Engineer 201043

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
deployment, inference, and retraining workflows. Implement infrastructure-as-code solutions using tools such as Terraform, Bicep, or similar technologies. Establish monitoring, logging, alerting, and observability frameworks to ensure reliability and performance. Support scalable deployment across multiple client environments while maintaining security and operational excellence. Collaborate with data scientists, engineers … pipelines using Azure DevOps or similar tools. Knowledge of infrastructure-as-code approaches using Terraform, Bicep, Pulumi, or equivalent technologies. Experience implementing monitoring and observability solutions using tools such as Prometheus, Grafana, or similar. Familiarity with orchestration platforms including Dagster, Airflow, Prefect, or related technologies. Desirable experience includes: Model serving ...

Lead DevSecOps Engineer

Location
Greater London, England, United Kingdom
DevSecOps capabilities that help engineers build, release, and operate services across our multi-cloud platform. You will shape CI/CD, Infrastructure as Code, observability, reliability, cloud governance, and secure engineering across Microsoft Azure and Google Cloud Platform (GCP). You will also manage cloud-vendor service reviews, security advisories … Evolve CI/CD pipelines and reusable workflows that enable fast, reliable, and repeatable testing and deployment, with automated quality and security controls Build observability and operational readiness through monitoring, logging, tracing, runbooks, and incident response, improving the reliability and performance of production services Define and track service reliability objectives ...

Software Engineering Manager

Hiring Organisation
Halian Technology Limited
Location
Central London, London, United Kingdom
Employment Type
Permanent
making sound architectural and design decisions. Champion software quality, security, performance, and operational excellence. Encourage modern engineering practices, including CI/CD, automated testing, observability, and cloud-native development. Stakeholder Management Build strong relationships with business and technology stakeholders. Communicate progress, risks, and dependencies effectively. Align engineering activities with organisational … would be beneficial: .NET, Java, Python, or Node.js, React Microservices architecture RESTful APIs Kubernetes and Docker AWS, Azure, or GCP CI/CD tooling Observability and monitoring platforms Modern data platforms and event-driven architectures There is a 2 - 3 stage interview process, with interview slots now available with ...

Senior Platform Engineer /Manager London

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £115000/annum
platforms supporting analytical and risk workloads. Developing orchestration frameworks for batch and on-demand calculations. Managing containerised environments and distributed compute services. Implementing monitoring, observability and operational tooling. Building resilient services with fault tolerance, recovery and retry capabilities. Creating CI/CD pipelines and deployment automation. Ensuring security, governance … orchestration, workflow engines and message queues. Strong SQL and data integration experience. Experience building highly available, production-grade systems. Understanding of monitoring, observability and operational support. Strong stakeholder engagement and communication skills. Desirable Experience Risk systems, portfolio analytics or financial services platforms. Large-scale batch processing or calculation-heavy environments. ...