1,876 to 1,900 of 2,452 Observability Jobs in London

Software Engineer, Cloud Infrastructure

Location
Greater London, England, United Kingdom
high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable … solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that ...

Senior Software Engineer (GO)

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
services using Golang within distributed and microservices-based architectures. Build enterprise-grade AI platform capabilities supporting LLMOps, model deployment, inference services, prompt management, model observability, and governance. Engineer reusable backend blueprints and reference architectures that serve as standardized patterns across the organization. Develop globally scalable systems capable of handling high … highly available data services leveraging MongoDB and other enterprise data platforms. Optimize data flows, event processing, and backend performance for large-scale distributed systems. Observability & Reliability Implement comprehensive observability solutions using: PrometheusOpenTelemetryGrafanaEstablish monitoring, tracing, logging, alerting, and SRE best practices. Drive operational excellence through performance tuning, reliability engineering, and proactive ...

Lead Observability Engineer

Hiring Organisation
Tria
Location
London, United Kingdom
Employment Type
Contract
Location: London, onsite 3 days per week (Sheffield as an alternative) Rate: £tbd/day inside IR35 Duration: 6 months+ Are you a Senior Observability Engineer/SRE Lead, with demonstrable experience of assessing and defining observability and monitoring roadmaps within enterprise scale environments? If so, apply now for this … contract opportunity. The Lead Observability Engineer/SRE Lead will be required to assess a complex hybrid estate, understand how services, platforms, infrastructure and networks should be monitored, and work across multiple internal teams, partners and suppliers to build a consolidated view of existing telemetry, monitoring and alerting capabilities. ...

Staff Software Engineer, Observability & Profiling

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the roleAnthropic is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at Anthropic depends on—from metrics … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Cloud Advisory Architect

Location
Greater London, England, United Kingdom
where GenAI and Agentic play a role. Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps. Manage & Mentor: Lead teams of architects and engineers, providing technical coaching, career counselling, performance management, and coaching ...

AVP, Observability & SRE Engineer

Location
Greater London, England, United Kingdom
Citi is seeking a Site Reliability Engineer - Assistant Vice President in London to drive end‐to‐end observability, migrate legacy monitoring to Google Cloud Observability and Grafana, and implement OpenTelemetry instrumentation across OpenShift/Kubernetes environments. The role emphasizes hands‐on deployment, automation (Ansible/Terraform), and collaboration with application ...

Staff Software Engineer, Observability & Profiling

Location
Greater London, England, United Kingdom
policy experts, and business leaders working together to build beneficial AI systems. About the role the company is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at the company depends on—from … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Staff Software Engineer, Observability & Profiling

Location
Greater London, England, United Kingdom
engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at Anthropic depends on—from metrics … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Lead Site Reliability Engineer (Dynatrace)

Location
London, United Kingdom
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. This is not a role for someone who has simply used Dynatrace dashboards. We're looking for an engineer who has been involved in the implementation, configuration … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners Admin
Location
London, UK
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. Increase your chances of an interview by reading the following overview of this role before making an application. This is not a role for someone … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Machine Learning Operations Engineer

Location
Greater London, England, United Kingdom
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services You’ll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams This is not a research role. … model metadata and reproducibility across research and production Build reusable tooling and platform capabilities that support multiple models and engineering teams Model serving and observability Deploy and operate batch and online inference services in containerised cloud environments Define and meet availability, latency, throughput and recovery objectives for ML services Monitor ...

SRE Engineer

Location
Greater London, England, United Kingdom
automate the deployment of our software Automate the provisioning and management of our infrastructure using Infrastructure as Code (IaC) tools Define, implement, and maintain observability solutions for our applications to ensure we can proactively detect system degradation, easily understand system state, and quickly diagnose issues Diagnose and resolve production issues … must have and one should be good at coding in terraform CICD Tools hands on : Jenkins , GitHub , GitHub Actions, Cloud Deployment pipelines Observability Tools - Splunk//Graphana/Datadog and Distributed Tracing, ELF, & Dynatrace Problem-Solving : Proven ability to troubleshoot complex issues in distributed systems and debug problems effectively. ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets … operational tooling to reduce manual processes Requirements Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices Deep understanding ...

Senior Software Engineer-AI

Location
Greater London, England, United Kingdom
operation with limited supervision Hands‐on experience with cloud-native technologies, serverless applications, event‐driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade‐offs Experience mentoring engineers … operation with limited supervision Hands‐on experience with cloud-native technologies, serverless applications, event‐driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade‐offs Experience mentoring engineers ...

Senior Database Platform Engineer

Location
Greater London, England, United Kingdom
services and modern lakehouse architectures.This is a hands-on engineering role. You'll be troubleshooting performance issues, validating recovery strategies, automating operational processes, improving observability and helping shape the future of our database estate.You will act as the team's database SME, working closely with Platform Engineers, Data Engineers … recovery and disaster recovery capabilitiesOwn restore testing and recovery readiness across critical platformsSupport high availability solutions and service resilience initiativesImplement proactive monitoring, alerting and observability for database servicesParticipate in incident response, problem management and post-incident reviewsDrive continual improvement through automation, root cause analysis and operational learningReduce operational toil through ...

Lead Databricks Engineer

Hiring Organisation
EPAM Systems
Location
London, UK
Employment Type
Full-time
guide developers, establishing best practices for data engineering on Azure and DatabricksParticipate in technical decision-making and review designs for scalability and maintainabilityImplement observability for critical pipelines and maintain quality assurance standardsEngage directly with stakeholders to ensure alignment on strategic platform initiativesMinimum 8+ years of experience in data engineering, including … demonstrated SQL performance optimization experienceDeep understanding of Azure Data Platform services and cloud-native engineering patternsProven experience implementing data governance, quality management and observability frameworksAbility to manage complex data pipelines for batch and streaming use cases in mission-critical environmentsStrong communication and leadership skills to work with distributed teams ...

Senior Observability Solution Architect – Pre-Sales

Location
Greater London, England, United Kingdom
leading observability platform in Greater London is seeking an experienced Solution Engineer to join their team. This role involves collaborating with account executives on technical sales cycles, delivering impactful presentations, and overseeing technical aspects of the process. The ideal candidate will have a minimum of 5 years in a customer ...

Staff Analytics Platform Engineer

Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Platform Engineer - Common Platform

Location
Greater London, England, United Kingdom
develop automation and platform tooling Designing and implementing AWS-native infrastructure and services Building and maintaining standard CI/CD pipelines and workflows Supporting observability and monitoring capabilities across the platform Managing Kubernetes upgrades and platform improvements Supporting production incidents and resolving complex infrastructure issues Reviewing and approving infrastructure … design Enterprise security and governance Compliance frameworks and organisational policies Implementing platform changes across distributed teams Cloud cost optimisation Helm and Kubernetes management tooling Observability and monitoring You’ll be someone who enjoys building platforms and solving engineering problems rather than simply maintaining existing infrastructure. Strong communication and collaboration skills ...

Junior Azure Engineer

Hiring Organisation
Computacenter
Location
London, United Kingdom
Salary
£ 60 K
Services, and containersImplement security and governance controls (RBAC, Azure Policy, Management Groups)Build and support landing zones and foundational cloud environmentsManage monitoring and observability using Azure Monitor, Log Analytics, and alertsSupport CI/CD pipelines using Azure DevOps or GitHub ActionsAutomate operational tasks using PowerShell, Azure CLI, or FunctionsMaintain … Hands-on experience with Infrastructure-as-Code (Bicep, ARM, or Terraform)Solid understanding of Azure networking and cloud architecture principlesExperience with monitoring, logging, and observability toolsAbility to troubleshoot and resolve complex cloud issuesExperience with automation and scripting (PowerShell, Azure CLI)Strong collaboration and communication skillsDesirableExperience with Azure Kubernetes Service ...

Network Automation Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
Python, Ansible, Terraform and Jinja2Integrating network automation into CI/CD pipelines for reliable, repeatable deploymentsCreating APIs and self-service tooling for engineering teamsImplementing observability and telemetry solutions for performance and reliabilityPartnering with network, platform and security teams to deliver resilient, scalable systemsContributing to incident response and production reliabilityOn-call … tools such as Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude CodeFamiliarity with Docker and KubernetesExposure to monitoring, observability or telemetry in distributed systemsPragmatic problem solver who can operate in ambiguity and take ownershipComfortable working in collaborative, fast-paced engineering teamsDeep understanding of networking ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

Senior Software Engineer-AI

Hiring Organisation
Moody's Corporation
Location
London, UK
Employment Type
Full-time
ongoing operation with limited supervisionHands-on experience with cloud-native technologies, serverless applications, event-driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelinesSolid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade-offsExperience mentoring engineers through code … integrationContribute to technical designs, participate in design reviews, and identify risks, constraints, trade-offs, and alternative approachesMaintain engineering excellence through automated testing, code reviews, observability, monitoring, alerting, operational readiness, and participation in on-call supportApply machine learning operations practices, including prompt versioning, automated evaluation, deployment pipelines, monitoring, and production issue ...

Network Automation Engineer

Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self‐service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with network, platform and security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude Code Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Deep understanding ...

Platform Specialist - PLS

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
health monitoring Collaborate with compute, networking, and application teams to ensure storage solutions meet performance and reliability requirements Implement and improve storage-related observability, alerting, and incident response processes Evaluate and integrate new storage technologies, including cloud-based storage services (AWS S3, EBS, EFS, GCP Cloud Storage, Filestore, etc.) Ensure … language (Go, Rust, etc.) for automation and tooling development Experience with modern software development practices: version control, agile development, CI/CD Experience with observability in distributed systems (e.g., Elasticsearch, Logstash, Kibana, Datadog, Prometheus, Grafana) Experience working with cloud storage services across various cloud providers (AWS and GCP) Bachelor ...