1,726 to 1,750 of 2,250 Observability Jobs in London

Technical Lead, Observability - London

Location
Greater London, England, United Kingdom
observability across H: the strategy, technical design and rollout. Your work will give visibility of the entire stack and surface the information the company relies on to make key decisions and quantify progress. H builds computer-use agents that take actions in browsers and desktop applications. This role covers every … responsible for ensuring our teams measure what matters, and for learning our products and services in depth. What you'd be doing Setting one observability strategy that can be applied across teams and technical stacks Owning the implementation of the systems that manage metrics, logs, traces, alerts and dashboards Defining ...

Staff Software Engineer, Observability & Profiling

Location
Greater London, England, United Kingdom
policy experts, and business leaders working together to build beneficial AI systems. About the role the company is seeking Software Engineers to join our Observability team within the Infrastructure organization. The Observability team owns the monitoring and telemetry infrastructure that every engineer and researcher at the company depends on—from … growing by orders of magnitude—and an increasing share of the hardest problems live below the application layer. We're building next-generation observability systems—high-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing and network visibility, and agentic diagnostic tools—so engineers can detect, diagnose, and resolve ...

Lead Site Reliability Engineer (Dynatrace)

Location
London, United Kingdom
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. This is not a role for someone who has simply used Dynatrace dashboards. We're looking for an engineer who has been involved in the implementation, configuration … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners Admin
Location
London, UK
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. Increase your chances of an interview by reading the following overview of this role before making an application. This is not a role for someone … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Machine Learning Operations Engineer

Location
Greater London, England, United Kingdom
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services You’ll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams This is not a research role. … model metadata and reproducibility across research and production Build reusable tooling and platform capabilities that support multiple models and engineering teams Model serving and observability Deploy and operate batch and online inference services in containerised cloud environments Define and meet availability, latency, throughput and recovery objectives for ML services Monitor ...

SRE Engineer

Location
Greater London, England, United Kingdom
automate the deployment of our software Automate the provisioning and management of our infrastructure using Infrastructure as Code (IaC) tools Define, implement, and maintain observability solutions for our applications to ensure we can proactively detect system degradation, easily understand system state, and quickly diagnose issues Diagnose and resolve production issues … must have and one should be good at coding in terraform CICD Tools hands on : Jenkins , GitHub , GitHub Actions, Cloud Deployment pipelines Observability Tools - Splunk//Graphana/Datadog and Distributed Tracing, ELF, & Dynatrace Problem-Solving : Proven ability to troubleshoot complex issues in distributed systems and debug problems effectively. ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets … operational tooling to reduce manual processes Requirements Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices Deep understanding ...

Senior Software Engineer-AI

Location
Greater London, England, United Kingdom
operation with limited supervision Hands‐on experience with cloud-native technologies, serverless applications, event‐driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade‐offs Experience mentoring engineers … operation with limited supervision Hands‐on experience with cloud-native technologies, serverless applications, event‐driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade‐offs Experience mentoring engineers ...

Senior Database Platform Engineer

Location
Greater London, England, United Kingdom
services and modern lakehouse architectures.This is a hands-on engineering role. You'll be troubleshooting performance issues, validating recovery strategies, automating operational processes, improving observability and helping shape the future of our database estate.You will act as the team's database SME, working closely with Platform Engineers, Data Engineers … recovery and disaster recovery capabilitiesOwn restore testing and recovery readiness across critical platformsSupport high availability solutions and service resilience initiativesImplement proactive monitoring, alerting and observability for database servicesParticipate in incident response, problem management and post-incident reviewsDrive continual improvement through automation, root cause analysis and operational learningReduce operational toil through ...

Senior Observability Solution Architect – Pre-Sales

Location
Greater London, England, United Kingdom
leading observability platform in Greater London is seeking an experienced Solution Engineer to join their team. This role involves collaborating with account executives on technical sales cycles, delivering impactful presentations, and overseeing technical aspects of the process. The ideal candidate will have a minimum of 5 years in a customer ...

Staff Analytics Platform Engineer

Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Platform Engineer - Common Platform

Location
Greater London, England, United Kingdom
develop automation and platform tooling Designing and implementing AWS-native infrastructure and services Building and maintaining standard CI/CD pipelines and workflows Supporting observability and monitoring capabilities across the platform Managing Kubernetes upgrades and platform improvements Supporting production incidents and resolving complex infrastructure issues Reviewing and approving infrastructure … design Enterprise security and governance Compliance frameworks and organisational policies Implementing platform changes across distributed teams Cloud cost optimisation Helm and Kubernetes management tooling Observability and monitoring You’ll be someone who enjoys building platforms and solving engineering problems rather than simply maintaining existing infrastructure. Strong communication and collaboration skills ...

Junior Azure Engineer

Hiring Organisation
Appcast
Location
London, UK
Services, and containersImplement security and governance controls (RBAC, Azure Policy, Management Groups)Build and support landing zones and foundational cloud environmentsManage monitoring and observability using Azure Monitor, Log Analytics, and alertsSupport CI/CD pipelines using Azure DevOps or GitHub ActionsAutomate operational tasks using PowerShell, Azure CLI, or FunctionsMaintain … Hands-on experience with Infrastructure-as-Code (Bicep, ARM, or Terraform)Solid understanding of Azure networking and cloud architecture principlesExperience with monitoring, logging, and observability toolsAbility to troubleshoot and resolve complex cloud issuesExperience with automation and scripting (PowerShell, Azure CLI)Strong collaboration and communication skillsDesirableExperience with Azure Kubernetes Service ...

Lead Databricks Engineer

Location
Greater London, England, United Kingdom
guide developers, establishing best practices for data engineering on Azure and DatabricksParticipate in technical decision-making and review designs for scalability and maintainabilityImplement observability for critical pipelines and maintain quality assurance standardsEngage directly with stakeholders to ensure alignment on strategic platform initiativesRequirementsMinimum 8+ years of experience in data engineering, including … demonstrated SQL performance optimization experienceDeep understanding of Azure Data Platform services and cloud-native engineering patternsProven experience implementing data governance, quality management and observability frameworksAbility to manage complex data pipelines for batch and streaming use cases in mission-critical environmentsStrong communication and leadership skills to work with distributed teams ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
pairing and code review. Help define and evolve the target state architecture for the platform: service boundaries, data models, integration patterns, and reliability/observability standards. Use AI tooling (e.g.Claude and similar) as a force multiplier to accelerate discovery, prototyping, and implementation, while maintaining a high bar for code quality … room for the right tool for specific problems. Data: Postgres as a core relational store. APIs: gRPC and HTTP‐based services. Tooling: GitOps, IaC, observability tooling. You don’t need prior experience with every technology above, but you should be comfortable learning and working across this kind of stack. What ...

Network Automation Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
Python, Ansible, Terraform and Jinja2Integrating network automation into CI/CD pipelines for reliable, repeatable deploymentsCreating APIs and self-service tooling for engineering teamsImplementing observability and telemetry solutions for performance and reliabilityPartnering with network, platform and security teams to deliver resilient, scalable systemsContributing to incident response and production reliabilityOn-call … tools such as Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude CodeFamiliarity with Docker and KubernetesExposure to monitoring, observability or telemetry in distributed systemsPragmatic problem solver who can operate in ambiguity and take ownershipComfortable working in collaborative, fast-paced engineering teamsDeep understanding of networking ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

Senior Software Engineer-AI

Hiring Organisation
Moody's Corporation
Location
London, UK
Employment Type
Full-time
ongoing operation with limited supervisionHands-on experience with cloud-native technologies, serverless applications, event-driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelinesSolid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade-offsExperience mentoring engineers through code … integrationContribute to technical designs, participate in design reviews, and identify risks, constraints, trade-offs, and alternative approachesMaintain engineering excellence through automated testing, code reviews, observability, monitoring, alerting, operational readiness, and participation in on-call supportApply machine learning operations practices, including prompt versioning, automated evaluation, deployment pipelines, monitoring, and production issue ...

Network Automation Engineer

Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self‐service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with network, platform and security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude Code Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Deep understanding ...

Platform Specialist - PLS

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
health monitoring Collaborate with compute, networking, and application teams to ensure storage solutions meet performance and reliability requirements Implement and improve storage-related observability, alerting, and incident response processes Evaluate and integrate new storage technologies, including cloud-based storage services (AWS S3, EBS, EFS, GCP Cloud Storage, Filestore, etc.) Ensure … language (Go, Rust, etc.) for automation and tooling development Experience with modern software development practices: version control, agile development, CI/CD Experience with observability in distributed systems (e.g., Elasticsearch, Logstash, Kibana, Datadog, Prometheus, Grafana) Experience working with cloud storage services across various cloud providers (AWS and GCP) Bachelor ...

Forward Deployed Agentic AI Engineer

Location
Greater London, England, United Kingdom
solutions that solve real-world business challenges. You will bring deep expertise across modern full-stack technologies, distributed systems, cloud-native architectures, and observability, combined with hands‐on experience developing enterprise‐grade AI applications. You will design, build, and deploy intelligent AI agents, copilots, and automation solutions using Anthropic Claude … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...

Senior Full Stack Engineer - Lyst Shop (11 Month FTC - Maternity Cover)

Location
Greater London, England, United Kingdom
rely heavily on experimentation to validate ideas and guide decisions. Technical Excellence: You will help maintain a high bar for code quality, testing, observability, and system reliability. You’ll contribute to architectural discussions, improve developer experience, and proactively address technical debt where needed. Team Contribution: You will mentor and support … working relationships across Product, Design, QA, Analytics, and Engineering teams while actively participating in team ceremonies and technical discussions. Technical Impact: Improve the stability, observability, and maintainability of our systems through better monitoring, resilient code, and thoughtful testing practices. Growth & Ownership: Gain confidence working across our platform and infrastructure while ...

Embedded Software Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
security best practices across the device lifecycleEdge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issuesCompliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded softwarePractical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platformsStrong problem-solving skills and the ability to work across hardware and software boundariesBonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services on AWS, Azure ...