1,801 to 1,825 of 2,452 Observability Jobs in London

Senior Database Platform Engineer

Hiring Organisation
Howden Group Holdings
Location
London, UK
Employment Type
Full-time
modern lakehouse architectures. This is a hands-on engineering role. You'll be troubleshooting performance issues, validating recovery strategies, automating operational processes, improving observability and helping shape the future of our database estate. You will act as the team's database SME, working closely with Platform Engineers, Data Engineers … recovery capabilities Own restore testing and recovery readiness across critical platforms Support high availability solutions and service resilience initiatives Implement proactive monitoring, alerting and observability for database services Participate in incident response, problem management and post-incident reviews Drive continual improvement through automation, root cause analysis and operational learning Reduce ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 70,000 Annual
automate deployments using Terraform and Ansible, and build CI/CD pipelines using GitHub Actions. You'll also work across cloud security, networking and observability, while having the opportunity to develop your technical expertise through training and professional certifications. This role would suit an experienced DevOps Engineer looking to work … private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
Camden, London, Camden Town, United Kingdom
Employment Type
Permanent
Salary
£65000 - £70000/annum + Remote + Progression
automate deployments using Terraform and Ansible, and build CI/CD pipelines using GitHub Actions. You'll also work across cloud security, networking and observability, while having the opportunity to develop your technical expertise through training and professional certifications. This role would suit an experienced DevOps Engineer looking to work … private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join ...

Lead Software Engineer - Platform

Location
Greater London, England, United Kingdom
need to scale accordingly, and our platform foundations need to support faster, more reliable delivery. You’ll be central to making that happen – from observability and developer experience to the infrastructure patterns that underpin everything we ship. What you’ll do Alongside the other Lead Engineers you’ll support … infrastructure. You’ll be responsible for the reliability, scalability, and operability of our systems. That means CI/CD pipelines, infrastructure‐as‐code, observability, incident response, and the day‐to‐day health of production. You’ll make sure we can ship with confidence and sleep at night. Shape technical direction. ...

Principal DevSecOps Engineer - LONDON

Hiring Organisation
83zero Limited
Location
Central London, London, United Kingdom
Employment Type
Permanent, Work From Home
workflows * Establish secure-by-design engineering practices and enforce security and technical standards * Lead Infrastructure as Code (IaC) practices across teams and environments * Drive observability, monitoring, logging and audit controls * Support incident response, patching, compliance reporting and technical debt remediation * Partner with developers and delivery teams to improve engineering quality … Security & Compliance: Trivy, vulnerability management, HashiCorp Vault, cert-manager * Containers & Cloud: Docker, AWS EKS, AWS IAM, S3 and network policies * Infrastructure as Code: Terraform * Observability: Grafana, Loki * Automation: Python and Bash * Experience delivering within the UK Government Digital Service (GDS) lifecycle on a public sector engagement Why join ...

platform engineer for legal AI

Location
Greater London, England, United Kingdom
environments; Implement and evolve Infrastructure-as-Code using Terraform; Collaborate with developers to improve CI/CD pipelines, deployment strategies, and developer experience; Enhance observability and reliability through alerting, monitoring, and incident response; Collaborate on cloud optimization projects to improve performance, cost efficiency, and security posture; Mentor and guide team … management; Experience with CI/CD systems such as Buildkite and GitHub Actions; Experience with relational or non-relational datastores in production; Familiarity with observability platforms such as Datadog, Prometheus, and Grafana; Familiarity with Linux systems administration, networking, and troubleshooting; Excellent communication and documentation abilities, with a focus on knowledge ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
scalability, and performance. Develop and maintain automation workflows using Ansible, Salt, and related frameworks to reduce operational toil. Build and operate monitoring, alerting, and observability dashboards using tools such as Grafana and Splunk. Proactively identify network bottlenecks, performance issues, and reliability risks, implementing long‐term fixes rather than reactive solutions. … network protocols (BGP, OSPF, EIGRP, STP, VXLAN, VPNs, QoS, MPLS, etc.). Strong experience with infrastructure automation using Ansible and Salt. Proficiency with observability tooling such as Grafana, Splunk, or equivalent. Solid understanding of SRE practices including SLIs, SLOs, error budgets, and proactive reliability. Strong troubleshooting, analytical, and performance optimization ...

ML Ops Engineer

Hiring Organisation
CMC Markets
Location
London, UK
Employment Type
Full-time
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services. You'll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams. This is not a research role. … production issues across model, application, infrastructure and critical data-dependency layers. Reliability, security and engineering qualityImprove system robustness, scalability and cost efficiency through automation, observability and infrastructure as code. Write production-grade Python for long-running services, deployment tooling and ML workflows. Establish testing, validation, release and incident-management practices ...

Junior Azure Engineer

Location
Greater London, England, United Kingdom
containers Implement security and governance controls (RBAC, Azure Policy, Management Groups) Build and support landing zones and foundational cloud environments Manage monitoring and observability using Azure Monitor, Log Analytics, and alerts Support CI/CD pipelines using Azure DevOps or GitHub Actions Automate operational tasks using PowerShell, Azure … experience with Infrastructure-as-Code (Bicep, ARM, or Terraform) Solid understanding of Azure networking and cloud architecture principles Experience with monitoring, logging, and observability tools Ability to troubleshoot and resolve complex cloud issues Experience with automation and scripting (PowerShell, Azure CLI) Strong collaboration and communication skills Experience with Azure Kubernetes ...

Platform Engineer

Hiring Organisation
Springer Nature
Location
London, United Kingdom
Salary
£ 70 K
engage on providing a unified and standardised platform by using modern and open standards. Our approach is encompassed by defined core capabilities, such as observability, continuous integration, security and storage. We therefore closely collaborate in our department so that the core capabilities are not only tightly integrated but also provide … Engineering department at Springer Nature Technology, we provide platform engineering expertise to support core capabilities and use cases across run time to databases to observability, ci/cd, and SRE. You will join a multidisciplinary team with different nationalities, backgrounds and experience levels. We are a very distributed department, sometimes ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
scalability. Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions. Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively. Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling, and improve infrastructure efficiency. Improve … Kubernetes operational experience (on-prem and AWS EKS).Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows. Deep understanding of observability and telemetry (monitoring, logging, tracing).Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD.Scripting proficiency in Python, Bash, or Go. Proven ...

Security Platform Engineer, UK Security Operations

Hiring Organisation
Google
Location
London, UK
Employment Type
Full-time
Python, Go, or Bash. Experience with Kubernetes security, including workload isolation, Role-Based Access Control (RBAC), and network policies, containerisation, orchestration, and Kubernetes observability tools (e.g., Falco, Prometheus, Grafana).Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Helm, ArgoCD).Active, or the ability to obtain, a Developed … Python, Go, or Bash. Experience with Kubernetes security, including workload isolation, Role-Based Access Control (RBAC), and network policies, containerisation, orchestration, and Kubernetes observability tools (e.g., Falco, Prometheus, Grafana).Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Helm, ArgoCD).Active, or the ability to obtain, a Developed ...

Senior Network Engineer- IP

Location
Greater London, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Software Engineer, Senior

Hiring Organisation
Infor
Location
Orpington, Greater London, UK
Employment Type
Full-time
integration patterns using modern cloud-native approaches. Evaluating and implementing emerging AI integration technologies, frameworks, and engineering practices. Helping establish engineering standards for security, observability, resiliency, and maintainability. Contributing to technical design discussions and architecture reviews. Mentoring and supporting other engineers as AI capabilities become more broadly adopted across … applications and services. Experience with cloud-native development on AWS, Azure, or similar cloud platforms. Strong understanding of software architecture, security, scalability, reliability, and observability principles. Experience working with event-driven architectures and asynchronous integration patterns. Experience working with relational databases and modern data access technologies. Experience developing reusable frameworks ...

Forward Deployed AI Engineer

Location
Greater London, England, United Kingdom
enabled systems. You’ll bring deep expertise across modern full-stack technologies (.NET, Azure, SQL, React/Angular), along with experience in distributed systems, observability, and AI tooling such as LLMs, retrieval pipelines, agentic workflows, and platforms such as Anthropic Claude. Experience designing and deploying AI agents, leveraging Model Context … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Software Engineer (Next.js Playwright)

Location
Greater London, England, United Kingdom
maintain applications within Azure. Contribute to CI/CD pipelines using GitHub Actions and Azure DevOps. Monitor application performance and reliability using modern observability tools. Collaborate Across the Business Work closely with clinicians, product managers and fellow engineers to understand user needs. Contribute ideas that improve both the product … Experience with Playwright or other automated testing frameworks. Experience working within Azure. Experience in healthcare, NHS, or regulated environments. Familiarity with application monitoring and observability tools. Experience working in a SaaS or scale-up environment. Mindset Product-minded and user-focused. Pragmatic, collaborative and delivery-oriented. Takes ownership and sees ...

DevOps Engineer

Location
Greater London, England, United Kingdom
build and operate the infrastructure that every engineering team at Invisible depends on — Kubernetes, CI/CD, identity, secrets management, networking, and observability — along with the internal tooling that allows engineers and AI coding agents to ship safely and quickly. The Platform team is small relative to the surface area … entire engineering organisation. We are looking for engineers who don’t just apply best practices in IaC, Kubernetes, CI/CD and observability, but understand the problems those practices were designed to solve. Most were built around the pace and failure modes of human engineers. Those assumptions are changing ...

Senior Software Engineer

Hiring Organisation
Daniel James Resourcing Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
engineers. Where AI enters the product itself, youll help establish the controls required to take it into production responsibly, including validation, guardrails, permissions, observability, evaluation, failure behaviour and appropriate human oversight. What you'll bring Youll be an accomplished Software Engineer who has designed and delivered complex production systems … data modelling Docker, Kubernetes and Infrastructure as Code Event-driven architecture CI/CD, automated testing and modern engineering practices Production ownership, reliability and observability Technical leadership and mentoring other engineers The underlying brief prioritises strong C#/.NET experience but can consider engineers from another modern backend language ...

Applied AI Engineering Lead - VP, Markets Operations

Location
Greater London, England, United Kingdom
information Build and enhance robust AI services and infrastructure using modern engineering practices, including APIs, event‐driven patterns, CI/CD, Infrastructure-as-Code, observability, and automated testing Partner with AI researchers, data scientists, and software engineers to translate emerging AI capabilities into practical, reliable, and compliant enterprise applications Establish … data engineering concepts, ETL and data pipelines, structured and unstructured data, and integration with enterprise data platforms Experience with CI/CD, automated testing, observability, production monitoring, and operational readiness practices Familiarity with Infrastructure-as-Code solutions such as Terraform and cloud or container‐based deployment patterns Working knowledge ...

Senior Associate, Full-Stack Engineer

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … programming concepts and microservicesProficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker).Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real‐time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Sales Engineer, EMEA

Location
City Of London, England, United Kingdom
business value Build and deliver tailored demos and POCs using Python, Airflow, and related tools Guide customers through modern DataOps strategies - from orchestration to observability and automation Influence product feedback loops, collaborate cross-functionally, and contribute to go-to-market strategy What you bring to the role: You’re technically … Architect, etc.) Strong communication skills across technical and business stakeholders Startup-friendly mindset: proactive, adaptable, fast-moving Bonus points if you have: Experience with observability, Kubernetes, or helping teams scale modern data stacks. Why Astronomer? We’re defining the DataOps layer of the modern data stack - not just orchestration Airflow ...

Global Platform Team Lead & Senior Director - Identity Security

Hiring Organisation
The Boston Consulting Group
Location
London, UK
Employment Type
Full-time
scale. Align identity security strategies with broader digital transformation, BCG X product delivery, and AI/ML platform initiatives. Establish a global identity observability and telemetry capability, enabling real-time visibility into identity risks, access anomalies, and compliance posture. Directory ServicesOversee the strategy and execution of BCG's enterprise directory … Partner with HR, Legal, and ISRM to ensure identity governance processes remain compliant with evolving regulatory and audit requirements. Secure Data (Secrets Management, DLP & Observability)Lead the engineering and operations of BCG's secrets management platform, ensuring credentials, API keys, certificates, and tokens are centrally vaulted, rotated, and audited — with ...

Core AI Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
running centralised MCP servers, ensuring secure, reliable access to enterprise tools and dataOwning platform reliability, performance and scalability across Kubernetes-based infrastructure, including observability, capacity planning and incident responseBuilding self-service tooling and APIs to enable teams to provision and consume AI infrastructure independentlyIntegrating platform services with existing technology stacks … similar platform servicesFamiliarity with sandboxing and workload isolation technologiesExperience in quantitative finance or low-latency systemsAWS experience particularly in hybrid environmentsExperience with observability tooling such as Prometheus, Grafana or OpenTelemetryContributions to open-source projects in relevant domainsWhy join us? Highly competitive compensation plus annual discretionary bonusLunch provided (via Just ...

Software Engineer

Hiring Organisation
CISCO Systems
Location
London, UK
Employment Type
Full-time
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality: Own end-to-end configuration quality … encourage applications without these: Hands-on with Helm or KustomizeExperience with GitOps (e.g., Argo CD)Knowledge of secrets management (e.g., HashiCorp Vault)Experience with observability (metrics/logs/tracing)CollabHiringWhy Cisco? At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations ...