126 to 150 of 2,214 Observability Jobs in London

Senior Data Engineer

Location
Greater London, England, United Kingdom
Unity Catalog, cluster policies, permissions, and environment separation. Work with architects to evolve Databricks platform patterns, security, and performance optimisation. Implement monitoring, logging, and observability across pipelines and clusters. Contribute to standards for versioning, DevOps, CI/CD, and testing for Databricks notebooks, workflows, and SQL assets. Support the rollout ...

Principal Site Reliability Engineer

Location
Greater London, England, United Kingdom
hold across environments Drive automation of infrastructure provisioning and configuration management using Terraform and related IaC tooling Establish and maintain comprehensive monitoring, alerting, and observability practices Cross-train and mentor other engineers, with the explicit goal of broadening GCP and multi-cloud capability across the team Partner closely with development ...

QA Automation Engineer / Consultant

Location
Greater London, England, United Kingdom
using technologies such as Apache ActiveMQ Artemis. Performance, load and stress testing using JMeter, k6 or similar tools. Experience using Splunk, Grafana or comparable observability tools. Experience testing applications hosted in Microsoft Azure and working with containerised environments, including Kubernetes. Browser automation using Playwright or WebdriverIO. Experience with accessibility testing ...

Senior Platform Engineer IRC296090

Location
Greater London, England, United Kingdom
chart/module systems (Helm, Terraform modules) Background in developer experience research — understanding how engineers consume platform tooling and designing for adoption Experience with observability and monitoring (OpenTelemetry, Grafana, Datadog) — particularly instrumenting developer workflows Experience in financial services or similarly regulated environments Job responsibilities Design and build reusable CI/ ...

Data Platform Engineer

Hiring Organisation
Quilter
Location
London, UK
Employment Type
Full-time
cloud data platforms through proactive monitoring, alerting, and incident management. Drive continuous improvement in deployment pipelines, development workflows, and engineering productivity. Support platform observability through logging, metrics, and tracing to enable rapid troubleshooting and insights. About YouQualifications Degree in a technical discipline (e.g. computer science, engineering, maths, physics) or evidence ...

Platform Engineering Manager (Cloud Foundations) London, UK

Location
Greater London, England, United Kingdom
recover predictably. The platform is reliable, incident detection and resolution are smooth due to proactive monitoring, well‐maintained alerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategy is defined and operational, working across the business with validated requirements and outcomes. Incident management processes ...

Platform Engineering Manager (Cloud Foundations)

Hiring Organisation
Rightmove
Location
London, UK
Employment Type
Full-time
Cloud platform and key services are reliable, incident detection and resolution are smooth due to proactive monitoring, well‐maintained alerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategy is defined and operational, working across the business with validated requirements and outcomes. Backup, restore ...

Platform Engineering Manager (Cloud Foundations)

Location
Greater London, England, United Kingdom
motivations. Platform Engineering, Operations&Reliability Cloudplatform and keyservices arereliable,incident detection and resolutionaresmooth due to proactive monitoring, well‐maintainedalerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategyisdefined and operational,working across the business with validated requirements and outcomes. Backup,restoreand resilience mechanisms are regularly ...

Technical Site Reliability Engineer

Hiring Organisation
Anduril Industries
Location
London, UK
Employment Type
Full-time
building automated validation pipelines, not just running them. Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).On-prem and cloud deployment experience. Observability tooling: Prometheus, Grafana, ELK, or equivalent. Linux systems administration depth; comfort in mixed Linux/Windows environments. Prior work in a defense, aerospace, or classified ...

Staff Infrastructure Engineer (GCP) - Engine by Starling

Location
Greater London, England, United Kingdom
Dedicated/Partner Interconnect Experience with Workload Identity and Workload Identity Federation for keyless authentication of workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity ...

Senior Infrastructure Engineer (GCP) - Engine by Starling

Hiring Organisation
Starling Bank
Location
London, UK
Employment Type
Full-time
Cloud VPN and Dedicated/Partner InterconnectExperience with Workload Identity and Workload Identity Federation for keyless authentication of workloads and CI/CDExperience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana)Experience setting up Google Workspace/Google Cloud ...

SC / NPPV3 DevOps Engineer - Azure

Location
Greater London, England, United Kingdom
using Docker, Kubernetes/AKS and Helm. Develop secure cloud landing zones, governance and infrastructure aligned with security and compliance requirements. Implement monitoring and observability using Azure Monitor, Log Analytics, Application Insights, Prometheus and Grafana. Automate infrastructure and operational processes using Python, PowerShell and Bash. Troubleshoot complex production issues ...

Tech Lead - Contract

Location
Greater London, England, United Kingdom
GitHub. Engineering quality & operational excellence – strong understanding of automated testing, code quality and security best practices, with experience troubleshooting distributed systems and using observability to improve performance and reliability. Technical leadership & collaboration – able to contribute to and influence technical decisions, communicate effectively with technical and non-technical stakeholders, and mentor ...

AI Platform Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, Castle Baynard, United Kingdom
Employment Type
Permanent
Salary
£80000 - £100000/annum
platform components. Deploying infrastructure using Terraform and supporting containerised applications. Building and maintaining CI/CD pipelines using GitHub and Azure DevOps. Improving observability, monitoring and platform resilience. Supporting vector search, embedding pipelines and knowledge ingestion. Applying security and governance best practice across the AI platform. What we're looking ...

Senior Network Automation Engineer

Hiring Organisation
FBI &TMT
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £600 per day
Desirable Exposure to OpenStack or VMware environments. Experience working with infrastructure-as-code tools. Knowledge of telecommunications or networking environments. Experience with monitoring and observability platforms. Understanding of cloud-native technologies and microservices architectures. What We're Looking For We're looking for a proactive engineer who enjoys tackling technical ...

Lead DevSecOps Engineer

Hiring Organisation
CBSbutler Holdings Limited trading as CBSbutler
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£645/day inside ir35
Code using Terraform Work across Kubernetes/Docker/AWS EKS Implement secrets, identity and access management using technologies such as HashiCorp Vault Establish observability, monitoring, logging and audit controls Develop secure-by-design engineering practices aligned to MOD security requirements Coordinate DevSecOps engineers, developers, testers and infrastructure teams Create ...

Google Cloud Platform GCP Architect Developer

Hiring Organisation
Telstra Associates
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£450/day
automation frameworks on Kubernetes, including authoring reconciliation loops, admission controllers, webhooks, and custom controllers; strong understanding of Kubernetes internals, API machinery, RBAC, multi-tenancy, observability, and operational best practices for production environments. In-depth Google Cloud development experience -architecting, building, deploying, and operating production cloud-native workloads using core ...

Machine Learning Engineer (Applied AI ML)

Location
Greater London, England, United Kingdom
learning architectures such as transformers, CNNs, and autoencoders Specialism or well-researched interest in NLP Broad knowledge of MLOps tooling for versioning, reproducibility, and observability Experience monitoring, maintaining, and enhancing existing models over an extended period Extensive experience with PyTorch and related data science Python libraries such as pandas Experience ...

data engineer in data platforms

Location
Greater London, England, United Kingdom
technical advisor Set engineering standards, patterns, and best practices across teams Review designs and code, providing technical direction and mentorship Improve data quality, testing, observability, and operational excellence Требования Strong Python and SQL skills Deep experience with Spark and modern data platforms such as Databricks and Snowflake Solid understanding ...

Platform Engineer - Enterprise Platforms

Location
Greater London, England, United Kingdom
hosting Managing AWS services including S3, FSx, EBS and WorkSpaces/AppStream Developing configuration management solutions using tools such as Puppet Maintaining monitoring and observability capabilities Supporting enterprise backup and recovery platforms Building and maintaining CI/CD pipelines using GitHub Actions Supporting production incidents, problem management and platform requests ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
security, performance, maintainability, and testing. Contribute to technical standards and continuous improvement initiatives within the team. Operational Excellence & Collaboration: Maintain reliable production environments, improve observability and monitoring systems, and respond to proactive and reactive support alerts. Work collaboratively across the Technology department while mentoring and supporting other engineers. What ...

TypeScript Engineer

Location
Greater London, England, United Kingdom
these too, but support will be offered if not: Familiarity with infrastructure as code, for example CDK or Terraform. Exposure to observability tooling, incident response, or production monitoring practices. A broader understanding of LLM concepts such as tokens, embeddings, hallucinations, and safe use cases in software delivery. What ...

Senior Lead Software Engineer - Java / Python - Risk Technology Data Strategy

Location
Greater London, England, United Kingdom
debugging, and maintaining code with modern programming languages and database querying languages Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, Observability tools (Dynatrace, Splunk, Grafana), and Orchestration tools (Airflow, Temporal) Proficiency in automation, continuous delivery methods, and all aspects of the Software Development Life Cycle Advanced ...

Engineering Manager

Location
Greater London, England, United Kingdom
backend engineering, including Node.js, TypeScript, AWS, APIs, and production systems supported by SQL or NoSQL data stores. Good understanding of CI/CD, observability, incident support, and modern engineering practices such as Docker‐based workflows and infrastructure as code. Practical understanding of AI‐enabled engineering and product use cases, including ...

Senior Site Reliability Engineer (LON)

Location
Greater London, England, United Kingdom
Azure Bicep, ARM and Azure DevOps. The client is moving to Terraform, which is essential , moving from Bicep (desirable). Experience with Full Stack Observability using tools such as Grafana Stack, Log Analytics, AppInsights Excellent knowledge of DevOps processes and principles Knowledge of IT Service Management and automation ...