726 to 750 of 2,338 Observability Jobs in London

Platform Engineering Director - SaaS & Observability

Location
Greater London, England, United Kingdom
ITRS is seeking an experienced Director of Platform Engineering to lead our Analytics SaaS platform, focusing on Kubernetes-based services, cloud-native operations, and reliable service outcomes. You will guide hands-on SaaS engineering, set ...

Senior Network SRE: Automation, Reliability & Observability

Location
Greater London, England, United Kingdom
A leading IT solutions provider in London is seeking a Senior Network Site Reliability Engineer (SRE) with extensive experience in network engineering and automation. The ideal candidate will design and maintain high-availability network solutions ...

Site Reliability Engineer, Studios

Location
Uxbridge, England, United Kingdom
embed reliability engineering practices across systems that support live, business‐critical environments. The successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders. Key Responsibilities And Accountabilities Design, build … platform services across on‐premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering‐led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators. Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical ...

Site Reliability Engineer, Studios

Location
Greater London, England, United Kingdom
embed reliability engineering practices across systems that support live, business-critical environments. The successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders. Key Responsibilities and Accountabilities Design, build … platform services across on-premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering-led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators. Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
governance, automation, security controls, and developer experience established with our GCP platform. Develop and enhance platform services, self‐service capabilities, CI/CD tooling, observability solutions, and automation frameworks that improve developer productivity and operational effectiveness. Contribute to our AI transformation by building the platform capabilities, tooling, operational patterns … teams to define technical direction, establish engineering standards, and ensure platform capabilities meet current and future business requirements. Drive continuous improvement across platform reliability, observability, automation, security, cost efficiency, and developer experience through data‐driven engineering practices. Mentor and support engineers across the organisation, promoting platform engineering best practices, technical ...

Senior DevOps Engineer (DV)

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £750 per day + Inside IR-35
/CD pipelines to support rapid and reliable delivery. Champion DevOps culture, automation and engineering best practice across delivery teams. Improve platform resilience, observability, security and operational performance. Work closely with software engineers, architects and senior stakeholders to deliver platform capabilities. Support containerised workloads and cloud-native deployments. Lead technical … similar. Containerisation and orchestration experience, including Kubernetes and Docker. Strong scripting and automation skills using Python, Bash or similar. Experience implementing monitoring, logging and observability solutions. Excellent communication skills with the ability to engage technical and non-technical stakeholders. Experience operating in secure or regulated environments. Desirable Experience Previous Defence ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
team. Design, build and operate a secure, resilient and scalable Azure infrastructure estate, including the associated CI/CD, infrastructure-as-code, container and observability capabilities. Establish and evolve engineering standards, reference architecture, guardrails and reusable platform services, promoting consistency, automation, security, maintainability and operational excellence. Lead the design … cloud and/or on-premises infrastructure, including the development of reusable, maintainable infrastructure-as-code patterns. Experience implementing or operating monitoring, logging and observability capabilities using tools such as Splunk, Prometheus, Grafana or equivalent. Strong scripting and automation capabilities using PowerShell, Bash and/or similar languages, together with ...

Senior Full Stack/Native Cloud Engineer (Vice President)

Hiring Organisation
Jefferies Financial Group
Location
London, UK
Employment Type
Full-time
technical design and architecture through to production deployment, monitoring, support, and continuous improvement. Champion engineering best practices across code quality, testing, CI/CD, observability, maintainability, and operational resilience. Mentor and guide junior and mid-level engineers, setting high technical standards through code reviews, design reviews, and hands-on technical … reliable, maintainable, and responsive UI delivery. Strong understanding of software engineering best practices, including automated testing, clean code, code reviews, CI/CD, observability, and secure development practices. Experience deploying and operating production-grade systems in cloud or hybrid-cloud environments. Demonstrated ability to lead technical delivery, mentor engineers ...

Platform Staff Engineer - UK

Hiring Organisation
DISCO
Location
London, UK
Employment Type
Full-time
flexible, self-service artifact generation. Researches and evaluates file processing software based on fidelity, reliability, and performance criteria. Models file processing outcomes within an observability framework to provide insights for engineering teams and business leaders. Develops engineering systems that facilitate rapid and automatic evaluation of file processors. Data Management: Crafts … projects, showing an ability to work effectively both independently and as part of a team. Demonstrated expertise in designing, implementing, and maintaining (through operational observability) high-availability, high-performance, distributed data processing systems. Experience with gRPC and Protocol Buffers for efficient, language-agnostic service-to-service communication. Proven ability ...

Senior AWS DevOps Engineer

Location
Greater London, England, United Kingdom
services such as ECS and EKS Work with architects and security specialists on cloud architecture, networking, identity, security and integration requirements Implement effective monitoring, observability, logging and alerting to support resilient services Provide technical leadership during complex troubleshooting and resolution of infrastructure and deployment issues Engage with clients and technical … operating containerised workloads in cloud environments Good understanding of DevSecOps, OWASP principles and security-by-design Experience designing and implementing monitoring, logging and observability solutions Strong scripting skills using Python, Bash, PowerShell or similar Experience making technical design decisions and providing engineering leadership within multidisciplinary teams Strong troubleshooting and problem ...

Google Cloud Platform Engineer

Hiring Organisation
Damia Group Ltd
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£375 - £450 per day
Definitions (CRDs). Experience with technologies such as Kubebuilder, Operator SDK or equivalent . Strong understanding of Kubernetes internals, API machinery, RBAC, multi-tenancy, observability and production operational practices. Proven experience building cloud-native control planes and automation frameworks on Kubernetes, including reconciliation loops, admission controllers, webhooks and custom controllers. … trade-offs to both technical and non-technical audiences. Desirable Skills Advanced Infrastructure-as-Code experience, particularly Terraform, CloudFormation or CDK . Strong observability and monitoring experience with technologies such as Prometheus, Grafana, Datadog or CloudWatch . Relevant Google Cloud certifications , such as Developer, Solutions Architect or Security. Experience implementing ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive … enterprise or production environments. Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication ...

Senior Engineer

Location
Greater London, England, United Kingdom
roadmap. Design, build and operate a secure, resilient and scalable Azure infrastructure estate, including the associated CI/CD, infrastructure-as-code, container and observability capabilities that support application delivery. Establish and evolve engineering standards, reference architectures, guardrails and reusable platform services, promoting consistency, automation, security, maintainability and operational excellence. … cloud and/or on-premise infrastructure, including the development of reusable, maintainable infrastructure-as-code patterns. Experience implementing or operating monitoring, logging and observability capabilities using tools such as Splunk, Prometheus, Grafana or equivalent, with the ability to use operational data to improve reliability and performance. Strong scripting ...

Lead Platform Engineer

Location
Greater London, England, United Kingdom
Platform Engineering & Cloud Technology This role offers the opportunity to work across a modern technology landscape, including cloud platforms, infrastructure as code, DevOps toolchains, observability, security, and automation technologies. You will help shape platform capabilities that accelerate software delivery, improve reliability, and enhance the developer experience. Working alongside highly skilled … design and platform governance principles. Experience with AWS and multi‐cloud architectures. Experience operating and governing Kubernetes platforms at enterprise scale. Knowledge of modern observability and monitoring tools such as Grafana, Prometheus, OpenTelemetry and Azure Monitor. Experience designing and supporting hybrid cloud environments. Familiarity with Internal Developer Platforms and self ...

VodafoneThree - Senior SRE

Location
Greater London, England, United Kingdom
reliability of VodafoneThree's customer-facing platforms and developer experience. The role combines strong software engineering with operational excellence, taking ownership of platform reliability, observability, automation and incident reduction. It works closely with software engineers, architects, security teams and product teams so services meet availability, performance and resilience targets while … platforms. – Essential Strong Infrastructure as Code experience using Terraform and experience building CI/CD pipelines and deployment automation. – Essential Strong understanding of monitoring, observability, telemetry, SLOs, SLIs and error budgets. – Essential Experience participating in on-call, incident management and root cause analysis processes. – Essential Ability to write production-quality ...

Data Technical Lead

Location
Greater London, England, United Kingdom
analyticaland domain-oriented modelling Data governance:Catalogue, metadata, lineage, quality, security,privacyand access controls Data products:Reusable, discoverable and well-governed dataproductsand marketplaces DataOpsand observability:Testing, monitoring, operationalcontrolsand platform reliability Platform engineering:Terraform, CloudFormation, Azure Bicep and infrastructure-as-code CI/CD:GitHub Actions, Azure DevOps,Jenkinsand equivalent tooling … serving. Applying strong software engineering practices to data platforms, including testing, CI/CDand infrastructure-as-code. Establishing effective approaches to data quality, observability, governance,metadataand security. You can lead data platform modernisation and migration, including coexistence,cutoverand decommissioning. Making pragmatic technology choices and understanding the trade-offs between platform ...

Production Engineer

Hiring Organisation
Liquidnet
Location
London, UK
Employment Type
Full-time
Group OverviewThe TP ICAP Group is a world leading provider of market infrastructure. Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data ...

Lead Data Engineer

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
Shape how hundreds of thousands of UK investors use data to make confident, informed investment decisions. Join a team building modern, cloud-native data platforms that enable analytics, regulatory reporting, and data-driven products at ...

Senior Platform & Automation Engineer (Python/IaC) - Hybrid

Location
City Of London, England, United Kingdom
automation across the stack through unified CI/CD pipelines. The role emphasizes IaC with Terraform and Ansible, Python-based automation, and provisioning across observability, security, and container ecosystems, with a hybrid work arrangement. #J-18808-Ljbffr ...

Senior DB Platform Engineer - Remote, Cloud & Automation

Location
City Of London, England, United Kingdom
toward PostgreSQL and cloud-native services, and partner with Platform, Data and Operations teams to deliver resilient data services. The role emphasizes automation, improved observability, incident response and continual improvement as we evolve our data platforms and lakehouse #J-18808-Ljbffr ...

Senior Kubernetes & GitOps Reliability Engineer

Location
Greater London, England, United Kingdom
production in a hybrid work model. You will drive GitOps workflows with Argo CD, implement canary releases, manage secrets with HashiCorp Vault, and improve observability with Prometheus and Grafana. #J-18808-Ljbffr ...

Senior Azure Platform Engineer - Cloud-Native & GitOps

Location
Greater London, England, United Kingdom
GitOps, delivering scalable multi-tenant environments. The position is hybrid (2 days in the London office) with responsibilities spanning CI/CD pipelines, security, observability and DevSecOps practices to improve reliability and performance. #J-18808-Ljbffr ...

Senior GCP Platform Engineer: Cloud Infra & Automation

Location
Greater London, England, United Kingdom
data platform solutions with autonomy and close collaboration with Data Engineering and Platform teams. The role emphasizes IaC, Terraform, CI/CD, security and observability, with a hybrid model (1 day in the office) and a competitive package up to £90,000 plus a 30% bonus. #J-18808-Ljbffr ...

Senior Cloud Platform Engineer - Hybrid Role

Location
Greater London, England, United Kingdom
Platform Tribe in London. The role focuses on designing and evolving a reliable cloud platform, improving AWS, Kubernetes, Terraform, CI/CD, and observability, while partnering with engineers to deliver safe, maintainable solutions. You will own work from problem definition to delivery, contribute to platform standards, and support on-call ...