1 to 25 of 3,178 Observability Jobs in the UK

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
produce, distribute and promote the most critically acclaimed and commercially successful music to delight and entertain fans around the world.As a Senior Observability Engineer, you will be a driving force for technical excellence and strategic vision within our global team. You will be instrumental in architecting, building, and leading … comprehensive observability strategy to ensure the reliability, performance, and scalability of our critical IT systems. This senior role demands a passion for data-driven strategy, a commitment to automation, and the ability to mentor and lead. You will not only solve complex technical challenges but also influence the direction ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
services, including AWS.Knowledge of containerisation and orchestration technologies, including Docker, Kubernetes, and OpenShift.Experience implementing Infrastructure as Code using Terraform or similar solutions.Experience with observability and monitoring platforms such as Splunk, Dynatrace, Grafana, and Prometheus.Understanding of DevSecOps principles, secure software development practices, and security-focused engineering approaches.Experience working within regulated Financial ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
reliability, performance and continuous improvement of the Bango Platform end-to-end — from the infrastructure and pipelines that build and deploy it, to the observability and incident response that keep it running to agreed service levels. You combine three things that have historically sat in separate teams: platform and cloud … delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. You design, build and operate the automation, observability and platform capabilities that other engineering teams rely on, and you are equally comfortable diagnosing a live incident as you are designing the Terraform module that ...

DevOps Engineer

Location
Greater London, England, United Kingdom
maintaining Infrastructure as Code (IaC). Support and improve CI/CD pipelines and release automation processes. Implement and maintain monitoring, alerting, logging, and observability solutions. Participate in technical projects including platform upgrades, migrations, and cloud transformation initiatives. Contribute to disaster recovery, business continuity, and high‐availability strategies. Create … Experience with Infrastructure as Code tools such as Terraform or CloudFormation. Knowledge of CI/CD pipelines and deployment automation. Experience with monitoring and observability tools such as CloudWatch, Prometheus, Grafana, ELK, or similar. Good understanding of Linux systems administration. Experience troubleshooting complex application and infrastructure issues. Knowledge of networking ...

DevOps Engineer

Location
Greater London, England, United Kingdom
Experience with Infrastructure as Code tools such as Terraform or CloudFormation. Knowledge of CI/CD pipelines and deployment automation. Experience with monitoring and observability tools such as CloudWatch, Prometheus, Grafana, ELK, or similar. Good understanding of Linux systems administration. Experience troubleshooting complex application and infrastructure issues. Knowledge of networking … maintaining Infrastructure as Code (IaC). Support and improve CI/CD pipelines and release automation processes. Implement and maintain monitoring, alerting, logging, and observability solutions. Participate in technical projects including platform upgrades, migrations, and cloud transformation initiatives. Contribute to disaster recovery, business continuity, and high‐availability strategies. Create ...

Lead Java Developer

Location
Swansea, Wales, United Kingdom
Strong experience implementing cloud-native and event-driven architectures . Experience with Infrastructure as Code tools such as Terraform or CloudFormation . Experience with observability and monitoring tools such as ELK, Grafana, Prometheus, or Splunk . Experience with Domain-Driven Design (DDD) and API-first development. Experience with distributed systems ...

Senior Lead SRE: Reliability, Observability & Resiliency

Location
Auchentibber, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability/DevOps Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part … significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices to support the business ...

Azure CloudOps Engineer

Location
Greater London, England, United Kingdom
resilient, scalable, and highly automated cloud platforms. The successful candidate will be responsible for designing and operating cloud infrastructure, implementing Infrastructure as Code, enhancing observability and AIOps capabilities, and driving automation across both application and infrastructure lifecycles. This role combines Cloud Engineering, DevOps, SRE, and AIOps practices, leveraging automation … optimise CI/CD pipelines supporting both application and infrastructure deployments. Develop and maintain automation scripts, deployment tooling, and operational workflows. Implement and support observability solutions, monitoring platforms, logging frameworks, and telemetry capabilities. Develop AIOps capabilities for anomaly detection, event correlation, alert noise reduction, and proactive issue identification. Create automated ...

Senior Backend Developer

Hiring Organisation
Protein Works
Location
Liverpool, Merseyside, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
connect e-commerce, ERP, WMS, finance, and manufacturing systems. Operations, Security & AI: Own end-to-end delivery including containers, CI/CD, IaC, full observability (metrics, traces, logs), security/privacy compliance (OWASP, GDPR, pen testing), and day-to-day use of AI-assisted/agentic tools. Cross-Functional Leadership … microservices and monoliths. Data & Async Systems: Solid background in relational databases (PostgreSQL, MySQL, SQL Server) and asynchronous messaging (queues, events, retries, idempotency). DevOps & Observability: Practical experience with Docker, Linux/Windows, Git/GitHub, Jira/Agile, monitoring/alerting, and diagnostic/profiling tools in high-volume ...

DevOps Engineer

Hiring Organisation
Infinity Quest
Location
Greater Edinburgh Area, United Kingdom
integration platforms. - Drive incident management, root cause analysis, and post-incident reviews. - Improve platform availability, scalability, and performance through automation and engineering improvements. - Monitoring & Observability - Implement enterprise monitoring, alerting, logging, and observability solutions. Utilize tools including: - Prometheus - Grafana - Google Cloud Monitoring - ELK/Elastic Stack - OpenTelemetry - Proactively identify issues before … with API Gateway technologies such as: - Apigee - Azure API Management - Kong (desirable) - Strong understanding of: - REST APIs - OpenAPI Specifications - OAuth2 - JWT - API Security Standards - Observability & Monitoring Experience with: - Prometheus - Grafana - ELK Stack - OpenTelemetry - Google Cloud Monitoring - Scripting & Automation Proficiency in one or more of: - Python [optional] - Bash - Go - PowerShell - Networking ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud … Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner ...

Engineering lead

Location
Greater London, England, United Kingdom
security scanning, and performance engineering. Drive CI/CD adoption and DevSecOps practices. Monitor engineering metrics, technical debt, and platform stability. Improve system reliability, observability, and operational readiness. Stakeholder Management Act as the primary technology contact for business and delivery stakeholders. Present engineering updates, risks, dependencies, and mitigation plans. Work … Kubernetes Docker OpenShift DevOps & Automation Jenkins GitHub/Bitbucket Terraform Ansible CI/CD Pipelines Engineering Practices TDD/BDD Secure Coding Performance Engineering Observability & Monitoring SRE Principles < h3>Tools JIRA Confluence Azure DevOps Splunk Dynatrace SonarQube Banking Domain Knowledge Preferred Experience In Retail Banking Commercial Banking Digital Banking Identity ...

DevOps Engineer (Security Cleared)

Location
Greater London, England, United Kingdom
automation using Python, Bash, PowerShell, or similar languages. Experience with configuration management tools such as Ansible, Puppet, or Chef. Knowledge of monitoring, logging, and observability platforms such as Prometheus, Grafana, ELK Stack, Splunk, or Datadog. Strong understanding of Linux administration, networking, cloud security, and DevSecOps principles. Experience working in Agile ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
hands-on, Linux-focused DevOps role where you will take ownership of large-scale production environments, working across Linux systems, automation, Kubernetes, containerisation, networking, observability and cloud infrastructure. Location: London, hybrid working with a minimum of 2 days per week in the office Salary: Up to £100,000 per annum … networking, security and performance troubleshooting within Linux environments Hands-on experience operating containerised workloads using Docker and Kubernetes Experience with monitoring, logging and observability technologies such as Prometheus, Grafana, Loki, OpenTelemetry or the ELK Stack Experience with Infrastructure as Code and automation tooling such as Terraform and Ansible Experience building ...

Senior DevOps Engineer

Location
City of Edinburgh, Scotland, United Kingdom
integration platforms. Drive incident management, root cause analysis, and post-incident reviews. Improve platform availability, scalability, and performance through automation and engineering improvements. Monitoring & Observability Implement enterprise monitoring, alerting, logging, and observability solutions. Grafana ELK/Elastic Stack OpenTelemetry Proactively identify issues before they impact customers and engineering teams. Develop ...

DevOps Solution Architect

Hiring Organisation
Frontier Resourcing Ltd
Location
South East, United Kingdom
Employment Type
Permanent
Ability to influence, guide, and align multiple teams and stakeholders. Desirable Skills Experience defining platform strategies or roadmaps at an organisational level. Knowledge of observability and monitoring solutions (Prometheus, Grafana, ELK, Splunk). Experience implementing SRE practices (SLIs, SLOs, error budgets). Familiarity with compliance frameworks and continuous compliance tooling. ...

DevOps Solution Architect

Location
London, United Kingdom
Ability to influence, guide, and align multiple teams and stakeholders. Desirable Skills Experience defining platform strategies or roadmaps at an organisational level. Knowledge of observability and monitoring solutions (Prometheus, Grafana, ELK, Splunk). Experience implementing SRE practices (SLIs, SLOs, error budgets). Familiarity with compliance frameworks and continuous compliance tooling. ...

Senior Lead Software Engineer - Python / Go

Location
Auchentibber, Scotland, United Kingdom
eliminate platform bottlenecks. Define and promote paved paths and self-service workflows for developers. Implement real-time telemetry pipelines and workflows for platform observability and analytics. Champion adoption of productivity tools through documentation, training, and engagement with the developer community. Standardize use of AI-assisted coding tools and AI-powered ...

Principal Cloud Engineer (Terraform), London

Location
Greater London, England, United Kingdom
Apply data quality and validation frameworks to ensure accuracy, completeness, and freshness of cost and usage data across all cloud providers; instrument pipelines with observability tooling to surface issues proactively. Build and maintain reusable data assets — curated datasets, aggregations, and data marts — that power FinOps dashboards, showback/chargeback reporting ...

Monitoring & Observability Engineer (Dynatrace)

Location
Greater London, England, United Kingdom
Monitoring & Observability Engineer (Dynatrace) Location: UK - London, UK - Reading, UK - Hatfield, UK - Nottingham, UK - Manchester, UK - Milton Keynes, UK - Birmingham | Job-ID: 214264 | Contract type: Standard | Business Unit: IT Consulting Life on the team At Computacenter, you’ll be joining a world-class team of over 1,000 skilled professionals … across the UK, Germany, France, and India, delivering complex, enterprise-grade IT solutions and consultancy across infrastructure, cloud, and modern operations. As a Monitoring & Observability Engineer, you'll work in high-impact delivery teams that support some of the world’s most well-known organisations. You’ll play ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
Site Reliability Engineer (SRE) - Assistant Vice President is a technical professional responsible for the hands‐on execution, technical implementation, and deployment of SRE and observability principles in a complex, critical, and large-scale multi-disciplinary environment. In this role, you will apply a deep understanding of multiple technology domains … technical contributor, you will drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors ...

Senior Site Reliability Engineer

Hiring Organisation
GCS
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£75000 - £95000/annum Bonus
drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem-solving activities. * Work ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams.It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling.**Responsibilities*** Lead the design and evolution of scalable, secure, and highly available trading infrastructure.* Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows.* Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering.* Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices.* Lead and participate ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...