1 to 25 of 3,928 Observability Jobs

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
produce, distribute and promote the most critically acclaimed and commercially successful music to delight and entertain fans around the world.As a Senior Observability Engineer, you will be a driving force for technical excellence and strategic vision within our global team. You will be instrumental in architecting, building, and leading … comprehensive observability strategy to ensure the reliability, performance, and scalability of our critical IT systems. This senior role demands a passion for data-driven strategy, a commitment to automation, and the ability to mentor and lead. You will not only solve complex technical challenges but also influence the direction ...

Senior Software Engineer

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
that keep critical platforms running globally. This is a hands-on engineering role where you'll design and build software that improves reliability, automation, observability, and resilience across Visa's middleware ecosystem. If you enjoy writing production-grade code, solving complex distributed systems problems, and seeing your work operate … Build systems for deployment orchestration, testing, validation, and rollback Develop self-service platforms that reduce operational toil and improve engineering productivity Improve Reliability & Observability Enhance monitoring, alerting, and system visibility across critical middleware platforms Troubleshoot production issues and implement durable, long-term engineering solutions Improve performance, resiliency, and operational readiness ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
services, including AWS.Knowledge of containerisation and orchestration technologies, including Docker, Kubernetes, and OpenShift.Experience implementing Infrastructure as Code using Terraform or similar solutions.Experience with observability and monitoring platforms such as Splunk, Dynatrace, Grafana, and Prometheus.Understanding of DevSecOps principles, secure software development practices, and security-focused engineering approaches.Experience working within regulated Financial ...

Senior Site Reliability Engineer

Location
Cambridge, England, United Kingdom
improvement of the Bango Platform. You hold everything expected of a Site Reliability Engineer — owning reliability end-to-end across the infrastructure, delivery pipelines, observability and incident response that keep the platform running to agreed service levels — but at greater scope, complexity and influence, and you take responsibility for lifting … than only responding to live issues. Work with Product and Integration Engineering to get recurring and systemic reliability issues onto roadmaps and permanently resolved. Observability, Monitoring & Security Set the standard for observability across teams — raising signal quality, driving down alert fatigue, and introducing the SLI/SLO maturity that others ...

Site Reliability Engineering Lead

Location
City Of London, England, United Kingdom
Risk at https://risk.lexisnexis.com/insurance About our Team The IC (Insurance Core Services) team is responsible for establishing and driving reliability, observability, automation, and operational excellence standards across Insurance technology platforms. The team partners closely with application, infrastructure, database, and cloud engineering teams to improve platform availability … scalability, performance, and resilience. ICS leads strategic initiatives including SLO/SLI implementation, observability platform adoption, cloud modernization, operational readiness reviews, performance engineering, and reliability automation. The team also develops reusable engineering frameworks, standards, and best practices that enable product teams to build and operate highly reliable cloud-native services ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
reliability, performance and continuous improvement of the Bango Platform end-to-end — from the infrastructure and pipelines that build and deploy it, to the observability and incident response that keep it running to agreed service levels. You combine three things that have historically sat in separate teams: platform and cloud … delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. You design, build and operate the automation, observability and platform capabilities that other engineering teams rely on, and you are equally comfortable diagnosing a live incident as you are designing the Terraform module that ...

DevOps Engineer

Location
Greater London, England, United Kingdom
Experience with Infrastructure as Code tools such as Terraform or CloudFormation. Knowledge of CI/CD pipelines and deployment automation. Experience with monitoring and observability tools such as CloudWatch, Prometheus, Grafana, ELK, or similar. Good understanding of Linux systems administration. Experience troubleshooting complex application and infrastructure issues. Knowledge of networking … maintaining Infrastructure as Code (IaC). Support and improve CI/CD pipelines and release automation processes. Implement and maintain monitoring, alerting, logging, and observability solutions. Participate in technical projects including platform upgrades, migrations, and cloud transformation initiatives. Contribute to disaster recovery, business continuity, and high‐availability strategies. Create ...

Senior Software Engineer - NWX Networx

Hiring Organisation
Iris Software
Location
United Kingdom
Salary
£ 70 K
Code, etc.) in engineering practices to accelerate design, development, testing and debugging. Using them critically and responsibly to improve quality, productivity and decision-making.· Observability: Advanced DataDog, Application Insights or Amazon CloudWatch implementation with performance tuning and cost optimisation· Infrastructure as Code: Infrastructure as Code with either Terraform, Bicep … internet-facing traffic levels· Performance & Scalability: Profiling and benchmarking applications.· Application Security: Confident vulnerability management, thread modelling and tracking· Production Support: Knowledge of observability and production support practices. Proficient in debugging complex issues, performance optimization, and production troubleshootingExperience Requirments· 5-7 years of professional software development experience· Proven ability ...

Senior Software Engineer

Hiring Organisation
Iris Software
Location
United Kingdom
Salary
£ 70 K
highly scalable solutions and internet-facing traffic levelsPerformance & Scalability: Profiling and benchmarking applications.Application Security: Confident vulnerability management, thread modelling and trackingProduction Support: Knowledge of observability and production support practices. Proficient in debugging complex issues, performance optimization, and production troubleshootingExperience Requirements5-7 years of professional software development experienceProven ability in delivering … systems with tools and services for richer, automated workflowsExpertise with advanced monitoring and APM strategies using Datadog, including custom dashboards and alerting (or similar observability platform)Expertise with modern UI architecture patterns (micro-frontends, SSR/SSG)Knowledge of security best practices and compliance requirements (OAuth2, OIDC, RBAC)Experience with ...

Lead Java Developer

Location
Swansea, Wales, United Kingdom
Strong experience implementing cloud-native and event-driven architectures . Experience with Infrastructure as Code tools such as Terraform or CloudFormation . Experience with observability and monitoring tools such as ELK, Grafana, Prometheus, or Splunk . Experience with Domain-Driven Design (DDD) and API-first development. Experience with distributed systems ...

Azure CloudOps Engineer

Location
Greater London, England, United Kingdom
resilient, scalable, and highly automated cloud platforms. The successful candidate will be responsible for designing and operating cloud infrastructure, implementing Infrastructure as Code, enhancing observability and AIOps capabilities, and driving automation across both application and infrastructure lifecycles. This role combines Cloud Engineering, DevOps, SRE, and AIOps practices, leveraging automation … optimise CI/CD pipelines supporting both application and infrastructure deployments. Develop and maintain automation scripts, deployment tooling, and operational workflows. Implement and support observability solutions, monitoring platforms, logging frameworks, and telemetry capabilities. Develop AIOps capabilities for anomaly detection, event correlation, alert noise reduction, and proactive issue identification. Create automated ...

Senior Backend Developer

Hiring Organisation
Protein Works
Location
Liverpool, Merseyside, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
connect e-commerce, ERP, WMS, finance, and manufacturing systems. Operations, Security & AI: Own end-to-end delivery including containers, CI/CD, IaC, full observability (metrics, traces, logs), security/privacy compliance (OWASP, GDPR, pen testing), and day-to-day use of AI-assisted/agentic tools. Cross-Functional Leadership … microservices and monoliths. Data & Async Systems: Solid background in relational databases (PostgreSQL, MySQL, SQL Server) and asynchronous messaging (queues, events, retries, idempotency). DevOps & Observability: Practical experience with Docker, Linux/Windows, Git/GitHub, Jira/Agile, monitoring/alerting, and diagnostic/profiling tools in high-volume ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud … Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner ...

GCP Data Engineer

Hiring Organisation
Teksystems
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£400/day
governance standards are consistently met across AWS environments, including IAM, encryption, and network security controls. Monitor system health, performance, and cost using appropriate observability tools, and proactively drive optimisation initiatives. Work with relational databases and data integration patterns to support reference data and other financial data domains. Collaborate with global … Kubernetes/EKS. Good understanding of DevOps practices, CI/CD pipelines, and automation tools to support continuous delivery. experience with monitoring, logging, and observability tools such as CloudWatch, Prometheus, and Grafana. Strong understanding of security best practices in cloud environments, including IAM, encryption, and network security. Significant experience working ...

Engineering lead

Location
Greater London, England, United Kingdom
security scanning, and performance engineering. Drive CI/CD adoption and DevSecOps practices. Monitor engineering metrics, technical debt, and platform stability. Improve system reliability, observability, and operational readiness. Stakeholder Management Act as the primary technology contact for business and delivery stakeholders. Present engineering updates, risks, dependencies, and mitigation plans. Work … Kubernetes Docker OpenShift DevOps & Automation Jenkins GitHub/Bitbucket Terraform Ansible CI/CD Pipelines Engineering Practices TDD/BDD Secure Coding Performance Engineering Observability & Monitoring SRE Principles < h3>Tools JIRA Confluence Azure DevOps Splunk Dynatrace SonarQube Banking Domain Knowledge Preferred Experience In Retail Banking Commercial Banking Digital Banking Identity ...

GCP Cloud Engineer

Location
Warwick, England, United Kingdom
requirements by implementing controls, monitoring, and evidence collection for cloud environments. Respond to security incidents, perform root cause analysis, and drive continuous improvement. Monitoring, Observability & Cost Optimization: Deploy and manage cloud-native monitoring, logging, and tracing solutions (e.g., Azure Monitor, GCP Operations Suite, AWS CloudWatch, Prometheus, Grafana). Analyse resource … frameworks, IAM, encryption, vulnerability management, and compliance (ISO, NIST, CIS) Networking: Expertise in cloud networking, VPNs, load balancers, DNS, CDN, and hybrid connectivity Monitoring & Observability: Experience with cloud-native and open-source monitoring, logging, and alerting tools Certifications: Azure Solutions Architect, GCP Professional Cloud Architect. Experience with policy-as-code ...

Senior Site Reliability Engineer

Location
Southampton, England, United Kingdom
such as Jenkins, GitLab CI/CD, or CircleCI. Strong knowledge of containerization technologies (e.g., Docker, Kubernetes) and microservices architecture. Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack, Cloudwatch). Excellent problem-solving skills and the ability to troubleshoot complex issues in distributed systems. Experience of Incident … advantage if you also have: Handson experience of working with large Kubernetes Cluster. Certification will be an added plus. Working experience of Grafana Observability Suite (Loki, Mimir, Tempo). Administration and/or development experience of standard monitoring and automation tools such as Splunk, Datadog, Pagerduty Rundeck. Familiarity with configuration ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
such as Jenkins, GitLab CI/CD, or CircleCI. Strong knowledge of containerization technologies (e.g., Docker, Kubernetes) and microservices architecture. Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack, Cloudwatch). Excellent problem-solving skills and the ability to troubleshoot complex issues in distributed systems. Experience of Incident … advantage if you also have: Handson experience of working with large Kubernetes Cluster. Certification will be an added plus. Working experience of Grafana Observability Suite (Loki, Mimir, Tempo). Administration and/or development experience of standard monitoring and automation tools such as Splunk, Datadog, Pagerduty Rundeck. Familiarity with configuration ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
fintech, payments, or enterprise SaaS platforms* Exposure to event-driven architecture (Kafka, RabbitMQ)* Familiarity with infrastructure-as-code tools (Terraform, CloudFormation)* Understanding of observability tools (Prometheus, Grafana, ELK stack) #J-18808-Ljbffr ...

DevOps Engineer (Security Cleared)

Location
Greater London, England, United Kingdom
automation using Python, Bash, PowerShell, or similar languages. Experience with configuration management tools such as Ansible, Puppet, or Chef. Knowledge of monitoring, logging, and observability platforms such as Prometheus, Grafana, ELK Stack, Splunk, or Datadog. Strong understanding of Linux administration, networking, cloud security, and DevSecOps principles. Experience working in Agile ...

Senior Software Engineer

Location
City of Westminster, England, United Kingdom
fintech, payments, or enterprise SaaS platforms Exposure to event-driven architecture (Kafka, RabbitMQ) Familiarity with infrastructure-as-code tools (Terraform, CloudFormation) Understanding of observability tools (Prometheus, Grafana, ELK stack) L’état d’esprit Edenred - Nous sommes une entreprise unique. Nous recherchons de nouveaux collaborateurs prêts à prendre part ...

Senior Software Engineer - Backend

Hiring Organisation
Fitch Ratings
Location
London, UK
Employment Type
Full-time
cross-functional stakeholders to prioritize work, align technical investments, and achieve business outcomes. Ensure high-quality software delivery through automated testing, code reviews, observability, and engineering governance. Lead resolution of complex technical and operational challenges while improving platform performance, resiliency, and operational excellence. Champion DevSecOps, CI/CD, and automation ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
hands-on, Linux-focused DevOps role where you will take ownership of large-scale production environments, working across Linux systems, automation, Kubernetes, containerisation, networking, observability and cloud infrastructure. Location: London, hybrid working with a minimum of 2 days per week in the office Salary: Up to £100,000 per annum … networking, security and performance troubleshooting within Linux environments Hands-on experience operating containerised workloads using Docker and Kubernetes Experience with monitoring, logging and observability technologies such as Prometheus, Grafana, Loki, OpenTelemetry or the ELK Stack Experience with Infrastructure as Code and automation tooling such as Terraform and Ansible Experience building ...

Senior Platform/Dev Ops Engineer

Location
Manchester, England, United Kingdom
platform reliability across multi-cloud environments. The ideal candidate has strong experience building enterprise-scale CI/CD platforms, Kubernetes ecosystems, Infrastructure as Code, observability solutions, and developer enablement capabilities. Key Responsibilities Architect, build, and evolve scalable platform engineering and DevOps solutions for enterprise engineering teams. Lead the design … manage scalable Infrastructure as Code frameworks using Terraform and Terraform Enterprise. Lead Kubernetes platform architecture, Helm-based deployments, and container orchestration best practices. Establish observability standards and operational excellence using Grafana, Loki, Prometheus, OpenTelemetry, and related tooling. Collaborate with engineering, security, cloud, and architecture teams to improve platform capabilities ...

Senior Machine Learning Engineer (MLOps)

Location
Greater London, England, United Kingdom
search relevance, personalisation and emerging AI applications. This is a highly engineering-focused role with an emphasis on cloud-native systems, platform architecture, automation, observability and operational excellence. What You’ll Be Doing: Design and build scalable machine learning platforms and infrastructure supporting model training, deployment and serving. Develop highly … data products. Design batch and real-time inference architectures using modern cloud-native technologies. Improve reliability, resilience and performance across ML workloads through monitoring, observability and automation. Build tooling and frameworks that enable data scientists and ML engineers to deploy models safely and efficiently. Own production services, infrastructure and operational ...