151 to 175 of 4,022 Observability Jobs

Senior DevSecOps Engineer (Remote)

Hiring Organisation
Integrated Data Services
Location
United States
Employment Type
Permanent
Salary
USD 160,000 Annual
monitoring, troubleshooting) is required 2+ years of experience in Configuration management (Ansible, Chef, or Puppet) is preferred 2+ years of experience in Monitoring and observability (Prometheus, Grafana, ELK stack, CloudWatch) is preferred Security Tools & Practices: 3+ years of experience in DevSecOps practices (security automation, shift-left security) is required 2+ ...

AWS Devops Engineer

Location
West of England, England, United Kingdom
disaster recovery strategies to ensure data protection and business continuity] Ability to implement monitoring and logging solutions e.g., CloudWatch, to ensure system reliability, observability, and proactive incident response Comfortable working in Agile development teams, translating business requirements into technical solutions, and actively participating in sprint planning, retrospectives, and daily stand ...

AI Full-Stack Engineer

Location
United Kingdom
## AI Full-Stack EngineerApply: Various: Full time: Posted Today: R527**Position Summary**The AI Full-Stack Engineer is responsible for designing, developing, implementing, and supporting enterprise AI solutions and modern business applications. This role ...

Lead Data Engineer

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
Shape how hundreds of thousands of UK investors use data to make confident, informed investment decisions. Join a team building modern, cloud-native data platforms that enable analytics, regulatory reporting, and data-driven products at ...

Production Engineer

Location
City Of London, England, United Kingdom
Group OverviewThe TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through ...

Managing Engineer - Observability, Pipeline & Analytics (Hybrid)

Location
Belfast City District, Northern Ireland, United Kingdom
identity protection. Your role in the team Core Engineering & Platforms (CEP) group is seeking a highly skilled and hands‐on Software Engineering Manager for Observability Pipeline & Analytics, who can lead the engineering strategy, architecture, delivery, and operational excellence of Allstate’s enterprise telemetry and analytics platforms, including Cribl, Azure Data … Explorer (ADX), and related observability data services, and own the end‐to‐end observability pipeline ecosystem that enables scalable collection, transformation, routing, governance, enrichment, retention, and analytics of operational telemetry across various applications, infrastructure platforms, cloud services, and technology domains. The role combines deep technical expertise with people leadership, requiring ...

Senior Engineering Manager - SRE

Location
Birmingham, England, United Kingdom
systems, and how fast we can act on what we see. As Senior Manager, Site Reliability Engineering, you'll own the monitoring and observability roadmap across our full product portfolio — spanning Health, Legal, Education, Workforce Management, and other sectors — and lead a team of 15-20 SRE engineers to deliver … dashboard, and moving us toward a more autonomous operating model. Mandate: To lead OneAdvanced's Site Reliability Engineering function — building the monitoring and observability roadmap that gives every product across Health, Legal, Education, Workforce Management, and other sectors real visibility into its own health, and increasingly uses AI agents ...

Senior Lead Site Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part of an agile … significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices to support the business ...

Platform Engineer (DevOps / MLOps Focus)

Hiring Organisation
The Portfolio Group
Location
London, United Kingdom
Employment Type
Permanent
Salary
£100000/annum
environments for production workloads. Developing and managing Infrastructure as Code using Terraform. Supporting CI/CD pipelines and platform automation initiatives. Improving platform reliability, observability and scalability. Collaborating closely with Software Engineers, Data Engineers and ML teams to optimise deployment workflows and infrastructure performance. Essential experience: Strong commercial experience with … available, scalable production environments. Nice to have: Experience with Kubeflow and ML platform tooling. Exposure to AI, machine learning or GenAI projects. Experience with observability tooling such as Prometheus, Grafana or OpenTelemetry. Experience working within regulated or enterprise environments. This is a fantastic opportunity to join a high-profile ...

Senior DevOps Engineer – Build Pipelines

Location
Greater London, England, United Kingdom
Application Engineering Teams To Design, Implement And Roll Out These Standards, With a Strong Focus On Repeatable CI/CD Automation Security Quality gates Observability Deployment consistency Developer experience Platform engineering principles This is a hands-on engineering role , rather than a traditional operational DevOps position. The Role … signing Security and quality thresholds Automated quality gates The objective is for security and quality controls to enabled by default across the application estate. Observability Integrate pipeline and deployment processes with appropriate monitoring and observability tooling. Help teams understand build and deployment health. Work with technologies such as: Prometheus Grafana ...

Lead Site Reliability Engineer

Location
Southampton, England, United Kingdom
enforce SLOs, SLAs, and error budgets Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor. Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc. Developing bicep modules for monitoring infrastructure and deploy it. Optimize system performance, cost … reproducing issues in a local environment. Multi-tasking and time-management to prioritise and switch between varied tasks. Significant experience in platform engineering, observability, and provisioning. Proven ability to develop and implement a strategic vision for platform services, observability, and provisioning. Strong understanding of cyber security principles, governance, and compliance ...

Test Environment Manager

Location
Greater London, England, United Kingdom
configuration, and teardown of test environments. Integrate environment automation seamlessly into CI/CD pipelines to enable on-demand, self-service environment delivery. Reliability & Observability Define and maintain Service Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning time, and stability metrics. Monitor environment health … using observability tools (Prometheus, Grafana, Splunk, etc.) and proactively identify and resolve performance issues or bottlenecks. Incident & Problem Management Lead incident response for environment-related issues, driving quick resolution and facilitating blameless post-mortems. Implement permanent fixes based on root cause analysis and reduce repeat incidents. Automation & Toil Reduction Identify ...

DevOps Engineer

Location
Horsham, England, United Kingdom
security, resilience and operational excellence. Working within multidisciplinary Agile teams, your work will include building secure CI/CD pipelines, automating infrastructure, implementing observability and monitoring solutions, and supporting live environments through SRE‐aligned ways of working. Typical responsibilities include: Designing, building and maintaining secure DevSecOps pipelines. Developing Infrastructure … Code and automated platform capabilities. Supporting and optimising AWS, Azure, Kubernetes, OpenShift, EKS and ECS environments. Implementing monitoring, logging, observability and alerting solutions for mission‐critical services. Contributing to secure, scalable and resilient cloud architectures within regulated environments. Due to the nature of our work, some assignments may require higher ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
Global Sales and Customer Success teams to help organizations navigate the complexities of "day-2" operations in containerized environments, ensuring they achieve full-stack observability across dynamic, ephemeral architectures.**Core Responsibilities**1. Cloud Native Domian Expertise:* Act as the technical "Subject Matter Expert" (SME) for all things Kubernetes … Concepts (POCs) specifically targeting "Born-in-the-Cloud" accounts or legacy enterprises undergoing digital transformation.* Build "Cloud-Native ROI" models, illustrating how automated observability reduces "Kubernetes Tax" (the operational overhead of managing K8s at scale).* Support the sales team in positioning Dynatrace against niche cloud-native tools by highlighting ...

Senior Data Engineer London

Location
Greater London, England, United Kingdom
them into robust technical solutions. Contribute to solution architecture, platform design, and technology selection decisions. Implement software engineering and DataOps best practices including testing, observability, CI/CD, version control, and infrastructure automation. Develop and optimise data models, transformations, and storage layers to ensure performance, reliability, and scalability. Collaborate with … Solid understanding of data lakehouse, warehouse, and medallion architecture patterns. Experience applying DevOps and DataOps best practices including CI/CD, automated testing, monitoring, observability, and release management. Experience with Infrastructure as Code technologies such as Terraform, Pulumi, CloudFormation, or Bicep. Strong understanding of distributed systems, data modelling, data governance ...

Senior Platform Engineer

Hiring Organisation
Workable Software Limited
Location
South East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Agents: Develop the infrastructure required to run increasingly agent-driven AI workflows reliably in production, including state management, task execution, model routing, retries and observability Support AI Infrastructure: Work closely with our AI engineers supporting production NLP, LLM and quantitative model workloads, including our in-house GPU infrastructure. Automate Infrastructure … Build and manage Infrastructure-as-Code using Pulumi and automate deployments through GitHub Actions. Own Reliability & Observability: Build monitoring, logging, tracing and alerting across our data, model, service and agent infrastructure. Requirements Degree in a related technical field: Computer science, software engineering or similar. 7+ years of professional experience ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
overhead. Support and enhance CI/CD platforms and engineering workflows, including the safe promotion of infrastructure and platform changes. Improve platform reliability, security, observability, performance and operational excellence. Partner with software engineers, quantitative developers, research teams, Technology Operations and security colleagues to deliver secure, scalable and resilient platforms. Education … experience with AWS and cloud‐native infrastructure. Experience with Kubernetes, including Amazon EKS, and Docker or other container runtimes. Experience with monitoring and observability tools such as Grafana, Prometheus, Loki or Datadog. Experience with artifact‐management platforms such as JFrog Artifactory. Knowledge of AWS Batch, AWS Step Functions, AWS Identity ...

Custody Support Applications Support - Assistant Vice President

Location
Belfast City District, Northern Ireland, United Kingdom
operational excellence of a suite of business-critical custody and settlement applications. The role focuses on distributed systems, cloud-native technologies, microservices, and modern observability platforms supporting securities processing and settlement functions. This position combines traditional application support responsibilities with Site Reliability Engineering (SRE) principles, automation, resiliency engineering, and operational … enterprise relational databases Database performance analysis and tuning Linux/Unix fundamentals Application troubleshooting and performance diagnostics Incident and Problem Management processes Monitoring and observability platforms Desirable: OpenShift/Kubernetes Cloud technologies (Google Cloud Platform, Azure, AWS) Microservices and distributed architectures Event-driven architectures and messaging platforms (Kafka, MQ) Elastic ...

Senior Cloud Solution Architect Manager

Hiring Organisation
FM
Location
Johnston, Rhode Island, United States
Employment Type
Permanent
Salary
USD Annual
business and technology transformation. This role requires deep expertise in Microsoft Azure, hybrid cloud architecture, Kubernetes, Terraform, cloud networking, DevSecOps, identity and access management, observability, and platform engineering. You will play a key role in advancing cloud-native architectures, infrastructure automation, self-service developer platforms, and modern engineering practices while … infrastructure services end-to-end as reliable, secure, cost-efficient, self-service capabilities across hybrid environments. Own production outcomes including reliability (SLIs/SLOs), observability, incident response, security, and FinOps (cost visibility, optimization, and accountability). Lead the evolution of the Internal Developer Platform into an AI-first platform ...

LiveOps Engineer

Hiring Organisation
Civica
Location
United Kingdom
Salary
£ 70 K
Civica Our LiveOps team sits at the front line of this mission: monitoring systems, responding to incidents, and continuously improving the automation and observability that keep our services healthy. Every action you take directly contributes to the stability and trust our customers depend on. As a LiveOps Engineer, you will … cloud environments. Working closely with SRE, Platform, and Development teams, interfacing with the support and services teams, you will support uptime, automation, and observability improvements that make our services more reliable. You’ll contribute to incident response, environment management, and the ongoing evolution of how Civica delivers and supports ...

DevOps Engineer – Security & Intelligence

Location
Manchester, England, United Kingdom
pipelines to support continuous delivery Developing Infrastructure as Code and automated platform capabilities Supporting AWS-based environments, including Kubernetes, OpenShift, EKS and ECS Implementing observability, monitoring, logging and alerting for live services Supporting SRE practices, cloud migration activities and production platform operations Job Responsibilities Design, implement and maintain secure … Ansible) Support containerised platforms and environments, including Kubernetes, OpenShift, EKS and ECS Embed security controls, quality gates and compliance checks into DevSecOps workflows Implement observability, monitoring, logging and alerting to support reliable live services Contribute to SRE practices, improving service reliability, performance and operational resilience Automate build, deployment and platform ...

DevOps Engineer - Security & Intelligence

Hiring Organisation
Hackajob Ltd
Location
Cheltenham, Gloucestershire, South West, United Kingdom
Employment Type
Permanent
pipelines to support continuous delivery Developing Infrastructure as Code and automated platform capabilities Supporting AWS-based environments, including Kubernetes, OpenShift, EKS and ECS Implementing observability, monitoring, logging and alerting for live services Supporting SRE practices, cloud migration activities and production platform operations Job Responsibilities Design, implement and maintain secure … Ansible) Support containerised platforms and environments, including Kubernetes, OpenShift, EKS and ECS Embed security controls, quality gates and compliance checks into DevSecOps workflows Implement observability, monitoring, logging and alerting to support reliable live services Contribute to SRE practices, improving service reliability, performance and operational resilience Automate build, deployment and platform ...

DevOps Engineer – Security & Intelligence

Location
Greater London, England, United Kingdom
pipelines to support continuous delivery Developing Infrastructure as Code and automated platform capabilities Supporting AWS-based environments, including Kubernetes, OpenShift, EKS and ECS Implementing observability, monitoring, logging and alerting for live services Supporting SRE practices, cloud migration activities and production platform operations Job Responsibilities Design, implement and maintain secure … Ansible) Support containerised platforms and environments, including Kubernetes, OpenShift, EKS and ECS Embed security controls, quality gates and compliance checks into DevSecOps workflows Implement observability, monitoring, logging and alerting to support reliable live services Contribute to SRE practices, improving service reliability, performance and operational resilience Automate build, deployment and platform ...

Senior DevOps Engineer

Location
Ashford, England, United Kingdom
Python/Go and modern CI/CD frameworks to shape how software is delivered at scale. Your impact will be visible: from improving observability and reliability across our payment platforms to driving an automate‐first culture that removes manual effort and accelerates delivery for engineering teams. Because this role … automate manual maintenance and repeatable tasks and critical business requests in addition to environment deployment Recommend, design and execute optimizations to improve the observability, performance and overall reliability of Verifone’s payment platforms Develop and maintain monitoring and logging solutions to ensure visibility into application performance and security Advocate DevOps ...