76 to 100 of 3,940 Observability Jobs

Software Engineer - DevOps

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
issues and partnering with development teams to resolve them. Maintain and monitor asset inventory across environments. Monitor, troubleshoot, and remediate issues using Splunk and observability/monitoring platforms such as Datadog, Dynatrace, or Grafana. Support cost rationalization efforts; partner with architects to gather, clarify, and translate technical requirements. Proficient with ...

Senior Platform Engineer

Hiring Organisation
Willis Towers Watson
Location
Reigate, Surrey, United Kingdom
Salary
£ 70 K
OIDC, JWT and claims‐based authorisation.Experience in platform engineering or SRE roles: building internal platforms, defining SLIs/SLOs, managing error budgets, and implementing observability (centralised logging, metrics, distributed tracing).Strong awareness of emerging cloud, AI, DevOps and platform technologies, with an understanding of their applicability to SaaS platforms.General knowledge ...

Senior Platform Engineer

Hiring Organisation
WTW
Location
Surrey, United Kingdom
Employment Type
Full Time
claims‐based authorisation. Experience in platform engineering or SRE roles: building internal platforms, defining SLIs/SLOs, managing error budgets, and implementing observability (centralised logging, metrics, distributed tracing). Strong awareness of emerging cloud, AI, DevOps and platform technologies, with an understanding of their applicability to SaaS platforms. General knowledge ...

AI Engineer

Hiring Organisation
Ten Group
Location
London, United Kingdom
Salary
£ 70 K
platforms (AWS, GCP, Azure) and infrastructure-as-code (Terraform etc). Hands-on with DevOps/Infra tooling (CI/CD, Docker, K8s) and observability (Prometheus, Grafana, Datadog etc) Experience building distributed systemsKnowledge and hands-on experience with multiple datastores (both SQL and NoSQL) Desired experience in building agents ...

Production Engineer

Location
Greater London, England, United Kingdom
identify patterns, and drive intelligent automation solutions* Hands-on experience with containerization technologies such as Docker and Kubernetes, including cluster management, deployment, scaling, and observability* Deep practical knowledge of algorithmic trading workflows, including the behaviour, lifecycle, and risk controls of execution algos used across the EMEA markets* Experience designing ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
ingress controllers. Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud‐native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications ...

Cloud Native Senior Architect (Portworx)

Hiring Organisation
Pure Storage
Location
United Kingdom
Salary
£ 70 K
optimized for performance, scale, and reliabilityLead API/microservices solution deployment and day-2 operations, including monitoring via Prometheus/Grafana/other Observability and alerting best practicesSupport CI/CD pipelines and Infrastructure-as-Code workflows (Jenkins, GitOps, Terraform, Ansible, Chef, Puppet) relevant to Portworx deliveryCustomer Support and EnablementAct ...

Production Engineer

Location
City Of London, England, United Kingdom
identify patterns, and drive intelligent automation solutions Hands-on experience with containerization technologies such as Docker and Kubernetes, including cluster management, deployment, scaling, and observability Deep practical knowledge of algorithmic trading workflows, including the behaviour, lifecycle, and risk controls of execution algos used across the EMEA markets Experience designing ...

Site Reliability Engineer

Location
Slough, England, United Kingdom
Site Reliability Engineer (SRE) DevSecOps | Cloud Engineering | Observability | Production Environments | London SR2 is supporting a major 3-year programme and looking for an experienced Site Reliability Engineer (SRE) to join the Production Engineering team. This function underpins the reliability, security, and performance of all live environments, from production systems … native infrastructure. Beyond supporting live systems, this team also acts as a centre of excellence, guiding project teams in adopting best practices across DevSecOps, observability, and cost optimisation. Key Responsibilities: Build, maintain, and support production and demo environments Automate infrastructure provisioning and deployment workflows (Terraform, GitHub Actions, GitOps) Package ...

Senior Lead Site Reliability / DevOps Engineer

Location
Auchentibber, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part of an agile … significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices to support the business ...

Senior Lead Site Reliability / DevOps Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability/DevOps Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part … significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilitiesRegularly provides technical guidance and direction on site reliability practices to support the business ...

Lead Software Engineer - Java, Go

Location
Bournemouth, England, United Kingdom
will work with technologies including Java, Spring Boot, Go, Kubernetes, Terraform, Ansible, PostgreSQL, CockroachDB, Redis, CI/CD tooling, public cloud services, and observability platforms to solve complex engineering challenges and improve operational stability across critical infrastructure services. As a Lead Software Engineer at JPMorgan Chase within Infrastructure Platforms, Enterprise … product areas include Automation as a Service (AaaS), Ansible Automation Platform (AAP), AutoM8, hybrid cloud automation, public cloud and mainframe enablement, SRE and observability capabilities, and AI-driven service management accelerators. Job Responsibilities Executes software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches ...

Senior DevOps Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£80,000
Terraform Creating, deploying and managing Kubernetes-based platforms Building and improving automated CI/CD pipelines and deployment processes Implementing monitoring, logging and observability solutions across cloud environments Driving infrastructure best practices around security, resilience, compliance and disaster recovery Optimising platforms for performance, availability and cost efficiency Troubleshooting complex infrastructure … with Kubernetes, Docker and containerised applications Strong CI/CD experience using tools such as GitLab CI, Jenkins or similar Knowledge of monitoring and observability tools such as Prometheus, Grafana or Elastic Stack Experience delivering highly available, secure and resilient cloud solutions Familiarity with Agile delivery environments, including Scrum ...

Senior DevOps Engineer

Hiring Organisation
Anson Mccade
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£80,000
Terraform Creating, deploying and managing Kubernetes-based platforms Building and improving automated CI/CD pipelines and deployment processes Implementing monitoring, logging and observability solutions across cloud environments Driving infrastructure best practices around security, resilience, compliance and disaster recovery Optimising platforms for performance, availability and cost efficiency Troubleshooting complex infrastructure … with Kubernetes, Docker and containerised applications Strong CI/CD experience using tools such as GitLab CI, Jenkins or similar Knowledge of monitoring and observability tools such as Prometheus, Grafana or Elastic Stack Experience delivering highly available, secure and resilient cloud solutions Familiarity with Agile delivery environments, including Scrum ...

Director of DevOps & SRE

Location
Greater London, England, United Kingdom
high-severity incidents and major client environment changes, ensuring appropriate change management and CAB governance, with occasional off-hours support for the teams. Observability and service levels. SLO/SLA and monitoring strategy across Grafana and our wider observability tooling, building alert coverage that is meaningful rather than noisy. … shared-responsibility operating model between DevOps/SRE and product engineering. Vendor and cost management. Relationships across cloud, security, CI/CD, and observability platforms, including spend and renewal negotiation. You Have: 8+ years in DevOps, SRE, or infrastructure engineering, including 3+ years leading and developing engineering teams. Deep hands ...

Senior DevOps Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
involved in software releases by fully automating delivery pipelines that includes testing* Improve production stability and resiliency through process improvement and adoption of better observability tools* Support application teams in adopting a consistent set of tools, practices and technology to support an agile delivery model, including automation and incremental delivery … principles* Knowledge of third-party automation tools and technologies with relative pros and cons for all stages of SDLC, CI/CD pipelines and observability* Platforms; Windows Server, Amazon Linux, RHEL, Ubuntu* Proficiency in at least one of the following scripting languages; Python, GO, PowerShell, Bash, Groovy* Programming language with ...

Lead Software Engineer - Java, Go

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, United Kingdom
Salary
£ 70 K
will work with technologies including Java, Spring Boot, Go, Kubernetes, Terraform, Ansible, PostgreSQL, CockroachDB, Redis, CI/CD tooling, public cloud services, and observability platforms to solve complex engineering challenges and improve operational stability across critical infrastructure services.As a Lead Software Engineer at JPMorgan Chase within Infrastructure Platforms, Enterprise Production … product areas include Automation as a Service (AaaS), Ansible Automation Platform (AAP), AutoM8, hybrid cloud automation, public cloud and mainframe enablement, SRE and observability capabilities, and AI-driven service management accelerators.Job ResponsibilitiesExecutes software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build ...

Software Engineer/ SRE (Linux)

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
platform strategy. In this role, you'll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. You'll promote observability best practices and automate resolution of recurring issues, working closely with software engineering teams to support security, availability, and performance. Responsibilities include triaging issues, collaborating … implement, and maintain systems for high availability, scalability, and performance. Monitor and improve application reliability through proactive measures and incident response. Develop and maintain observability solutions (metrics, logging, tracing). Participate in on-call rotations and drive root cause analysis for incidents. Collaboration & Continuous Improvement Partner with engineering teams ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
team. Design, build and operate a secure, resilient and scalable Azure infrastructure estate, including the associated CI/CD, infrastructure-as-code, container and observability capabilities. Establish and evolve engineering standards, reference architecture, guardrails and reusable platform services, promoting consistency, automation, security, maintainability and operational excellence. Lead the design … cloud and/or on-premises infrastructure, including the development of reusable, maintainable infrastructure-as-code patterns. Experience implementing or operating monitoring, logging and observability capabilities using tools such as Splunk, Prometheus, Grafana or equivalent. Strong scripting and automation capabilities using PowerShell, Bash and/or similar languages, together with ...

Site Reliability Engineer, Studios

Location
Uxbridge, England, United Kingdom
embed reliability engineering practices across systems that support live, business‐critical environments. The successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders. Key Responsibilities And Accountabilities Design, build … platform services across on‐premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering‐led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators. Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical ...

Site Reliability Engineer, Studios

Hiring Organisation
iMG world
Location
London, United Kingdom
Salary
£ 70 K
help embed reliability engineering practices across systems that support live, business-critical environments.The successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders.Key Responsibilities and AccountabilitiesDesign, build, and maintain … infrastructure and platform services across on-premises and cloud environments.Improve service availability, latency, performance, and operational efficiency through engineering-led reliability practices.Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators.Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical services.Automate ...

Site Reliability Engineer, Studios

Location
Greater London, England, United Kingdom
embed reliability engineering practices across systems that support live, business-critical environments. The successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders. Key Responsibilities and Accountabilities Design, build … platform services across on-premises and cloud environments. Improve service availability, latency, performance, and operational efficiency through engineering-led reliability practices. Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators. Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
governance, automation, security controls, and developer experience established with our GCP platform. Develop and enhance platform services, self‐service capabilities, CI/CD tooling, observability solutions, and automation frameworks that improve developer productivity and operational effectiveness. Contribute to our AI transformation by building the platform capabilities, tooling, operational patterns … teams to define technical direction, establish engineering standards, and ensure platform capabilities meet current and future business requirements. Drive continuous improvement across platform reliability, observability, automation, security, cost efficiency, and developer experience through data‐driven engineering practices. Mentor and support engineers across the organisation, promoting platform engineering best practices, technical ...