1,376 to 1,400 of 4,717 Permanent Observability Jobs

Senior Associate, Full-Stack Engineer

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … programming concepts and microservicesProficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker).Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile ...

Context Plane Python Engineer

Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end-to-end — from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands-on experience using enterprise-authorized AI-assisted software development ...

Senior Network Engineer- IP

Location
Birmingham, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Senior Network Engineer- IP

Location
Ipswich, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Vice President, DevOps Production Services

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
enterprise applications and ensure platform stability, resiliency, and availability. Monitor application health, system performance, batch jobs, interfaces, and alerts using enterprise monitoring and observability tools. Investigate, troubleshoot, and resolve production incidents within defined SLAs. Perform root cause analysis (RCA) for recurring issues and drive permanent fixes. Analyze production logs, identify … Cloud experience preferred. Knowledge of automation/scripting using Python, Shell, or PowerShell. Exposure to DevOps/SRE practices, CI/CD pipelines, and observability tooling. Strong communication skills with the ability to provide concise incident and executive status updates. Our culture allows us to run our company better ...

Senior Network Engineer- IP

Location
Greater London, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Remote, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Nottingham, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Buckley, Flintshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Cambridge, Cambridgeshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Bromsgrove, Worcestershire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Highbridge, Somerset, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Bedford, Bedfordshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Location
Cumbria, United Kingdom
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Location
Carmarthenshire, United Kingdom
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Location
Sussex, United Kingdom
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Dinnington, South Yorkshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Location
Consett, Durham, United Kingdom
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Lead Software Engineer

Location
Manchester, England, United Kingdom
high standards through code reviews, mentoring engineers, and contributing to CI/CD and testing best practices Applying site reliability engineering principles to improve observability, resilience, and incident response Monitoring, profiling, and optimising application performance, including capacity planning to meet SLAs Identifying technical debt and recommending improvements to systems, processes … containerisation Strong debugging, performance optimisation, and problem-solving skills across distributed systems Experience mentoring engineers and contributing to technical leadership within teams Knowledge of observability, monitoring, and site reliability engineering practices Experience with Java and/or telecoms technologies (e.g. VoIP, WebRTC) is highly beneficial What do we offer ...

DevOps Engineer

Location
Greater London, England, United Kingdom
build and operate the infrastructure that every engineering team at Invisible depends on — Kubernetes, CI/CD, identity, secrets management, networking, and observability — along with the internal tooling that allows engineers and AI coding agents to ship safely and quickly. The Platform team is small relative to the surface area … entire engineering organisation. We are looking for engineers who don’t just apply best practices in IaC, Kubernetes, CI/CD and observability, but understand the problems those practices were designed to solve. Most were built around the pace and failure modes of human engineers. Those assumptions are changing ...

Kubernetes Platform Engineer

Location
Greater London, England, United Kingdom
pipelines (ArgoCD, FluxCD) for safe, auditable changes. Drive Infrastructure as Code practices with Terraform and Helm for reliable and repeatable builds. Heavily embed observability using Prometheus, Grafana, and OpenTelemetry to make systems measurable and reliable. Stay ahead of Kubernetes evolution by testing and adopting new versions and features early. Collaborate … including namespace isolation, network policies, and Open Policy Agent. Proven ability to troubleshoot complex performance and reliability issues across infrastructure and workloads. Experience with observability tools such as Prometheus, Grafana, and OpenTelemetry to monitor cluster metrics and health. Great communication skills, with experience collaborating with internal platform users to gather ...

Software Engineer (EMS)

Location
Greater London, England, United Kingdom
improve performance. Distributed Systems: Maintain distributed, fault-tolerant systems with high availability across multiple regions, ensuring data consistency and low-latency communication. Monitoring & Observability: Implement real-time monitoring, telemetry and alerting systems for proactive issue detection and resolution. Collaboration: Work closely with product, QA and infrastructure teams to align technical … Azure) and tools for multi-region deployments . A working knowledge of databases and caches (PostgreSQL, SQLite, Redis, Zookeeper). Experience with observability tools (e.g. Prometheus, Grafana, ELK stack). Soft Skills: Strong analytical and problem-solving abilities. Excellent communication and collaboration skills, with a track record of working ...

Software Engineer - Software Delivery

Location
England, United Kingdom
involves things like Working closely with internal engineers to identify pain points Making sure the product experience is as good as possible Setting up observability around how the platform is performing but also how users are interacting with the platform Experience creating abstractions to simplify the local developer workflow … build tooling like kustomize or helm is also meriting. Experience integrating software with Google Cloud Platform, AWS and Azure Some experience in common software observability practices such as tracing, logging and metrics exporting. #J-18808-Ljbffr ...

Kubernetes Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
pipelines (ArgoCD, FluxCD) for safe, auditable changes. Drive Infrastructure as Code practices with Terraform and Helm for reliable and repeatable builds. Heavily embed observability using Prometheus, Grafana, and OpenTelemetry to make systems measurable and reliable. Stay ahead of Kubernetes evolution by testing and adopting new versions and features early. Collaborate … including namespace isolation, network policies, and Open Policy Agent. Proven ability to troubleshoot complex performance and reliability issues across infrastructure and workloads. Experience with observability tools such as Prometheus, Grafana and OpenTelemetry to monitor cluster metrics and health. Great communication skills, with experience collaborating with internal platform users to gather ...

Engineering & Delivery Lead - Risk & Compliance

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
enabled engineering practices across design, development, testing, security, documentation and operations with governance and traceability Strengthen production stability and service resiliency by improving observability, operability, technical debt management and controlled release practices Establish a strong technology control environment aligned to our client standards and regulatory expectations and address gaps … systems including low-latency architectures and synchronous and asynchronous execution models Use strong knowledge of containers, Kubernetes, cloud, service mesh, event streaming and production observability to improve resilience and operability Show experience automating delivery and operations including CI/CD, deployment automation, performance testing, rollback and controlled production releases Apply ...