1,476 to 1,500 of 5,223 Permanent Observability Jobs

Software Engineer, Senior

Hiring Organisation
Infor
Location
Orpington, Greater London, UK
Employment Type
Full-time
integration patterns using modern cloud-native approaches. Evaluating and implementing emerging AI integration technologies, frameworks, and engineering practices. Helping establish engineering standards for security, observability, resiliency, and maintainability. Contributing to technical design discussions and architecture reviews. Mentoring and supporting other engineers as AI capabilities become more broadly adopted across … applications and services. Experience with cloud-native development on AWS, Azure, or similar cloud platforms. Strong understanding of software architecture, security, scalability, reliability, and observability principles. Experience working with event-driven architectures and asynchronous integration patterns. Experience working with relational databases and modern data access technologies. Experience developing reusable frameworks ...

Platform Engineer

Hiring Organisation
Springer Nature
Location
London, United Kingdom
Salary
£ 70 K
engage on providing a unified and standardised platform by using modern and open standards. Our approach is encompassed by defined core capabilities, such as observability, continuous integration, security and storage. We therefore closely collaborate in our department so that the core capabilities are not only tightly integrated but also provide … Engineering department at Springer Nature Technology, we provide platform engineering expertise to support core capabilities and use cases across run time to databases to observability, ci/cd, and SRE. You will join a multidisciplinary team with different nationalities, backgrounds and experience levels. We are a very distributed department, sometimes ...

Senior Platform Reliability Engineer

Location
Greater London, England, United Kingdom
platforms, ensuring services meet defined availability, performance, security, and compliance standards. It combines deep operational expertise with strong automation, Infrastructure as Code (IaC), and observability capability to reduce toil, improve recovery, and enable predictable service outcomes. What you will be doing Deliver standards for availability, latency, performance, capacity, and scalability. … closed. Drive infrastructure-as-code and automation across Azure and co-lo environments. Evolve image bakery pipeline for secure, repeatable server images. Embed observability using metrics, logs, traces, and alerting tools. Partner with SRE and helpdesk teams to deliver service. Oversee automated patching, vulnerability remediation, and configuration compliance. Introduce KPIs ...

Software Engineer (Next.js Playwright)

Location
Greater London, England, United Kingdom
maintain applications within Azure. Contribute to CI/CD pipelines using GitHub Actions and Azure DevOps. Monitor application performance and reliability using modern observability tools. Collaborate Across the Business Work closely with clinicians, product managers and fellow engineers to understand user needs. Contribute ideas that improve both the product … Experience with Playwright or other automated testing frameworks. Experience working within Azure. Experience in healthcare, NHS, or regulated environments. Familiarity with application monitoring and observability tools. Experience working in a SaaS or scale-up environment. Mindset Product-minded and user-focused. Pragmatic, collaborative and delivery-oriented. Takes ownership and sees ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Live Operations Associate Roku, Inc.

Location
Cambridge, England, United Kingdom
role sits at the intersection of live operations and platform reliability, so you will work alongside infrastructure, engineering teams, and service providers to support observability, automation, and encoding pipeline health as the FAST ecosystem scales. As a trusted point of contact for service partners, you will manage expectations, communicate clearly … resolved efficiently Managing ongoing FAST channel configurations, including schedule updates, metadata changes, feed ingestion, and EPG alignment Monitoring channel and pipeline health using observability tools; validating ingestion, triaging transcode and CDN delivery issues, and escalating incidents with clear context to engineering or peer technical teams Driving automation of repetitive operational ...

Forward Deployed AI Engineer

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
enabled systems. You'll bring deep expertise across modern full-stack technologies (.NET, Azure, SQL, React/Angular), along with experience in distributed systems, observability, and AI tooling such as LLMs, retrieval pipelines, agentic workflows, and platforms such as Anthropic Claude. Experience designing and deploying AI agents, leveraging Model Context … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Senior Specialist, Production Services Application Support Analyst

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Analytical and problem-solving skills with the ability to identify root causes and implement sustainable ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelinesImplement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloadsTune GPU and accelerator capacity, autoscaling, and cost efficiency ...

Applied AI Engineering Lead - VP, Markets Operations

Location
Greater London, England, United Kingdom
information Build and enhance robust AI services and infrastructure using modern engineering practices, including APIs, event‐driven patterns, CI/CD, Infrastructure-as-Code, observability, and automated testing Partner with AI researchers, data scientists, and software engineers to translate emerging AI capabilities into practical, reliable, and compliant enterprise applications Establish … data engineering concepts, ETL and data pipelines, structured and unstructured data, and integration with enterprise data platforms Experience with CI/CD, automated testing, observability, production monitoring, and operational readiness practices Familiarity with Infrastructure-as-Code solutions such as Terraform and cloud or container‐based deployment patterns Working knowledge ...

Senior Software Engineer

Hiring Organisation
Daniel James Resourcing Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
engineers. Where AI enters the product itself, youll help establish the controls required to take it into production responsibly, including validation, guardrails, permissions, observability, evaluation, failure behaviour and appropriate human oversight. What you'll bring Youll be an accomplished Software Engineer who has designed and delivered complex production systems … data modelling Docker, Kubernetes and Infrastructure as Code Event-driven architecture CI/CD, automated testing and modern engineering practices Production ownership, reliability and observability Technical leadership and mentoring other engineers The underlying brief prioritises strong C#/.NET experience but can consider engineers from another modern backend language ...

Senior Associate, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile (Scrum ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real‐time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Lead Software Engineer

Hiring Organisation
StepChange Debt Charity
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
practices while fostering a culture of ownership, collaboration, and continuous improvement. You'll also champion operational excellence by improving developer onboarding, enhancing monitoring and observability, supporting incident management and root cause analysis, and driving improvements in reliability, productivity, and delivery velocity. Alongside this, you'll work with stakeholders across … strong understanding of RESTful APIs, relational databases, Git-based development workflows, and software engineering best practice. Experience with CI/CD pipelines, GitHub Actions, observability tooling, and cloud-native architectures will enable you to drive quality and consistency across the engineering function. You'll thrive in an Agile environment, balancing ...

Senior Associate, Full-Stack Engineer

Location
Westminster, West End, United Kingdom
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile (Scrum ...

Senior Associate, Full-Stack Engineer

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … programming concepts and microservicesProficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker).Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile ...

Context Plane Python Engineer

Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end-to-end — from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands-on experience using enterprise-authorized AI-assisted software development ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
Sophos
Location
Belfast, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Senior Network Engineer- IP

Location
Ipswich, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Senior Network Engineer- IP

Location
Birmingham, England, United Kingdom
improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing automation to reduce operational toil, and embedding observability and reliability‐first practices across the team. You will champion and build effective working relationships, both internally and externally, to deliver business outcomes … tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building and maintaining ...

Vice President, DevOps Production Services

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
enterprise applications and ensure platform stability, resiliency, and availability. Monitor application health, system performance, batch jobs, interfaces, and alerts using enterprise monitoring and observability tools. Investigate, troubleshoot, and resolve production incidents within defined SLAs. Perform root cause analysis (RCA) for recurring issues and drive permanent fixes. Analyze production logs, identify … Cloud experience preferred. Knowledge of automation/scripting using Python, Shell, or PowerShell. Exposure to DevOps/SRE practices, CI/CD pipelines, and observability tooling. Strong communication skills with the ability to provide concise incident and executive status updates. Our culture allows us to run our company better ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
grabjobs
Location
Dollar, Clackmannanshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
grabjobs
Location
Skegness, Lincolnshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
grabjobs
Location
Cambridge, Cambridgeshire, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
grabjobs
Location
Saltford, Somerset, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...