2,476 to 2,500 of 5,503 Permanent Observability Jobs

Remote Sr. Software Engineer, Fullstack (UK)

Location
Brora, Inverness-shire, United Kingdom
post-incident reviews in a "you build it, you run it" environment. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Partner with Product Management and Design to translate business requirements into scalable technical solutions. Minimum Qualifications Bachelor's degree in Computer Science … HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Kubernetes, Docker, and Helm. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Experience in enterprise SaaS organisations, particularly HR Tech or regulated ...

Remote Sr. Software Engineer, Fullstack (UK)

Location
Sutton, South West London, United Kingdom
post-incident reviews in a "you build it, you run it" environment. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Partner with Product Management and Design to translate business requirements into scalable technical solutions. Minimum Qualifications Bachelor's degree in Computer Science … HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Kubernetes, Docker, and Helm. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Experience in enterprise SaaS organisations, particularly HR Tech or regulated ...

Staff Engineer, iOS

Location
United Kingdom
ready mobile systems. Raise the bar for reliability and execution: Shape how mobile work is designed, reviewed, tested, shipped, and operated, including the automation, observability, rollout practices, and AI‐enabled validation needed to move quickly with confidence. Shape mobile system contracts: Bring enough Android and backend context to align teams … driven testing systems to improve engineering velocity, quality, and regression confidence. Strong reliability and release‐quality mindset, including experience with testing strategy, automation, observability, CI/CD, and production mobile quality practices. Track record of leading significant technical initiatives, setting engineering standards, mentoring engineers, and raising the technical bar across ...

Director of Software Engineering

Hiring Organisation
Spire Global
Location
Glasgow, UK
Employment Type
Full-time
infrastructureStay hands-on: review code, prototype solutions, and get into the details when it mattersEstablish engineering standards across code quality, system design, testing, and observability, and hold the team to themBe the person engineers come to when the problem is genuinely hardTeam Building & Culture Recruit, develop, and retain a team … monitoring problemsExperience writing performance software in RustBackground in space systems, aerospace, or highly constrained real-time environmentsExperience building data lakes, telemetry platforms, or observability infrastructure at scaleA history of leading teams through technical transformations and not just maintaining the status quoSpire operates a hybrid work model, and this position will ...

Staff Software Engineer

Location
Greater London, England, United Kingdom
from problem framing and design review through to production rollout, migration and decommissioning of what came before Drive operational excellence by improving observability, alerting, capacity planning, load and chaos testing, and incident response, and by closing the loop on the root causes you find Raise the engineering bar across teams … development (we use AWS & Azure), containers, orchestration and infrastructure as code Experience owning production systems on call, with a strong sense of ownership for observability, monitoring and incident response Experience leading complex migrations or re-architectures of live, business-critical systems with no loss of availability or data integrity Excellent ...

Staff Software Engineer - AI Agents (Satori)

Location
Belfast City District, Northern Ireland, United Kingdom
ready" means at Proofpoint: the eval suites, safety checks, telemetry, and AuthX that gate every rollout. The primitives you set — from prompt management to observability — will shape how mission‐critical agentic software ships across the company. You will partner with product tech leads to turn one‐off integrations into reusable … similar) and keep up with the field. Prior experience taking agentic or LLM‐based systems to production — not just prototypes — including the eval, observability, and rollback story. Deep experience with Python; TypeScript/Node a plus. Hands‐on work with agent frameworks and eval tooling. Track record of debugging distributed ...

Principal Software Engineer - Squad Lead Engineer

Location
Greater London, England, United Kingdom
complete complex bug fixes and performance improvements Define and uphold Definition of Ready/Done including code quality, automated test coverage, security checks, and observability Establish/maintain CI/CD pipelines, quality gates, and sensible branching/release strategies Drive a pragmatic quality strategy: test pyramid balance, contract tests … Windows Experience with relational and non‐relational data stores, performance tuning, and data modelling Knowledge of CI/CD platforms, containers, cloud technologies, observability, and monitoring practices Understanding of secure coding, performance optimisation, reliability engineering, and incident response Work in a Way That Works for You We promote a healthy ...

Principal Engineer

Location
Milton Keynes, England, United Kingdom
Azure services, containers, application services, messaging, API management, relational and NoSQL data platforms. Improve continuous integration, continuous delivery, automated quality controls, infrastructure management, observability and release safety. Promote cost‐effective architecture and operational practices, including capacity planning, service ownership and FinOps awareness. Lead the adoption of reusable global components … queues, API Management, Azure SQL, Cosmos DB, containers and monitoring services. Strong understanding of CI/CD, automated testing, code quality controls, observability, incident diagnosis, reliability engineering and DevSecOps practices. Knowledge of application security, identity and access management, OWASP guidance, threat modelling, privacy by design and relevant compliance requirements. Experience ...

Senior Cloud Platform Engineer

Location
Milton Keynes, England, United Kingdom
available cloud platform solutions Leading the implementation of cloud infrastructure, automation and Infrastructure as Code Enhancing CI/CD pipelines and deployment capabilities Driving observability, monitoring and platform security best practices Supporting cloud transformation and platform modernisation initiatives Providing technical leadership during major incidents and service restoration activities Mentoring engineers ...

ML Data & Platform Engineer

Location
United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

Cloud Advisory Architecture Associate Manager

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
where GenAI and Agentic play a role. Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps. Manage & Mentor: Lead teams of architects and engineers, providing technical coaching, career counselling, performance management, and coaching ...

ML Data & Platform Engineer

Location
Cambridge, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

Software Engineer - AI Agents (Satori)

Location
Belfast City District, Northern Ireland, United Kingdom
means at Proofpoint: the eval suites, safety checks, telemetry, and AuthX that gate every rollout. The primitives you help shape — from prompt management to observability — influence how mission-critical agentic software ships across the company. You will partner with product tech leads to turn one-off integrations into reusable platform … LangChain, or similar) and keep up with the field. Experience contributing to agentic or LLM-based systems in production, including familiarity with the eval, observability, and rollback story. Demonstrated learning velocity and curiosity; driven to understand systems beyond surface-level usage at enterprise scale. Development experience with Python; TypeScript/ ...

Data Scientist - BAU Analytics

Hiring Organisation
Executive Facilities
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£500.00 per day
sources that help build the AI/ML user experience Build and maintain scalable data pipelines for BigQuery using Cloud Composer and Airflow. Provide observability and monitoring using Monte Carlo and Looker as well as operational tools such as NewRelic and Splunk, driving reliability, data quality, data quality and robustness. ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
captures and governs new categories of data (e.g. smart badge telemetry, location/proximity signals) in a privacy-compliant, lawful-basis-aware way Build observability into data systems — structured logging, freshness and quality monitoring, SLOs, and pipeline health dashboards Champion high-quality technical communication: proposals, specifications, and documentation that other … product engineering teams, even if your focus is data and AI infrastructure DevOps fluency: AWS, Kubernetes/EKS, Terraform, CI/CD pipelines Excellent observability practices — structured logging, metrics, distributed tracing, SLOs (Datadog, Sentry) Feature flags, canary deployments, and gradual rollout patterns Track record of driving data quality, governance ...

Data Scientist - BAU Analytics

Location
City Of London, England, United Kingdom
sources that help build the AI/ML user experience Build and maintain scalable data pipelines for BigQuery using Cloud Composer and Airflow. Provide observability and monitoring using Monte Carlo and Looker as well as operational tools such as NewRelic and Splunk, driving reliability, data quality, data quality and robustness. ...

Software Engineer - Affiliate Operations

Location
Greater London, England, United Kingdom
remain highly reliable while evolving to support new markets and acquisition channels. Reliability is fundamental to everything we build. We invest heavily in automation, observability, and operational excellence to reduce manual effort. We are also exploring how AI can transform our engineering productivity and marketing platforms. We operate … improve platform effectiveness, launch new capabilities, and reduce manual operational effort across the business. Improve You'll continuously improve our systems through better observability, automation, and thoughtful refactoring. You'll help evolve our architecture and engineering practices to ensure our platforms remain resilient as they scale. Own You'll take ...

Cloud DevOps Platform Engineer

Location
Greater London, England, United Kingdom
their failure modes, and a good understanding of network concepts and fundamentals Expertise in managing and maintaining Kubernetes clusters in production Excellent knowledge on observability tooling and best practices Experience working within or alongside engineering teams to deliver Preferred qualifications, capabilities and skills: Systems and database experience, with an understanding ...

Technical Lead, Lending & Savings

Hiring Organisation
Blockchain
Location
London, UK
Employment Type
Full-time
performance, security, and maintainability. Remain hands-on, contributing production-quality code and reviewing critical changes. Drive engineering best practices across system design, testing, deployment, observability, and operational excellence. Mentor engineers through code reviews, technical coaching, and day-to-day leadership. Engineering DeliveryOwn the delivery of technical initiatives from design through … design. Experience working with Redis or other NoSQL technologies. Deep understanding of microservices architecture, APIs, distributed systems, and cloud-native applications. Experience with monitoring, observability, incident response, and production operations. Strong debugging and performance optimisation skills. Demonstrated experience shipping reliable production systems that process financial transactions. LeadershipExperience leading engineering teams ...

Site Reliability Engineer (Edv) - National Security

Location
Cheltenham, England, United Kingdom
keep broadening their technical remit. What you'll be working with AWS Kubernetes Terraform Linux CI/CD Python/Bash Monitoring and observability Automation Reliability and performance engineering You’ll be working across secure, live National Security environments, helping teams improve deployment, resilience, monitoring and operational performance. ...

DevOps and Infrastructure Engineer

Hiring Organisation
Sanderson Government and Defence
Location
Gloucestershire, South West, United Kingdom
Employment Type
Permanent
cloud-based solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver ...

Senior Full Stack Engineer (Realtime & Voice) Customer Experience Platform

Location
Greater London, England, United Kingdom
Build the safety and compliance plumbing enterprise partners audit, including guardrails, content filtering, and PII redaction integration points Keep revenue-critical deployments healthy through observability, alerting, incident response, and SLA performance Build the platform capabilities forward-deployed engineers configure for partner telephony integrations and go-lives Raise the engineering … standard part of their workflow, with the judgment to review, correct, and own everything that ships An operable-systems mindset, covering SLAs, observability, on-call rotations, and rollback plans The ability to break down complex problems, make pragmatic tradeoffs, and ship iteratively, backed by strong communication across product, design ...

MongoDB Site Reliability Engineer

Location
Knutsford, England, United Kingdom
paced environment, your role will be essential to ensuring our infrastructure remains resilient, secure, and scalable. You’ll work on automating operations, enhancing system observability, and driving continuous improvements that reduce downtime and improve efficiency. If you’re motivated by solving, multi-layered problems and building systems that perform reliably ...

Senior Applied AI Engineer (Defence Contractor)

Location
United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Senior Applied AI Engineer (Defence Contractor)

Location
Greater London, England, United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...