501 to 525 of 864 Remote/Hybrid Observability Jobs

Engineering Manager, Payments

Location
Greater London, England, United Kingdom
work. Bonus points if You are familiar with any one of the following technologies : Scala, Java or Typescript. You are familiar with observability, tracking and data pipeline tools and methodologies. You have previously worked in an App first business. Additional Information Health + Mental Wellbeing PMI and cash plan healthcare ...

Systems Engineer - Service Providers

Hiring Organisation
Hewlett Packard Enterprise
Location
Manchester, UK
Employment Type
Full-time
service provider routing, with useful experience in one or more adjacent areas: data centre switching, network security, automation, network virtualisation, cloud networking or observability and assurance. Experience designing resilient, scalable and operationally supportable network solutions, and the ability to explain architecture choices and trade-offs clearly. Knowledge of network automation ...

Systems Engineer - Service Providers

Hiring Organisation
Hewlett Packard Enterprise
Location
Bristol, UK
Employment Type
Full-time
service provider routing, with useful experience in one or more adjacent areas: data centre switching, network security, automation, network virtualisation, cloud networking or observability and assurance. Experience designing resilient, scalable and operationally supportable network solutions, and the ability to explain architecture choices and trade-offs clearly. Knowledge of network automation ...

Systems Engineer - Service Providers

Hiring Organisation
Hewlett Packard Enterprise
Location
London, UK
Employment Type
Full-time
service provider routing, with useful experience in one or more adjacent areas: data centre switching, network security, automation, network virtualisation, cloud networking or observability and assurance. Experience designing resilient, scalable and operationally supportable network solutions, and the ability to explain architecture choices and trade-offs clearly. Knowledge of network automation ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
play a key role in building highly reliable, scalable, and observable infrastructure. This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience. You'll have genuine ownership and the opportunity to influence … infrastructure using Terraform and Infrastructure as Code principles Develop and optimise CI/CD pipelines using GitHub Actions Build and improve comprehensive monitoring and observability across the platform Implement and maintain effective logging, metrics, tracing, alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve ...

Lead AI Software Engineer

Location
Newbury, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle. What you will do Lead build execution within a squad, platform capability, or engineering domain. Translate … generated and engineer‐written code to ensure correctness, maintainability, security, and alignment with specifications. Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification. Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership. Ensure AI-generated outputs are explainable, traceable, secure ...

Senior Backend Engineer | AI Platform

Location
Greater London, England, United Kingdom
high degree of autonomy and ownership, as you'll be responsible for designing scalable solutions that empower multiple engineering teams while ensuring reliability, observability, and cost efficiency. What are we looking for: 5+ years of experience in Software Engineering, Backend Engineering, or Platform Engineering. Strong experience building and maintaining backend … LangChain, LangGraph, CrewAI, or similar. Experience working with cloud platforms such as Google Cloud Platform (preferred), AWS, or Azure. Strong understanding of system reliability, observability, monitoring, and incident management. Experience with Infrastructure as Code and cloud-native architectures. Previous experience working within a Platform Engineering team is a strong plus. ...

Lead AI Software Engineer London, United Kingdom Value Stream Engineering Posted 12 hours ago

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle.## **What you will do*** Lead build execution within a squad, platform capability, or engineering domain.* Translate … generated and engineer-written code to ensure correctness, maintainability, security, and alignment with specifications.* Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification.* Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership.* Ensure AI-generated outputs are explainable, traceable, secure ...

AWS DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600/day
/CD pipelines utilising GitHub Actions and Jenkins. Deploy, manage and troubleshoot containerised applications with Docker and Kubernetes (EKS). Implement monitoring, logging and observability solutions to improve platform visibility. Automate operational and deployment processes using Python and Bash scripting. Ensure platform reliability, scalability and high availability. Maintain security, compliance … Python automation experience. Solid understanding of IAM and cloud security principles. Experience designing and implementing CI/CD pipelines. Exposure to monitoring, logging and observability tools. Experience working within Agile delivery environments. Contract Overview £600 per day, Outside IR35. Initial 6-month contract with strong extension potential. Hybrid working model ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Lead Software Engineer | AI & Agentic Systems

Location
Greater London, England, United Kingdom
applying AI‐assisted engineering practices throughout the software development lifecycle, including specification, implementation, testing and code review. Understanding of responsible AI principles, including governance, observability and production monitoring. Experience integrating AI‐assisted development tools such as GitHub Copilot, Cursor, Claude Code or similar into day‐to‐day engineering workflows. Software … PostgreSQL, MongoDB and cloud platforms such as AWS and Azure. Strong knowledge of CI/CD, automated testing (unit, integration and end‐to‐end), observability and production operations. Understanding of secure software development, authentication/authorisation and DevSecOps practices. Familiarity with AI‐assisted Spec‐Driven Development (SDD). Strong analytical ...

DevOps Developer

Hiring Organisation
Stott & May Professional Search Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
£65,000
role would suit a developer who is building their commercial experience, is eager to learn and has a particular interest in Microsoft Azure, automation, observability and reliable digital services. Key Responsibilities Development & Delivery Deliver small to medium-sized application changes, enhancements, defect fixes and technical improvements. Design, develop and maintain … training, practical delivery and knowledge sharing. Develop knowledge across Azure, secure software engineering and DevOps practices. Suggest incremental improvements to developer experience, application security, observability, performance and team processes. Key Performance Indicators Unreviewed Pull Requests - percentage of Pull Requests that are not peer reviewed. Median Cycle Time - how long ...

Senior Platform Engineering Manager at Prolific

Location
United Kingdom
operational excellence, and innovation. Champion SRE Culture: Own availability and embed SRE principles across the organization, including defining SLOs, SLAs, error budgets, and enhancing observability and incident remediation. Platform & Developer Experience: Own the developer-facing platform — golden paths, self-service infrastructure, and internal tooling — so teams can provision, deploy … Infrastructure-as-Code (Terraform/Terragrunt, Crossplane), with GitOps workflows using tools like ArgoCD. Reliability & Architecture: Solid understanding of complex infrastructure and application architecture, observability principles, and incident management. Stability & Velocity: Experience balancing the need for platform stability and reliability with the goal of increasing developer productivity and velocity. Bridge ...

Lead GCP Engineer

Hiring Organisation
SoftPapaya
Location
London Area, United Kingdom
across the platform. • Drive adoption and maturity of CI/CD pipelines, Infrastructure as Code and automated deployment processes. www.softpayapa.com• Implement and improve platform observability, monitoring, alerting and operational support capabilities. • Develop and maintain reusable automation patterns and engineering standards. • Leverage Terraform and Infrastructure as Code principles to ensure repeatable … Infrastructure as Code expertise. • Strong experience with CI/CD tooling and deployment automation. • Experience building and operating resilient cloud platforms. • Strong understanding of observability, monitoring, incident management and operational excellence. Data Engineering & Integration • Experience designing secure data integration solutions. • Knowledge of APIs, databases, data pipelines and cloud-based data ...

Platform Engineer III - (Pipelines & Developer Experience)

Location
Leeds, England, United Kingdom
least privilege, policy‐as‐code, scanning). Improve developer experience: faster feedback loops, quality gates, ephemeral/preview environments, and great documentation. Instrument pipeline observability (Datadog or equivalent) and define SLOs (queue time, lead time, change fail rate, MTTR) to drive reliability. Automate IaC workflows (Terraform/Terragrunt) and integrate … implement and maintain CI/CD release pipelines. Experience with scripting and programming (.NET preferred; familiarity with Go, Python, PowerShell beneficial). Knowledge of observability tooling, chaos testing, and incident management. Strong analytical and problem‐solving abilities, with the capability to closely collaborate with engineering teams. Highly outcome‐oriented, pragmatic ...

Software Engineer Senior Consultant II - (Hybrid)

Location
Belfast City District, Northern Ireland, United Kingdom
continuous improvement initiatives. Contribute to Agile delivery processes including sprint planning, estimation, backlog refinement, and execution of engineering commitments. Promote DataOps practices, automation, observability, and engineering excellence across development teams. Serve as a technical subject matter expert for modern data platforms, Spark processing, Microsoft Fabric, and lakehouse architectures. Essential Skills … Ability to mentor junior engineers and provide technical leadership without direct people management responsibilities. Knowledge of streaming and batch data processing patterns. Experience with observability solutions including logging, monitoring, tracing, and performance analysis. Experience working in cloud-based analytics and data platform environments. Experience working in Agile delivery environments. Experience ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational ExcellenceDesign and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry).Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skillsFamiliarity with infrastructure ...

Senior Software Engineer - ST&S

Hiring Organisation
BP Energy
Location
South West London, London, United Kingdom
solutions from prototypes into production. Develop production-grade agentic workflows, including multi-agent architectures and MCP-based services. Build high-performance evaluation pipelines and observability frameworks to assess the accuracy, safety, reliability and latency of AI-powered applications. Find opportunities where AI, automation and data can deliver measurable business value … Actor Model, Event Sourcing and CQRS architectures. Experience building AI/ML-enabled applications and agentic systems in production. Experience developing evaluation, monitoring and observability frameworks for AI applications. Experience leading or mentoring a squad of engineers. Experience running business-critical applications in production and participating in operational support models. ...

Engineering Manager: Platform & Growth Lead (Hybrid)

Location
Bristol, England, United Kingdom
client journeys on React and React Native. You'll coach engineers, own delivery, guide architectural decisions, partner with product, and drive observability and high-velocity delivery in a regulated financial services environment. #J-18808-Ljbffr ...

Remote Backend Engineer - Scalable Systems & APIs

Location
Greater London, England, United Kingdom
with Engineering, Product, Infrastructure, and Data Science/ML teams, owning software from design through production, building APIs and microservices, and improving reliability and observability in a distributed system. #J-18808-Ljbffr ...

Lead Software Engineer - Hybrid, Share Options, £80k+

Location
Wigan, England, United Kingdom
production. You will lead a small cross-functional team, delivering features in a fast-paced SaaS environment. You will help modernise microservices, improve observability, and enhance testability using modern architectural approaches. Hybrid working with two days in the office is available. #J-18808-Ljbffr ...

MLOps Engineering Manager — Lead Scalable ML (Hybrid)

Location
Greater London, England, United Kingdom
Trainline in London is seeking an experienced MLOps Engineering Manager to build and lead a new team of engineers. You will shape deployment, observability, and scalable machine learning systems across the platform. You will collaborate with ML Engineers, Data Engineers, Software Engineers, Data Scientists, Product Managers and stakeholders to deliver ...

Principal Engineer: Cloud Architecture & DevOps (Azure)

Location
Manchester, England, United Kingdom
Azure. You will translate high‐level designs into workable solutions and guide engineers across the value stream. You’ll mentor engineers, uphold SOLID, drive observability and quality, and collaborate with Architecture and Product teams. Flexible hybrid working from Salford Quays (Manchester) with occasional in‐person sessions. #J-18808-Ljbffr ...

Principal Engineer, CSRE Provisioning (Remote, UK)

Location
Greater London, England, United Kingdom
shape platform lifecycle work, lead high‐risk initiatives, and define standards across a globally distributed estate. You will mentor engineers, drive reliability, DR and observability improvements, and participate in on‐call rotations while staying hands‐on with platforms. A strong background in SRE, distributed systems, and IaC is required. #J ...

Senior AI Safety Engineer - Production Guardrails (Remote)

Location
United Kingdom
multimodal spaces. You’ll work with ML and full-stack engineers to build production-grade safety infrastructure from the ground up and ensure robustness, observability, and scalability. This is a product ownership role focusing on end-to-end execution, architecture, deployment, and monitoring of safety systems in a fast-moving ...