501 to 525 of 850 Remote Observability Jobs

AWS DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600/day
/CD pipelines utilising GitHub Actions and Jenkins. Deploy, manage and troubleshoot containerised applications with Docker and Kubernetes (EKS). Implement monitoring, logging and observability solutions to improve platform visibility. Automate operational and deployment processes using Python and Bash scripting. Ensure platform reliability, scalability and high availability. Maintain security, compliance … Python automation experience. Solid understanding of IAM and cloud security principles. Experience designing and implementing CI/CD pipelines. Exposure to monitoring, logging and observability tools. Experience working within Agile delivery environments. Contract Overview £600 per day, Outside IR35. Initial 6-month contract with strong extension potential. Hybrid working model ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Lead Software Engineer | AI & Agentic Systems

Location
Greater London, England, United Kingdom
applying AI‐assisted engineering practices throughout the software development lifecycle, including specification, implementation, testing and code review. Understanding of responsible AI principles, including governance, observability and production monitoring. Experience integrating AI‐assisted development tools such as GitHub Copilot, Cursor, Claude Code or similar into day‐to‐day engineering workflows. Software … PostgreSQL, MongoDB and cloud platforms such as AWS and Azure. Strong knowledge of CI/CD, automated testing (unit, integration and end‐to‐end), observability and production operations. Understanding of secure software development, authentication/authorisation and DevSecOps practices. Familiarity with AI‐assisted Spec‐Driven Development (SDD). Strong analytical ...

DevOps Developer

Hiring Organisation
Stott & May Professional Search Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
£65,000
role would suit a developer who is building their commercial experience, is eager to learn and has a particular interest in Microsoft Azure, automation, observability and reliable digital services. Key Responsibilities Development & Delivery Deliver small to medium-sized application changes, enhancements, defect fixes and technical improvements. Design, develop and maintain … training, practical delivery and knowledge sharing. Develop knowledge across Azure, secure software engineering and DevOps practices. Suggest incremental improvements to developer experience, application security, observability, performance and team processes. Key Performance Indicators Unreviewed Pull Requests - percentage of Pull Requests that are not peer reviewed. Median Cycle Time - how long ...

Senior Platform Engineering Manager at Prolific

Location
United Kingdom
operational excellence, and innovation. Champion SRE Culture: Own availability and embed SRE principles across the organization, including defining SLOs, SLAs, error budgets, and enhancing observability and incident remediation. Platform & Developer Experience: Own the developer-facing platform — golden paths, self-service infrastructure, and internal tooling — so teams can provision, deploy … Infrastructure-as-Code (Terraform/Terragrunt, Crossplane), with GitOps workflows using tools like ArgoCD. Reliability & Architecture: Solid understanding of complex infrastructure and application architecture, observability principles, and incident management. Stability & Velocity: Experience balancing the need for platform stability and reliability with the goal of increasing developer productivity and velocity. Bridge ...

Lead GCP Engineer

Hiring Organisation
SoftPapaya
Location
London Area, United Kingdom
across the platform. • Drive adoption and maturity of CI/CD pipelines, Infrastructure as Code and automated deployment processes. www.softpayapa.com• Implement and improve platform observability, monitoring, alerting and operational support capabilities. • Develop and maintain reusable automation patterns and engineering standards. • Leverage Terraform and Infrastructure as Code principles to ensure repeatable … Infrastructure as Code expertise. • Strong experience with CI/CD tooling and deployment automation. • Experience building and operating resilient cloud platforms. • Strong understanding of observability, monitoring, incident management and operational excellence. Data Engineering & Integration • Experience designing secure data integration solutions. • Knowledge of APIs, databases, data pipelines and cloud-based data ...

Platform Engineer III - (Pipelines & Developer Experience)

Location
Leeds, England, United Kingdom
least privilege, policy‐as‐code, scanning). Improve developer experience: faster feedback loops, quality gates, ephemeral/preview environments, and great documentation. Instrument pipeline observability (Datadog or equivalent) and define SLOs (queue time, lead time, change fail rate, MTTR) to drive reliability. Automate IaC workflows (Terraform/Terragrunt) and integrate … implement and maintain CI/CD release pipelines. Experience with scripting and programming (.NET preferred; familiarity with Go, Python, PowerShell beneficial). Knowledge of observability tooling, chaos testing, and incident management. Strong analytical and problem‐solving abilities, with the capability to closely collaborate with engineering teams. Highly outcome‐oriented, pragmatic ...

Software Engineer Senior Consultant II - (Hybrid)

Location
Belfast City District, Northern Ireland, United Kingdom
continuous improvement initiatives. Contribute to Agile delivery processes including sprint planning, estimation, backlog refinement, and execution of engineering commitments. Promote DataOps practices, automation, observability, and engineering excellence across development teams. Serve as a technical subject matter expert for modern data platforms, Spark processing, Microsoft Fabric, and lakehouse architectures. Essential Skills … Ability to mentor junior engineers and provide technical leadership without direct people management responsibilities. Knowledge of streaming and batch data processing patterns. Experience with observability solutions including logging, monitoring, tracing, and performance analysis. Experience working in cloud-based analytics and data platform environments. Experience working in Agile delivery environments. Experience ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational ExcellenceDesign and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry).Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skillsFamiliarity with infrastructure ...

Senior Software Engineer - ST&S

Hiring Organisation
BP Energy
Location
South West London, London, United Kingdom
solutions from prototypes into production. Develop production-grade agentic workflows, including multi-agent architectures and MCP-based services. Build high-performance evaluation pipelines and observability frameworks to assess the accuracy, safety, reliability and latency of AI-powered applications. Find opportunities where AI, automation and data can deliver measurable business value … Actor Model, Event Sourcing and CQRS architectures. Experience building AI/ML-enabled applications and agentic systems in production. Experience developing evaluation, monitoring and observability frameworks for AI applications. Experience leading or mentoring a squad of engineers. Experience running business-critical applications in production and participating in operational support models. ...

Engineering Manager: Platform & Growth Lead (Hybrid)

Location
Bristol, England, United Kingdom
client journeys on React and React Native. You'll coach engineers, own delivery, guide architectural decisions, partner with product, and drive observability and high-velocity delivery in a regulated financial services environment. #J-18808-Ljbffr ...

Remote Backend Engineer - Scalable Systems & APIs

Location
Greater London, England, United Kingdom
with Engineering, Product, Infrastructure, and Data Science/ML teams, owning software from design through production, building APIs and microservices, and improving reliability and observability in a distributed system. #J-18808-Ljbffr ...

Lead Software Engineer - Hybrid, Share Options, £80k+

Location
Wigan, England, United Kingdom
production. You will lead a small cross-functional team, delivering features in a fast-paced SaaS environment. You will help modernise microservices, improve observability, and enhance testability using modern architectural approaches. Hybrid working with two days in the office is available. #J-18808-Ljbffr ...

MLOps Engineering Manager — Lead Scalable ML (Hybrid)

Location
Greater London, England, United Kingdom
Trainline in London is seeking an experienced MLOps Engineering Manager to build and lead a new team of engineers. You will shape deployment, observability, and scalable machine learning systems across the platform. You will collaborate with ML Engineers, Data Engineers, Software Engineers, Data Scientists, Product Managers and stakeholders to deliver ...

Principal Engineer: Cloud Architecture & DevOps (Azure)

Location
Manchester, England, United Kingdom
Azure. You will translate high‐level designs into workable solutions and guide engineers across the value stream. You’ll mentor engineers, uphold SOLID, drive observability and quality, and collaborate with Architecture and Product teams. Flexible hybrid working from Salford Quays (Manchester) with occasional in‐person sessions. #J-18808-Ljbffr ...

Senior AI Safety Engineer - Production Guardrails (Remote)

Location
United Kingdom
multimodal spaces. You’ll work with ML and full-stack engineers to build production-grade safety infrastructure from the ground up and ensure robustness, observability, and scalability. This is a product ownership role focusing on end-to-end execution, architecture, deployment, and monitoring of safety systems in a fast-moving ...

Principal Engineer, CSRE Provisioning (Remote, UK)

Location
Greater London, England, United Kingdom
shape platform lifecycle work, lead high‐risk initiatives, and define standards across a globally distributed estate. You will mentor engineers, drive reliability, DR and observability improvements, and participate in on‐call rotations while staying hands‐on with platforms. A strong background in SRE, distributed systems, and IaC is required. #J ...

Staff Backend Engineer for AI-Driven Context Layer SaaS

Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is seeking a Staff-level Backend Engineer to build production services for an AI-native context layer. This remote role offers autonomy, collaboration across a global team, and the opportunity to shape foundational architecture. You will design ingestion, storage, and retrieval ...

Investment Tech Engineer — AI & Data Pipelines (Hybrid)

Location
United Kingdom
platform. You will work on cloud-based tools (Streamlit/ReactJS), apply AI-assisted tooling, and help drive CI/CD, data quality, and observability across the lifecycle. Collaborating with Cloud Teams, Enterprise Architecture and Security, you will translate business requirements into robust technical solutions, build data pipelines for analytics ...

Remote Backend Engineer — Real-Time MMO Scale & Systems

Location
Greater London, England, United Kingdom
players connected, using GKE, Kafka, MongoDB and PostgreSQL, building stateful game servers and event-driven services to 10x current scale. You will optimize deployments, observability, and resilience, implement IaC with Terraform/Helm, and collaborate with a small senior team. #J-18808-Ljbffr ...

Senior Backend Engineer - Remote UK (Python/Go)

Location
United Kingdom
Python (Flask), Go, Ruby, and cloud services. You’ll collaborate with product and design in a remote-friendly environment while maintaining strong testing and observability practices. #J-18808-Ljbffr ...

Remote Serverless Backend Engineer - Fintech TypeScript & AWS

Location
United Kingdom
reliable microservices. You design REST APIs with type-safe schemas, model data for DynamoDB and PostgreSQL, and orchestrate event-driven workflows. You will ensure observability and quality through tests and tracing, with a collaborative, English-speaking team. #J-18808-Ljbffr ...

DV-Cleared SRE — Hybrid Cheltenham Contract (Infra & Cloud)

Location
Cheltenham, England, United Kingdom
government organisations. You will work with multiple feature teams and the BAU/Support group to evolve our cloud and on‐prem infrastructure, improve observability, and mitigate reliability risks. The role involves designing and maintaining scalable infrastructure, monitoring performance, and automating processes using Ansible across the full #J-18808-Ljbffr ...

Hybrid Backend Engineer — Commercial Planning

Location
Greater London, England, United Kingdom
planning platforms for Fashion, Home & Beauty, serving millions of customers and thousands of colleagues. Join us to craft robust, scalable backend services with modern observability and secure design. You will lead design and delivery of cloud-native solutions, mentor engineers, and collaborate with architecture and product teams in a hybrid ...

Senior AI Software Engineer - Hybrid, High-Impact Delivery

Location
Greater London, England, United Kingdom
secure production code, and mentor engineers and AI coding agents. The role blends deep software engineering with AI-native practices, overseeing code quality, testing, observability, deployment readiness, and governance while ensuring outputs are explainable and compliant with standards. #J-18808-Ljbffr ...