401 to 425 of 858 Remote/Hybrid Observability Jobs

Senior Backend Developer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Software Engineering Manager

Hiring Organisation
Halian Technology Limited
Location
Central London, London, United Kingdom
Employment Type
Permanent
making sound architectural and design decisions. Champion software quality, security, performance, and operational excellence. Encourage modern engineering practices, including CI/CD, automated testing, observability, and cloud-native development. Stakeholder Management Build strong relationships with business and technology stakeholders. Communicate progress, risks, and dependencies effectively. Align engineering activities with organisational … would be beneficial: .NET, Java, Python, or Node.js, React Microservices architecture RESTful APIs Kubernetes and Docker AWS, Azure, or GCP CI/CD tooling Observability and monitoring platforms Modern data platforms and event-driven architectures There is a 2 - 3 stage interview process, with interview slots now available with ...

AWS Cloud Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
manage containerised environments using Docker and Kubernetes Build and improve CI/CD pipelines and automated delivery workflows Implement monitoring, logging, tracing and observability best practices Optimise cloud environments for performance, reliability and cost efficiency Support incident response, troubleshooting and root cause analysis Apply cloud security best practices and governance … Experience working with backend services developed in Python Experience with Docker and Kubernetes Experience building and maintaining CI/CD pipelines Strong understanding of observability, monitoring and logging practices Experience troubleshooting issues across both cloud infrastructure and application layers Knowledge of cloud security principles and best practices Strong communication ...

Site Reliability Engineer

Location
City of Westminster, England, United Kingdom
help drive a culture of continuous improvement. You'll also play a key role in scaling Curve's platform for millions of customers, improving observability, accelerating engineering teams through automation, and ensuring our services remain secure, resilient and highly available. Why Join us? If you think all banks … Experience working with Amazon Web Services and Google Cloud Platform Experience supporting Kubernetes (EKS) environments and service mesh technologies such as Istio Knowledge of observability tooling including Prometheus, Grafana or Coralogix Experience with PostgreSQL, MongoDB or HashiCorp Vault Experience using GitLab, Flux or Helm within CI/CD pipelines Knowledge ...

Lead Platform Operations Engineer

Location
Greater London, England, United Kingdom
maintain platform standards, patterns, and best practices Own Platform Reliability, Security & Performance Lead incident response, root cause analysis, and platform improvements Implement robust monitoring, observability, and alerting strategies Drive security improvements aligned to ISO27001, SOC2, and modern SDLC practices Ensure strong governance across infrastructure, applications, and data Deliver Scalable & Secure … Experience implementing security tooling (SAST, DAST, container scanning, WAF) Strong knowledge of cloud security, encryption, TLS/SSL, certificates, and access control Experience with observability, monitoring, and alerting tools Security & Compliance Practical experience implementing ISO27001 and SOC2 controls Knowledge of OWASP methodologies and secure development lifecycle practices Experience with vulnerability ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
help drive a culture of continuous improvement. You’ll also play a key role in scaling Curve’s platform for millions of customers, improving observability, accelerating engineering teams through automation, and ensuring our services remain secure, resilient and highly available. What we’re looking for? Experience deploying production-ready applications … Experience working with Amazon Web Services and Google Cloud Platform Experience supporting Kubernetes (EKS) environments and service mesh technologies such as Istio Knowledge of observability tooling including Prometheus, Grafana or Coralogix Experience with PostgreSQL, MongoDB or HashiCorp Vault Experience using GitLab, Flux or Helm within CI/CD pipelines Knowledge ...

Staff Engineer - Money, Risk & Payment Ancillaries Yuno Totalmente remoto · Mundial ayer

Location
Greater London, England, United Kingdom
velocity never compromises reliability You Build It, You Run It — champion a YBIYRI culture: your teams design for multi-tenancy, elastic scaling, and automated observability, and own their microservices in production Mentorship and cross-functional collaboration — mentor engineers and partner with Product, Data/ML, Compliance, and Staff Engineers across … Java APIs — gRPC, REST Frameworks — Spring Boot, Spring WebFlux; Go standard library Messaging — Apache Kafka, SQS Databases — PostgreSQL, Redis Infrastructure — AWS, Kubernetes, Docker, Terraform Observability — Datadog, OpenTelemetry CI/CD — GitHub Actions, ArgoCD Version Control — Git/GitHub What We Offer at Yuno Competitive Compensation Remote Work — you can work ...

Senior Devops Engineer

Location
Greater London, England, United Kingdom
pipelines for zero‐touch deployments. · Security hardening – lead security-posture reviews, implement GuardDuty, CloudWatch and IAM best practices. · SRE & monitoring – uphold SLAs through observability stacks, proactive alerting and performance tuning of distributed systems. · Collaboration & enablement – automate repetitive tasks, mentor developers and champion DevSecOps best practice across teams. · Policy & audit ownership …/Bonus Experience with MLOps/LLMOps (Softwares such as Sagemaker, Kubeflow or ZenML). Deployment of on‐premise Kubernetes Prometheus (or other stacks) observability Experience with AWS Karpenter & Compute Optimizer Compliance literacy - ISO 27001, NIST SSDF/OWASP SAMM, GDPR basics Why Oxford Dynamics? Join the most exciting growth ...

Service Reliability Engineer - London

Hiring Organisation
Fitch Ratings
Location
London, UK
Employment Type
Full-time
Core Engineering to architect and govern GitHub Actions CI/CD with quality gates, canary/blue‐green strategies, and AI‐assisted redeploy checksOwn observability in Datadog—define SLIs/SLOs, dashboards, alerting, and MS Teams integrations—and reduce incidents via telemetry-driven automation and blameless postmortems. Champion AI‐enabled … compliance implementation across CIS, NIST, ISO 27001, with automated remediation integrated via CSPM tools (e.g., Wiz). Applying AI in CI/CD, observability, and incident response using AWS Bedrock/SageMaker and Model Context Protocol (MCP).Hands-on Agile delivery experience, actively participating in stand-ups and sprint ceremonies. ...

Service Reliability Engineer - Manchester

Hiring Organisation
Fitch Group
Location
Manchester, United Kingdom
Employment Type
Full Time
Engineering to architect and govern GitHub Actions CI/CD with quality gates, canary/blue‐green strategies, and AI‐assisted redeploy checks Own observability in Datadog—define SLIs/SLOs, dashboards, alerting, and MS Teams integrations—and reduce incidents via telemetry-driven automation and blameless postmortems. Champion AI‐enabled … compliance implementation across CIS, NIST, ISO 27001, with automated remediation integrated via CSPM tools (e.g., Wiz). Applying AI in CI/CD, observability, and incident response using AWS Bedrock/SageMaker and Model Context Protocol (MCP). Hands-on Agile delivery experience, actively participating in stand-ups and sprint ...

IAM Developer - Sheffield/Hybrid - £475.00 Per Day Umbrella

Hiring Organisation
Click
Location
Sheffield, Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 475 Annual
We are recruiting for a IAM Developer for a leading financial organisation based in Sheffield, this will be a hybrid role with 3 days a week on site. Role Summary We are looking for a ...

Senior Hybrid Azure Network Reliability Engineer

Location
Wimbledon, England, United Kingdom
Wimbledon office. You will own the reliability, performance and security of our hybrid Azure network, lead major incident resolution, and drive automation and observability across our estate. You’ll work with Azure networking, SD-WAN, security appliances and IaC tools, mentoring engineers and collaborating with Cloud, Security and Engineering teams ...

Remote AWS Cloud Architect & Engineer: Modernize & Optimize

Location
United Kingdom
Kingdom. Design and deliver modern cloud solutions across Europe, blending hands-on engineering with architectural thinking. You will work across IaC, CI/CD, observability and cost optimization, mentoring teams and translating complex concepts for both technical and non-technical stakeholders. Fully remote across Europe. #J-18808-Ljbffr ...

Lead Java Developer — Real-Time Risk & Cloud (Hybrid)

Location
Greater London, England, United Kingdom
full lifecycle from design to production support, integrating new analytics and data sets across global teams. The role emphasizes scalable microservices, streaming data, and observability with ELK, Prometheus and Grafana. Hybrid work model and competitive benefits are offered. #J-18808-Ljbffr ...

Lead DevOps Engineer for Low-Latency Trading Platform (Hybrid)

Location
Greater London, England, United Kingdom
infrastructure across cloud and on-prem environments. The role emphasizes reliability, security, and fast delivery, with hands-on leadership across DevOps, platform engineering, and observability initiatives. You will build CI/CD pipelines, automate provisioning and deployment, and collaborate with engineering, security, and operations teams to improve platform readiness ...

Cloud SRE: Build Resilient, Scalable Platforms

Location
United Kingdom
improving reliability, scalability and performance of critical platforms and products. You’ll contribute to DevOps initiatives, automate processes, support incidents and strengthen monitoring and observability using tools like Dynatrace, Terraform and GitHub. Hybrid work model applies. #J-18808-Ljbffr ...

Senior Platform Engineer – Remote UK (Cloud/SRE)

Location
Greater London, England, United Kingdom
Hudl in London, United Kingdom, is seeking a Senior Engineer to join our Platform Engineering team. You’ll work on site reliability, cloud infrastructure, observability and production operations to keep Hudl’s platform highly available, scalable and secure. You’ll lead with technical excellence, mentor engineers and drive innovation using ...

Cloud Platform Operations Lead: Azure, Kubernetes & SRE

Location
Greater London, England, United Kingdom
will lead hands-on engineering and inspire a team of Platform Operations engineers. You'll shape platform strategy, manage incident response, and advance observability, automation, and governance across infrastructure, applications, and data. Remote/hybrid working options available with global teams. #J-18808-Ljbffr ...

Senior Security Engineer: AI-Driven Detection & Response

Location
Greater London, England, United Kingdom
native environment. You will design, build, and maintain detection capabilities, develop AI-powered tooling, and lead incident response while collaborating with engineering to improve observability and secure the development lifecycle. The role emphasizes data-driven decision making, adaptability to evolving threats, and strong coding fundamentals in Python, Go, or #J ...

Remote AWS SRE: Build Resilient Cloud Platforms

Location
England, United Kingdom
join a globally operating AI-driven cloud platform team. This fully remote UK role involves maintaining production systems on AWS, implementing automation and observability, and partnering with software, platform, cloud and security engineers to improve reliability. You will handle 24/7 incidents, build resilient cloud services, and drive continuous ...

Senior Java Platform Engineer - Hybrid London

Location
Greater London, England, United Kingdom
Java Software Engineer in London to join the Platform Team. You will work on a core platform of distributed services that ingest and transform observability data for visualisation, analytics and integrations. This is a permanent, full-time role with a hybrid working model from our London office. You will contribute ...

Global Pricing Platform Lead, Backend Engineer

Location
Greater London, England, United Kingdom
scale backend systems and integrations that power international markets, collaborating with architecture, product and operations. Join a brand-led engineering culture focused on reliability, observability and secure cloud-native solutions, while mentoring others and driving continuous improvement across teams. #J-18808-Ljbffr ...

DevEx Platform Engineer — Build Tooling & Reliability

Location
Greater London, England, United Kingdom
engineer in London to own end-to-end service delivery for internal developer platforms. You will work with US/UK teams on tooling, observability, and workflows to empower engineers across the firm. Role requires hands-on experience with Kubernetes, CI/CD, and infrastructure as code, plus strong communication ...

Senior Real-Time Data Platform Engineer

Location
Greater London, England, United Kingdom
messages daily, ensuring reliability at scale. You’ll design and implement scalable data processing with Flink and Kafka, using Java or .NET, while emphasizing observability, cost efficiency, and production readiness. The role features a hybrid work model in London. #J-18808-Ljbffr ...

Staff Software Engineer: Cloud Networking (Remote)

Location
Hursley, England, United Kingdom
operational excellence, working across AWS, Azure, and GCP to deliver scalable networking solutions. You will architect and drive cross-team projects, ensuring reliability, observability, and cost efficiency while collaborating with product, security, and platform teams to enable seamless integration for customers. #J-18808-Ljbffr ...