1,376 to 1,400 of 2,124 Remote/Hybrid Observability Jobs

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to take ownership of the cloud infrastructure and DevOps practices powering our BIM Platform, working across Azure, Kubernetes, CI/CD, observability, security and developer tooling. The Role This is a hands‐on senior engineering role with broad ownership across our cloud and platform infrastructure. … Code using Terraform, ARM templates and Helm Design and improve CI/CD pipelines using GitHub Actions, enabling fast, safe and repeatable deployments Build observability across our infrastructure and services through monitoring, logging, alerting and distributed tracing Improve platform reliability, scalability and performance through capacity planning, autoscaling and resource optimisation ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Runcorn, Cheshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
North End, Lincolnshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer – Associate - London

Location
City Of London, England, United Kingdom
build, run and continuously improve this business-critical service. The role combines software engineering, systems engineering and production expertise to improve the reliability, scalability, observability, incident response and operational efficiency of the CTL platform. Required Skills/Experience Strong programming ability in one or more modern languages – Java … building maintainable automation beyond simple scripts. Good understanding of networking, messaging, distributed systems, data structures, algorithms and software design fundamentals. Hands-on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes ...

Senior Tech Consultant

Location
Leeds, England, United Kingdom
direction, identifying opportunities to improve a client’s product or technology beyond the immediate brief. Implement and improve CI/CD, cloud infrastructure and observability, primarily using GitHub Actions, AWS and Azure. Promote strong engineering practices around code quality, automated testing, peer review and security, while mentoring and supporting other … building production‐grade AI‐native products, integrating frontier models and designing reliable agentic systems, with a strong understanding of evaluation, context engineering, tool use, observability, security, latency, cost and non‐deterministic behaviour. Strong problem‐solving skills, with the ability to take complex or ambiguous problems and turn them into well ...

DevOps Engineer - CloudOps

Location
Aberdeen City, Scotland, United Kingdom
production. Improve developer tooling, automation and internal platforms, with clear documentation and maintainable code that helps teams work efficiently. Strengthen reliability through monitoring and observability, helping teams detect and resolve issues early. Apply security and compliance practices throughout infrastructure and delivery pipelines. Create clear documentation and maintainable, production-ready code. … Terraform modules. As a DevOps Engineer, you’ll help develop and operate the platform while building your experience in cloud infrastructure, networking, security, observability, and DevSecOps practices. Who we think you are You are humble, open, collaborative, and constructive in how you work withothers. You have experience in DevOps, cloud ...

Remote Senior Platform/DevOps Engineers

Location
Leicester, Leicestershire, United Kingdom
productivity metrics. Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. Mentor junior and mid-level engineers, lead incident response for team-owned services … rigorous testing strategies. High proficiency in Java and Python (Golang is a plus). Experience building mature CI/CD pipelines and working with observability tools (e.g., Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. Experience with financial systems, real ...

Remote Senior Platform/DevOps Engineers

Location
Chelmsford, Essex, United Kingdom
productivity metrics. Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. Mentor junior and mid-level engineers, lead incident response for team-owned services … rigorous testing strategies. High proficiency in Java and Python (Golang is a plus). Experience building mature CI/CD pipelines and working with observability tools (e.g., Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. Experience with financial systems, real ...

Senior Platform Engineer for DevOps - Remote (m/f/d)

Location
Grantham, Lincolnshire, United Kingdom
productivity metrics. Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. Mentor junior and mid-level engineers, lead incident response for team-owned services … rigorous testing strategies. High proficiency in Java and Python (Golang is a plus). Experience building mature CI/CD pipelines and working with observability tools (e.g., Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. Experience with financial systems, real ...

Senior Platform Engineering Manager at Prolific

Location
United Kingdom
operational excellence, and innovation. Champion SRE Culture: Own availability and embed SRE principles across the organization, including defining SLOs, SLAs, error budgets, and enhancing observability and incident remediation. Platform & Developer Experience: Own the developer-facing platform — golden paths, self-service infrastructure, and internal tooling — so teams can provision, deploy … Infrastructure-as-Code (Terraform/Terragrunt, Crossplane), with GitOps workflows using tools like ArgoCD. Reliability & Architecture: Solid understanding of complex infrastructure and application architecture, observability principles, and incident management. Stability & Velocity: Experience balancing the need for platform stability and reliability with the goal of increasing developer productivity and velocity. Bridge ...

Lead GCP Engineer

Location
City Of London, England, United Kingdom
practices across the platform. Drive adoption and maturity of CI/CD pipelines, Infrastructure as Code and automated deployment processes. Implement and improve platform observability, monitoring, alerting and operational support capabilities. Develop and maintain reusable automation patterns and engineering standards. Leverage Terraform and Infrastructure as Code principles to ensure repeatable … Infrastructure as Code expertise. Strong experience with CI/CD tooling and deployment automation. Experience building and operating resilient cloud platforms. Strong understanding of observability, monitoring, incident management and operational excellence. Data Engineering & Integration Experience designing secure data integration solutions. Knowledge of APIs, databases, data pipelines and cloud‐based data ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Remote, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Norwich, Norfolk, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Lossiemouth, Moray, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Abingdon, Oxfordshire, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Preston, Lancashire, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Doncaster, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Hengoed, Glamorgan, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Rochester, Kent, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Lead DevOps Engineer

Location
Manchester, England, United Kingdom
lifecycle. Automation and reliability engineering Proficient in Python with solid scripting in PowerShell and Bash to drive automation and operational efficiency, combined with strong observability practice (for example Azure Monitor, Sentinel, Application Insights and Log Analytics) using SLIs, SLOs and distributed tracing to sustain and improve system reliability. Autonomy … Code and CI/CD. This is a hands‐on leadership role where you'll influence key technical decisions, champion DevSecOps, resilience and observability, and work closely with engineering teams to create high‐quality services that make a real impact. You'll use technologies including Azure, Azure DevOps and Terraform ...

Technology Lead (Remote - UK)

Hiring Organisation
Reonomy
Location
London, United Kingdom
Salary
£ 70 K
reliability, scalability, and long-term sustainabilityManage and reduce technical debt strategicallyDesign cost-efficient, cloud-native solutionsEngineering ExcellenceSet and uphold standards for code quality, testing, observability, security, and documentationChampion automated testing and “shift-left” quality practicesPromote DevSecOps principles to embed security and compliance into developmentLead through thorough code reviews, design documentation … Deep understanding of system design, distributed systems, scalability, APIs, and data modelingExperience with Infrastructure as Code (Terraform or similar)Strong CI/CD and observability experienceExperience implementing robust testing strategiesComfortable working cross-functionally in Agile environmentsStrong communication skills and ability to operate effectively in ambiguous situationsUnlock your Altus Experience ...

Integration Architect

Location
United Kingdom
Terraform or Bicep . Design integration patterns using REST APIs, microservices, Azure Service Bus, Event Grid and event-driven architectures . Establish monitoring and observability using Azure Monitor, Log Analytics and Application Insights . Present solution designs through architecture governance and review processes. Provide technical leadership and guidance to engineering … implementing APIOps/GitOps approaches for API lifecycle management. Knowledge of Azure API Center or similar API cataloguing/governance platforms. Experience with Azure observability tooling including Application Insights, Azure Monitor and Log Analytics . Knowledge of emerging AI integration patterns such as Model Context Protocol (MCP) or AI Gateway ...

Platform Engineer III - (Pipelines & Developer Experience)

Location
Leeds, England, United Kingdom
least privilege, policy‐as‐code, scanning). Improve developer experience: faster feedback loops, quality gates, ephemeral/preview environments, and great documentation. Instrument pipeline observability (Datadog or equivalent) and define SLOs (queue time, lead time, change fail rate, MTTR) to drive reliability. Automate IaC workflows (Terraform/Terragrunt) and integrate … implement and maintain CI/CD release pipelines. Experience with scripting and programming (.NET preferred; familiarity with Go, Python, PowerShell beneficial). Knowledge of observability tooling, chaos testing, and incident management. Strong analytical and problem‐solving abilities, with the capability to closely collaborate with engineering teams. Highly outcome‐oriented, pragmatic ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity ...