601 to 625 of 910 Remote Observability Jobs

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £83,778 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£67,547 - £83,778 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Salford, Greater Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Senior Software Engineer

Hiring Organisation
Hargreaves Lansdown
Location
Newport, UK
Employment Type
Full-time
applications using modern engineering practices. Develop integrations with internal and external systems. Enhance platform capabilities for managing and maintaining customer data. Improve platform resilience, observability, performance and scalability. Contribute across the full stack, including React-based applications where required. Take ownership of features throughout the software development lifecycle, from design … schema design. Experience migrating from REST to GraphQL architectures. Typesense, Elasticsearch or Contentful. OAuth and authentication flows. LaunchDarkly or feature management platforms. Grafana and observability tooling. Cross-account AWS architectures. TanStack Query, Zustand, Vite, Tailwind CSS. Micro-frontend architectures. AI-assisted and agentic development workflows. Financial Services experience. Interview process ...

Platform Chapter Lead - Engineering

Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Senior Software Engineer

Hiring Organisation
Hargreaves Lansdown
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Part Time
applications using modern engineering practices. Develop integrations with internal and external systems. Enhance platform capabilities for managing and maintaining customer data. Improve platform resilience, observability, performance and scalability. Contribute across the full stack, including React-based applications where required. Take ownership of features throughout the software development lifecycle, from design … schema design. Experience migrating from REST to GraphQL architectures. Typesense, Elasticsearch or Contentful. OAuth and authentication flows. LaunchDarkly or feature management platforms. Grafana and observability tooling. Cross-account AWS architectures. TanStack Query, Zustand, Vite, Tailwind CSS. Micro-frontend architectures. AI-assisted and agentic development workflows. Financial Services experience. Interview process ...

Software Engineering Manager

Location
Corby, England, United Kingdom
Collaboration: Collaborate effectively with agile teams, domains, architects, the engineering practice, scrum masters, and product owners to drive integrated solutions. Operational Excellence: Coordinate the observability and operational support of production infrastructure, ensuring system reliability and performance. Quality Assurance: Review engineering output for quality, completeness, and adherence to best practices. Production …/CD pipelines using tools such as Terraform and GitLab CI. Knowledge of cloud platform tooling including HashiCorp Vault, Nomad, Kong API Gateway, and observability platforms such as Datadog. Solid understanding of application security, performance optimisation, caching technologies (e.g. Redis), and test automation. Demonstrated knowledge of solution architecture, emerging technologies ...

Lead Java Developer

Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc. Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

DevOps Engineer

Location
Reigate, England, United Kingdom
infrastructure. Automate environment provisioning across development and production. Manage backend state, pipelines, and state-change detection integrations. Platform Engineering & SRE Own and improve reliability, observability, and performance of the platform. Implement SLOs, alerting, dashboards, and auto remediation where possible. Troubleshoot cluster level, networking, and workload deployment issues. Lead root cause … endpoints, Certificate/Secret management etc Strong debugging and operational experience (SRE mindset). Solid experience of DevSecOps architecture, processes & tooling Solid understanding of Observability Process & Tooling Logging, metrics, traces, dashboards Other highly desirable, but not essential skills are: Experience with: GitOps - ArgoCD or GitOps workflows Zero downtime deployments (blue ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services.You will work with technology, engineering, operations … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices including ...

AI Platform & Site Reliability Engineering Managing Consultant

Location
United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services. You will work with technology, engineering, operations … operational requirements. AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. Reliability Engineering & SRE: Establish SRE practices including ...

ML Data & Platform Engineer — Hybrid ML Ops & Pipelines

Location
Cambridge, England, United Kingdom
infrastructure to production ML—owning problems end-to-end to accelerate model delivery. You’ll collaborate with the ML team to improve data quality, observability, and MLOps practices, while scaling infrastructure for faster iteration and reliability. #J-18808-Ljbffr ...

Lead AI Software Engineer – Hybrid Delivery Leader

Location
Newbury, England, United Kingdom
secure, production-ready code while guiding engineers and AI agents. You''ll combine deep software engineering with AI-native practices, ensuring code quality, testing, observability, and compliance with standards. The role is hybrid, located in Newbury or Paddington, and offers the opportunity to shape #J-18808-Ljbffr ...

Backend Engineer, EMEA — Remote & AI-Driven Ownership

Location
United Kingdom
work. You will implement and ship backend features using Ruby/Rails, Go, or other modern languages, design APIs, and ensure maintainability, performance, and observability across services. #J-18808-Ljbffr ...

Remote Applied AI Engineer: Agentic Systems & Automation

Location
Greater London, England, United Kingdom
Applied AI Engineer to build agentic infrastructure and production-grade AI systems powering automation across its real estate platform. You will focus on reliability, observability, and scalable workflows to push Dwelly's AI capabilities into core business processes. As an early core member, you will influence architectural decisions, developer tooling ...

Product Manager - Health & Home AI Infrastructure

Location
United Kingdom
complex problems into actionable roadmaps that connect technology with user value. You will align offline data infrastructure and experimentation platforms to improve data access, observability, and analytics capabilities for products used by millions. The role emphasizes collaboration with engineering, data science, and marketing teams to define metrics, drive experiments ...

Remote Backend Engineer - Multi-Tenant App Platform

Location
United Kingdom
Grafana Labs, the leader in open observability, is hiring for a cloud/Grafana-as-a-service role. You will develop backend features in Golang, improve system reliability, and influence the product roadmap while collaborating across teams. The role supports a fully remote setup with compensation guided by UK market ...