551 to 575 of 850 Remote/Hybrid Observability Jobs

Principal Engineer

Location
Greater London, England, United Kingdom
front of the funnel through to production release at the back, and make sure we know in advance what would break first. Improve observability and operability across the flow from "buy" to "delivered", reducing "where is my order" contacts and manual interventions. AI-assisted development: Set the standard ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
secure supply chain configurations; Establish IaC and GitOps standards with automated testing for every infrastructure change; Prototype agentic infrastructure components, including deployment and observability platforms in service meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/… Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay current with CNCF and AI ecosystem innovations, including eBPF observability and agent‐aware orchestration; Lead and mentor a team of DevOps engineers while remaining hands‐on. Требования: Experience leading or mentoring engineering teams, setting direction ...

Head of Site Reliability Engineering (SRE)

Hiring Organisation
Computershare
Location
Bristol, UK
Employment Type
Full-time
define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape. Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable … while fostering a culture of shared ownership and continuous learning. Other key responsibilities: Drive adoption of SRE principles (SLOs, error budgets, toil reduction).Establish observability and monitoring standards. Lead automation-first operations. Improve incident and problem management maturity. Partner with software and infrastructure engineering teams to embed reliability into ...

Java Developer

Hiring Organisation
Hays Technology
Location
Telford, Shropshire, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £640/day £600 - £640 p/d Inside IR35
Agile delivery. What we're looking for -Java 17+/Spring Boot 3+ -Kubernetes and containerised application delivery -Helm -GitLab CI/CD pipelines -Observability tooling (OpenSearch, Grafana, Prometheus, Dynatrace etc.) -Technical leadership experience -Agile delivery experience Role 2 - Java Engineer with Helm What you'll be doing -Working closely … standards. What we're looking for -Java and Spring Boot -Kubernetes -Helm (key requirement) -Microservices design and development -GitLab CI/CD pipelines -Containerisation -Observability and monitoring -AWS Role 3 - Senior Java Engineer (HCP/Technical Leadership) What you'll be doing -Providing technical leadership within a Java engineering team. ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
environments Define and implement infrastructure-as-code using Terraform, CDK, or equivalent Monitor platform health, define SLOs/SLAs, and build alerting and observability tooling Respond to and lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security … Actions, Jenkins, or equivalent) Solid understanding of networking, security groups, load balancing, and DNS Experience with container orchestration (Docker, Kubernetes, or ECS) Familiarity with observability tooling (Datadog, Grafana, CloudWatch, or equivalent) Understanding of HIPAA infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site ...

Advanced Engineer, Investment Technology

Location
United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

AI Consulting –Sr. AI Architect & Client Partner

Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Dunfermline, Fife, UK
Employment Type
Full-time
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Livingston, West Lothian, UK
Employment Type
Full-time
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Salford, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £80,158 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, United Kingdom
Employment Type
Full-Time
Salary
£63,824 - £83,778 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£67,547 - £83,778 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Salford, Greater Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Senior Software Engineer

Hiring Organisation
Hargreaves Lansdown
Location
Newport, UK
Employment Type
Full-time
applications using modern engineering practices. Develop integrations with internal and external systems. Enhance platform capabilities for managing and maintaining customer data. Improve platform resilience, observability, performance and scalability. Contribute across the full stack, including React-based applications where required. Take ownership of features throughout the software development lifecycle, from design … schema design. Experience migrating from REST to GraphQL architectures. Typesense, Elasticsearch or Contentful. OAuth and authentication flows. LaunchDarkly or feature management platforms. Grafana and observability tooling. Cross-account AWS architectures. TanStack Query, Zustand, Vite, Tailwind CSS. Micro-frontend architectures. AI-assisted and agentic development workflows. Financial Services experience. Interview process ...