151 to 175 of 261 Observability Jobs in the East of England

Remote Senior Full-Stack Engineer

Location
Frinton-On-Sea, Essex, United Kingdom
easily working on a fully distributed team. Bonus points! Background and interest in Machine Learning and/or mathematics. Familiarity with monitoring and observability tools like DataDog. Experience or interest in implementing performant Web Applications and responsive design. \n £63,600 - £98,325 a year We base our salary ranges ...

Artificial Intelligence Engineer

Location
Cambridge, England, United Kingdom
solutions in enterprise or regulated environments — aviation, land transport, public safety, telecommunications, or government Familiarity with cloud platforms, APIs, CI/CD, analytics, and observability Experience shaping early-stage products with users and iterating based on feedback Additional information Why Join NCS? Grow with Us Work on cutting-edge ...

Remote Staff Software Engineer - Databases SRE | UK | Remote

Hiring Organisation
Grafana Labs
Location
Norfolk, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Tilbury, Essex, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
St Neots, Cambridgeshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Head of DevOps

Location
Sandy, Bedfordshire, United Kingdom
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Senior Security Platform Architect (SCA & Backend)

Location
Cambridge, England, United Kingdom
design and implement backend services, Python APIs, and workflow components to enable tool onboarding, analysis execution, results processing, and delivery, while improving scalability and observability across the platform. #J-18808-Ljbffr ...

Senior ML Infra Engineer for Scalable Conversational AI

Location
Cambridge, England, United Kingdom
product, data and platform partners. This hands-on role spans software architecture, ML lifecycle decisions and production operations to improve training, evaluation, deployment, and observability of models at scale. #J-18808-Ljbffr ...

AWS Solutions Architect - Microservices & Event-Driven

Location
Cambridge, England, United Kingdom
translate complex requirements into production-grade solutions. You will define end-to-end architectures, lead domain-driven design, and set patterns for resilience, observability, and security. #J-18808-Ljbffr ...

Remote SRE: Cloud Reliability Engineer (AI Tools)

Location
Hemel Hempstead, England, United Kingdom
holiday operator, is seeking an experienced Site Reliability Engineer to join our Product Technology team. This role focuses on cloud reliability, CI/CD, observability and database resilience across diverse engines, with occasional travel to Hemel Hempstead and off-site events. You will work with engineering teams to design, implement ...

Remote SRE: Cloud Reliability & AI-Driven Ops

Location
Hemel Hempstead, England, United Kingdom
Haven is seeking a hands-on Site Reliability Engineer to join our Product Technology team. This remote-first role involves shaping CI/CD, observability, and incident response while collaborating with engineers and tech leads to ensure reliable, scalable platforms for guests and colleagues. You’ll tackle infrastructure design, tooling ...

Backend SDE II — Real-Time Data & Event-Driven Systems

Location
Welwyn Garden City, England, United Kingdom
delivering personalised experiences at scale while collaborating with more senior engineers. You will gain hands-on experience with Kafka, Kubernetes, CI/CD, and observability tools, contributing to distributed systems and production readiness in a fast-paced environment. #J-18808-Ljbffr ...

Operations Team Lead (Production & Reliability)

Location
Norwich, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Operations Team Lead (Production & Reliability)

Location
Watford, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Remote Software Engineer

Location
Ipswich, Suffolk, United Kingdom
RFCs, build proof-of-concepts, and run experiments to validate approaches. Once we commit to building something, we care about maintainability more than cleverness, observability more than hoping it works, and scalability more than premature optimisation. We ship quality code over hitting arbitrary deadlines. Once something goes live, we refactor ...

Site Reliability Engineer (SRE)

Location
Cambridge, England, United Kingdom
availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications. Join Altium as a Senior Site Reliability Engineer … ensure the reliability and performance of the Altium Cloud Platforms. Key Responsibilities: Understanding how an Altium Cloud Platform works Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection. Develop and implement reliability frameworks and patterns that standardize and elevate ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
focused on their Azure platform, another focused on their AWS platform. Both roles require hands on experience with IaC, Containerisation and Monitoring/Observability tools. Experience required for the Senior Site Reliability Engineer: Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within … experience with both Azure, and on-premise virtual machines. Experience with Infrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Sawbridgeworth, Hertfordshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand ...

Remote Senior Platform/DevOps Engineers

Location
Cambridge, Cambridgeshire, United Kingdom
productivity metrics. Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. Mentor junior and mid-level engineers, lead incident response for team-owned services … rigorous testing strategies. High proficiency in Java and Python (Golang is a plus). Experience building mature CI/CD pipelines and working with observability tools (e.g., Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. Experience with financial systems, real ...

Senior Platform Engineer for DevOps - Remote (m/f/d)

Location
Dunstable, Bedfordshire, United Kingdom
productivity metrics. Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. Mentor junior and mid-level engineers, lead incident response for team-owned services … rigorous testing strategies. High proficiency in Java and Python (Golang is a plus). Experience building mature CI/CD pipelines and working with observability tools (e.g., Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. Experience with financial systems, real ...

Senior Platform Engineer for DevOps - Remote (m/f/d)

Location
Baldock, Hertfordshire, United Kingdom
productivity metrics. Build and optimize CI/CD pipelines and developer tooling to eliminate friction in the software delivery lifecycle. Contribute to the system observability roadmap by implementing monitoring, tracing, and alerting to ensure high operational reliability. Mentor junior and mid-level engineers, lead incident response for team-owned services … rigorous testing strategies. High proficiency in Java and Python (Golang is a plus). Experience building mature CI/CD pipelines and working with observability tools (e.g., Proven ability to take ownership of complex projects in a fast-paced environment, balancing speed with architectural health. Experience with financial systems, real ...

Senior Platform Engineer

Location
Welwyn Garden City, England, United Kingdom
GitOps workflows to enable safe, fast, and repeatable delivery. Championing DevSecOps principles, embedding security and compliance into the software delivery lifecycle. Establishing and improving observability, monitoring, and incident response practices, including vulnerability management and remediation. Mentoring engineers and contributing to a strong engineering culture through knowledge sharing, documentation, and technical … would be great if you have the following Experience with Helm, Kustomize, and Kubernetes ecosystem tooling. Familiarity with Azure and Azure DevOps. Experience with observability platforms and Kubernetes policy enforcement tools. Proficiency in scripting or programming (e.g. Bash, Python, PowerShell, C#). Experience designing multi‐region or highly available systems. ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Central bedfordshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US Our client ...