1 to 25 of 194 Observability Jobs in the East of England

Remote Senior Software Engineer (Go)

Hiring Organisation
grabjobs
Location
Mundesley, Norfolk, UK
resilience Taking full ownership of services: from initial design and implementation to deployment and production support Working with a mindset where cost-efficiency, observability, and operational excellence are core to how we build Collaborating closely with other engineers in a flat, autonomous team structure, with a strong focus on code ...

Head of Cloud

Hiring Organisation
Epos Now
Location
Norwich, Norfolk, United Kingdom
Salary
£ 70 K
improve velocity, reliability, and operational efficiency.Documentation & Continuous ImprovementChampion high standards for technical documentation and process optimization across cloud teams.Promote a culture of knowledge sharing, observability, and proactive problem-solving.Success Metrics (KPIs)Team Growth: High-performing cloud engineers recruited, onboarded, and retained.System Evolution: Delivery excellence in cross-team technical initiatives with ...

Head of Cloud

Hiring Organisation
Epos Now
Location
Norwich, Norfolk, UK
Employment Type
Full-time
reliability, and operational efficiency. Documentation & Continuous ImprovementChampion high standards for technical documentation and process optimization across cloud teams. Promote a culture of knowledge sharing, observability, and proactive problem-solving. Success Metrics (KPIs)Team Growth: High-performing cloud engineers recruited, onboarded, and retained. System Evolution: Delivery excellence in cross-team technical ...

Portfolio Software Full Stack Engineer (Contractor)

Location
Cambridge, England, United Kingdom
well-tested and performant code following engineering best practices. Design data models and work with relational and NoSQL databases where appropriate. Improve application reliability, observability and performance through monitoring, logging and optimisation. Software Engineering & Code Quality Write clean, maintainable and well-documented code following established coding standards. Perform thorough code ...

Senior Manager of SRE

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture. Drives continuous improvement in system observability, alerting, and capacity planning. Collaborates with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence. Performs platform design ...

Remote Staff Software Engineer - Data Platforms

Hiring Organisation
grabjobs
Location
Bishop's Stortford, Hertfordshire, UK
review, TDD, CI/CD and pairing using tools like Git and GitHub. Experience in operationally managing software components/service once live, including: observability best practises, logging best practises, error reporting, debugging and live incident management. Experience using tools such as Grafana, Prometheus, New Relic etc. Experience of working ...

Platform Engineer

Location
St Albans, England, United Kingdom
skills Hands-on experience of AI tools and knowledge of AI practices Desirable Experience with Kubernetes and related technologies (Linux/Docker) Knowledge of observability and monitoring tools Understanding of security principles and best practices An enthusiasm to learn different technologies and adapt to an evolving industry Benefits Competitive salary ...

Senior Software Infrastructure Engineer

Location
Cambridge, England, United Kingdom
Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu) Experience with GitHub Actions Experience with build tools (e.g. CMake) Experience with modern observability tooling (e.g. prometheus) Experience with Grafana Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance ...

Python Developer

Hiring Organisation
Addition
Location
Watford, Hertfordshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 per annum
GitHub Actions. • Significant experience with low-level Unix or Linux administration and troubleshooting. • Experience with application monitoring tools such as AppDynamics. • Administration experience with observability and infrastructure tools such as Sensu, Logz.io, Consul, HashiCorp Vault, Grafana, Graphite or InfluxDB, or equivalent technologies. • Experience with ServiceNow ITOM would be advantageous. ...

Developer - Python

Location
Watford, England, United Kingdom
same. What experience we’re looking for Experience with application monitoring tools such as AppDynamics, along with administration skills in Sensu, Logz.io: Modern Observability Powered by AI , Consul, Hashicorp Vault, Grafana, Graphite & InfluxDB, or equivalents. Some experience with ServiceNow ITOM would be a bonus. Comprehensive practical experience with ...

Head of Engineering New Hybrid (Hemel Hempstead)

Location
Hemel Hempstead, England, United Kingdom
engineering transformation initiatives. Champion modern engineering practices, DevOps, automation and AI adoption. Establish engineering standards, governance and technical best practice. Drive operational excellence through observability, reliability engineering and continuous improvement. Balance technology investment, technical debt and long-term platform evolution. Support strategic customer opportunities with technical leadership and expertise. Build ...

Senior Build & Cloud Infra Engineer

Location
Cambridge, England, United Kingdom
e.g. Terraform/OpenTofu) Desirable Experience using Kubernetes (k8s) or OpenStack Experience with GitHub Actions Experience with build tools (e.g. CMake) Experience with modern observability tooling (e.g. Prometheus) Experience with Grafana Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance ...

Senior Site Reliability Engineer

Location
Cambridge, England, United Kingdom
improvement of the Bango Platform. You hold everything expected of a Site Reliability Engineer — owning reliability end-to-end across the infrastructure, delivery pipelines, observability and incident response that keep the platform running to agreed service levels — but at greater scope, complexity and influence, and you take responsibility for lifting … than only responding to live issues. Work with Product and Integration Engineering to get recurring and systemic reliability issues onto roadmaps and permanently resolved. Observability, Monitoring & Security Set the standard for observability across teams — raising signal quality, driving down alert fatigue, and introducing the SLI/SLO maturity that others ...

Senior Storage Architect

Location
Cambridge, England, United Kingdom
registries. Strong analytical, troubleshooting, communication and documentation skills. We also value: Knowledge of GPU compute environments or AI training infrastructure. Experience with monitoring and observability tools (Prometheus, Grafana, etc.). Contributions to open-source storage, data management, or infrastructure projects. Familiarity with object storage systems (S3, RADOS Gateway, MinIO, etc. ...

Remote AWS DevOps Engineer

Hiring Organisation
grabjobs
Location
Great Yarmouth, Norfolk, UK
Automation (RPA) platforms. Experience delivering services within UK Government or highly regulated environments. Familiarity with accessibility standards, particularly WCAG 2.2. Knowledge of monitoring and observability tools such as CloudWatch, Grafana, or Prometheus. What We Offer Opportunity to work on impactful digital transformation programmes. Collaborative and supportive working environment. Exposure ...

Staff Full Stack Software Engineer

Location
Cambridge, England, United Kingdom
delivers a high-quality user experience. Design, build, and maintain our developer portal including CI/CD pipelines, documentation, automated testing, security upgrades, and observability integrations. Partner closely with platform, software and hardware teams to integrate services, tooling, and policies into the portal in a user-centric and automated manner. ...

Remote Principal Software Engineer Backend Technologies, Platform (UK Remote)

Location
Norfolk, United Kingdom
deadlines and business goals. Preferred Skills: Experience with frontend technologies such as React, Angular, or Web Components is a plus. Familiarity with monitoring and observability tools (e.g., CloudWatch, New Relic, Datadog). Knowledge of data modeling and working with both NoSQL databases. Understanding of agile methodologies, including Scrum ...

Remote Software Engineer (Cloud & Integration)

Hiring Organisation
grabjobs
Location
Epping, Essex, UK
code (IaC), automation, and CI/CD pipelines. Proficiency in writing and debugging code using TypeScript (Node.js) and Python. Experience with monitoring, logging, and observability tools such as Grafana, Prometheus, and AWS CloudWatch. Proficiency with Infrastructure as Code (IaC) tools, particularly AWS CDK and Terraform. Exposure to unit testing, integration ...

Senior AI Engineer

Hiring Organisation
PA Consulting
Location
Pimlico, Hertfordshire, UK
Employment Type
Full-time
testing agentic architectures through our own Genie Platform – enabling AI agents that reason, retrieve, and act across systems. Exploring LLMOps, evaluation tools, and model observability platforms like TruLens and LangSmith. Deploying solutions on modern cloud and DevOps environments (AWS, Azure, GCP) – but always choosing the right tools for the problem ...

Senior Software Engineer - Live & VOD Video Infrastructure

Hiring Organisation
Roku
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
GStreamer, FFmpeg, MediaMTX, or similar technologiesExperience with GPU-accelerated encoding or hardware media pipelinesFamiliarity with Kubernetes, ECS, Nomad, or other orchestration platformsExperience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or DatadogExperience building fault-tolerant ingest or transcoding platforms operating across multiple regions#LI-JC5What's Roku's approach ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
reliability, performance and continuous improvement of the Bango Platform end-to-end — from the infrastructure and pipelines that build and deploy it, to the observability and incident response that keep it running to agreed service levels. You combine three things that have historically sat in separate teams: platform and cloud … delivery pipeline ownership, and proactive/reactive reliability engineering including incident response and customer impact management. You design, build and operate the automation, observability and platform capabilities that other engineering teams rely on, and you are equally comfortable diagnosing a live incident as you are designing the Terraform module that ...

AWS Solution Architect

Location
Cambridge, England, United Kingdom
Design - defining bounded contexts, inter-service contracts, and REST/GraphQL/gRPC APIs - and set patterns for resilience (circuit breakers, retries, bulkheads) and observability (distributed tracing, structured logging, metrics). Design event streaming and async messaging using Kafka, Amazon SQS/SNS, EventBridge and Kinesis, including event schema standards ...

Developer Experience (DevEx) Engineer — Pipeline Squad

Location
Cambridge, England, United Kingdom
with the engineers you serve, and treat internal tooling as a real product.* Security-conscious by default: least privilege, secrets hygiene, supply-chain awareness.* Observability tooling (Grafana, Prometheus).Nice to have* Experience creating agentic development workflows or writing and maintaining AI skills.* npm workspaces/shared-package build systems.* Experience ...

Senior Software Engineer

Location
Luton, England, United Kingdom
design and architectural discussions, balancing speed, simplicity, maintainability and long-term sustainability when making engineering decisions. You'll continuously improve systems through better testing, observability, performance optimisation and operational practices, helping ensure our platforms remain reliable and scalable. You'll also: Write clean, maintainable and well‐tested code. Support engineering ...

Remote Full-Stack Software Engineer I

Location
Cambridgeshire, United Kingdom
mobile apps. Regularly releasing working software, using trunk-based development, automated test suites, and infrastructure-as-code principles. Incorporating requirements such as performance, resilience, observability, maintainability, security and accessibility. Collaborating with other disciplines, building effective working relationships. With your team, achieving a shared understanding of user needs, Kooth commercial ...