1 to 25 of 37 Observability Jobs in the East of England

Head of Cloud

Hiring Organisation
Epos Now
Location
Norwich, Norfolk, United Kingdom
Salary
£ 70 K
improve velocity, reliability, and operational efficiency.Documentation & Continuous ImprovementChampion high standards for technical documentation and process optimization across cloud teams.Promote a culture of knowledge sharing, observability, and proactive problem-solving.Success Metrics (KPIs)Team Growth: High-performing cloud engineers recruited, onboarded, and retained.System Evolution: Delivery excellence in cross-team technical initiatives with ...

Senior Software Engineer - Java, Cloud and AI

Hiring Organisation
NEEV LIMITED
Location
Waltham Cross, Hertfordshire, South East, United Kingdom
Employment Type
Permanent
medallion architecture Familiarity with enterprise data platforms and data mesh principles, including domain-aligned datasets and well-defined data contracts Exposure to data quality, observability, and metadata management practices and tools (data validation frameworks, lineage, monitoring, and alerting) Experience enabling analytics, reporting, and AI/ML workloads through curated datasets ...

Senior Private Cloud Engineer

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
ability to debug compute, networking, storage in large production environments.“Nice To Have” Skills and Experience:Exposure to modern cloud native principles.Familiarity with observability tools (Prometheus, Grafana)!Exposure to large-scale or multi-tenant environments.Exposure to GitOps driven and CI/CD pipelines (Jenkins, ArgoCD. FluxCD, etc)Contribution to Open ...

Staff Full Stack Software Engineer

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
delivers a high-quality user experience.Design, build, and maintain our developer portal including CI/CD pipelines, documentation, automated testing, security upgrades, and observability integrations.Partner closely with platform, software and hardware teams to integrate services, tooling, and policies into the portal in a user-centric and automated manner.We invest significant ...

Senior Software Engineer – Live & VOD Video Infrastructure

Hiring Organisation
Roku
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
GStreamer, FFmpeg, MediaMTX, or similar technologiesExperience with GPU-accelerated encoding or hardware media pipelinesFamiliarity with Kubernetes, ECS, Nomad, or other orchestration platformsExperience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or DatadogExperience building fault-tolerant ingest or transcoding platforms operating across multiple regions#LI-JC5What's Roku's approach ...

AI Platform Engineer - Cambridge

Hiring Organisation
Nextech Group Limited
Location
Cambridgeshire, East Anglia, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
/conversation systems Collaborate with data, product, and engineering teams to translate business requirements into technical solutions Implement CI/CD pipelines, monitoring, and observability for AI-driven services Own the full lifecycle: experimentation, prototyping, and production deployment What We're Looking For: Strong commercial experience with Python and/ ...

R&D Software Engineer

Hiring Organisation
Aveva Group
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
with AVEVA CONNECT.Operate and improve cloud environments: support desktop streaming (Amazon WorkSpaces Applications) and Windows-centric infrastructure (EC2, FSx, Active Directory, DynamoDB) with strong observability and cost/performance focus.Deliver securely with automation and teamwork: write clean, tested, documented, deployable code; contribute to CI/CD and infrastructure automation (Azure ...

Full Stack Developer – CT Office Incubation Team

Hiring Organisation
Aveva Group
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
others.Apply baseline security best practices (secrets handling, dependency hygiene, access control awareness, safe data handling) following team standards.Troubleshoot prototype applications using logs and basic observability tools; escalate issues early with clear context.Document your work clearly (setup steps, API notes, known limitations, assumptions) to help teammates and future reuse.Collaborate with cross ...

Senior Full-Stack Software Engineer (C#/.NET, React)

Hiring Organisation
Gearset
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 60 K
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Full-Stack Software Engineer (C#/.NET, React)

Hiring Organisation
Gearset
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 50 K
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Platform Engineer - Linux / Python / AWS

Hiring Organisation
Softweb Resourcing
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 60 K
building technically complex products used in demanding development environments.What you’ll do• Maintain CI systems, artifact repositories and delivery platforms • Implement monitoring and observability tools • Build machine, container and development images • Maintain internal toolchains and infrastructure • Collaborate with engineers and stakeholdersAbout you• 2–5 years’ experience in Platform, DevOps ...

Platform Software Engineer

Hiring Organisation
X-On Health
Location
Woodbridge, Suffolk, East Anglia, United Kingdom
Employment Type
Permanent
Salary
£50,000
pipeline development and maintenance Application deployment and release management Developer tooling administration (GitLab, Packagist, NPM, Dependabot) Dependency management and automated security updates Incident Management & Observability: Assist with the diagnosis and assessment of technical issues On-call engineering and incident response Monitoring, alerting, and error tracking (Zabbix, incident.io, Sentry) Application performance ...

Infrastructure Team Lead

Hiring Organisation
Speechmatics
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
deep into any subject and a genuine love of learning on the jobBroad knowledge of the infrastructure supporting distributed systems — operating systems, observability tooling, networking, storage, containers and container orchestrationComfortable troubleshooting Linux systems (we use Ubuntu)Solid experience with infrastructure-as-code and automation tooling — we use Terraform and Ansible ...

Service Design Specialist

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
needed to support the service, including incident, major incident, request, problem, change, release, knowledge and continual improvement activities.Embed availability, capacity, resilience, continuity, performance, security, observability, user experience and supportability requirements into service designs.Help define service levels and service health measures, including SLAs, SLOs, SLIs, operational metrics, monitoring, alerting, dashboards ...

DevOps Engineer - (Platform Engineering & Security)

Hiring Organisation
Sanderson Recruitment
Location
Cambridgeshire, East Anglia, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £520 per day
management solutions. Integrate security scanning, compliance validation, and vulnerability management into engineering workflows. Develop automation solutions that improve efficiency and reduce manual intervention. Support observability, monitoring, logging, and platform reliability initiatives. Collaborate with development, infrastructure, database, and security teams to deliver scalable platform solutions. Contribute to platform standardisation, engineering best ...

Senior MLOps Engineer

Hiring Organisation
Danaher Corporation
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
grade AI systems with focus on reliability, scalability, and performance.Cloud and infrastructure platforms and tools, including deployment, containerization, and CI/CD practices.Monitoring and observability tools, multi-agent orchestration tools, monorepo technologiesStrong inter-personal skills and ability to collaborate with cross-functional teams and translate business problems into technical solutions.Abcam ...

Senior CPU Performance Engineer

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
ResponsibilitiesDevice Performance AnalysisAnalyse performance on real devices using real-world workloadsSupport investigation of system-level performance issues, identifying CPU-related factorsOperate effectively in low-observability environments (e.g. no waveforms, partial counters, noisy systems)CPU-Focused AnalysisUse PMU counters and profiling tools to understand CPU behaviourContribute to identifying performance bottlenecks (e.g. ...

Infrastructure Monitoring Engineer

Hiring Organisation
BT Group
Location
Ipswich, Suffolk, United Kingdom
Salary
£ 70 K
like to see on your CVMandatoryFamiliarity with using Linux operating systemsExperience with one or more of the following areas Administering or deploying monitoring and observability tools such as CheckMK, Prometheus or ZabbixAdministering or deploying SIEM tools such as Elastic Security, Splunk or Trellix ESMWorking with dashboarding and visualisation tools such ...

Senior Data and AI Engineer

Hiring Organisation
Renewable Energy Systems
Location
Kings Langley, Hertfordshire, United Kingdom
Salary
£ 70 K
/ML engineering as a core workstream, designing feature-ready datasets, model pipelines, and AI-ready data products, and applying engineering rigour (testing, versioning, observability) to ML pipelines.Engineer data products and pipelines that support LLM and generative AI use cases, including retrieval-ready data structures and pipelines feeding AI applications.Drive … validating their output.MLOps — CI/CD for data and ML pipelines, infrastructure as code, containerisation, and orchestration tools such as Airflow or equivalent.Data quality & observability — hands-on experience with testing frameworks, monitoring, and quality controls in production environments.LLMs and generative AI — practical understanding of how to engineer data products ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, United Kingdom
Employment Type
Permanent
Salary
GBP 50,000 Annual
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, click apply for full ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, UK
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity xkybehq to make a real impact across a large, Please click ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Cloud Platform Tech Lead

Hiring Organisation
Danaher Corporation
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 60 K
Bring more to life.Are you ready to accelerate your potential and make a real difference within life sciences, diagnostics and biotechnology At Abcam, one of Danaher’s 15+ operating companies, our work saves lives—and ...

Site Reliability Engineering Professional

Hiring Organisation
BT Group
Location
Ipswich, Suffolk, United Kingdom
Salary
£ 60 K
heart of BT International's next-generation platforms, collaborating with engineering, product, supplier, and operational teams to improve service reliability through automation, observability, and SRE best practices.What you’ll be doingSupport the 24x7 operation of BT International's core network and platform services, ensuring high availability and performance.Proactively monitor services … improve platform reliability, resilience, and operational readiness.Build and maintain high-quality operational documentation, including runbooks, service maps, playbooks, and handover processes.Develop and enhance monitoring, observability, and operational tooling capabilities.Support the implementation of automation and CI/CD practices to improve operational efficiency.Coach and support your colleagues and customer-facing operational ...

Software Engineer, Agentic AI

Hiring Organisation
Roku
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development—from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement.You will contribute directly to Roku’s AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … with the systems around them.Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases.Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance.Build safeguards that improve production readiness and reliability, including testing pipelines, controlled ...