2,651 to 2,675 of 4,362 Permanent Observability Jobs

Remote DevOps Team Lead

Hiring Organisation
Runware
Location
Merthyr Tydfil, Wales, UK
operation of Runware’s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Remote DevOps Team Lead

Hiring Organisation
Runware
Location
Leamington Spa, Warwickshire, UK
operation of Runware’s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Remote DevOps Team Lead

Location
Egham, Surrey, United Kingdom
operation of Runware s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Remote DevOps Team Lead

Location
Immingham, Lincolnshire, United Kingdom
operation of Runware s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Remote DevOps Team Lead

Location
Stowmarket, Suffolk, United Kingdom
operation of Runware s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Azure Cloud Engineer

Location
Leeds, England, United Kingdom
Identities Azure Key Vault RBAC Secrets and certificate management Experience with Azure SQL connectivity and platform integration . Experience with Azure-native monitoring and observability tools . Comfortable designing and implementing solutions for: Resilience Failover High availability Recovery Secure production operations Ability to produce and contribute to detailed Low-Level ...

Global Cloud Platform Technical Design Authority

Location
Egham, England, United Kingdom
recovery activities, and significant infrastructure initiatives. Reviews technical solutions and validates alignment with functional and non-functional requirements including availability, resilience, disaster recovery, security, observability, supportability, and cost optimisation. Escalates material technical risks, unresolved design conflicts, and policy exceptions through appropriate governance channels where required. Technical Governance & Change Approval Operates … supporting governance, compliance, and regulatory requirements. Standards Compliance & Platform Assurance Enforces adherence to cloud platform standards covering networking, identity, security, segmentation, naming conventions, tagging, observability, backup, disaster recovery, and operational controls. Assures alignment with approved AWS and Azure landing zones, reference architectures, engineering standards, security frameworks, and Well-Architected principles. ...

Remote DevOps Team Lead

Location
Kingston upon Hull, East Yorkshire, United Kingdom
operation of Runware s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Network Automation Engineer

Location
United Kingdom
automated network assurance platforms, centred on Nautobot and IP Fabric. The role provides the data and automation foundation that underpins our NOC, Managed Services, observability and AI/ML (MARTINA) capabilities – turning trusted network data into automated outcomes that reduce manual effort and improve service quality. You will play … Managed Services teams by automating L1 triage, alert enrichment, inventory reporting and customer-facing analytics, and by feeding high-quality data into observability and AI/ML (MARTINA) capabilities. Skills, Qualifications and Experience Minimum 5 years' experience in network engineering roles, with demonstrable experience of network automation (NetDevOps) in production ...

Network Engineer

Hiring Organisation
Cboe Exchange
Location
London, UK
Employment Type
Full-time
leadership, translating complex issues into actionable outcomes. The role also emphasizes building efficient, modern operations through scripting, automation, and API-driven workflows to improve observability, accelerate root cause analysis, and reduce mean time to repair (MTTR). While familiarity with AI-assisted tooling is beneficial, the primary focus … cloud, and infrastructure domains, collaborating with engineering teams and vendors to restore service and resolve complex issues Develop and maintain automation, monitoring enhancements, and observability tooling using scripting and API-driven integrations to improve efficiency, visibility, and troubleshooting speed Perform deep packet analysis and contribute to root cause investigations, translating ...

Senior Sales Engineer - Public Sector (UK)

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscapeWho You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customerSomeone with strong written and oral communication … above may vary based on the country of your employment and the nature of your employment with Datadog. About Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Senior Full Stack Engineer

Hiring Organisation
Dignity Funerals Limited
Location
Maidenhead, Berkshire, South East, United Kingdom
Employment Type
Permanent
Salary
£90,000
challenges. Using discovery, experimentation and proof-of-concepts to shape technical direction. Improving engineering efficiency through best practice, standardisation and automation. Leading standards across observability, performance, testing and system quality. Communicating risks, options and technical trade-offs clearly to both technical and non-technical stakeholders. Mentoring junior and mid-level … TypeScript, Node.js, and React or Angular , together with strong knowledge of systems architecture, API design and databases. You will also understand the importance of observability, monitoring, testing, performance, security, accessibility and modern DevOps practices. You are comfortable working with ambiguous requirements, managing technical debt and making thoughtful decisions that balance ...

Remote Staff Software Engineer Identity and Access, Identity Squads UK Remote

Location
Kent, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … experience. Experience working with OpenFGA and KiFederate IDP. Experience working in security critical environments. Experience contributing to or maintaining Open Source projects. Familiarity with observability tooling (e.g., Grafana, Prometheus, OpenTelemetry). Compensation & Rewards: In the UK, the Base compensation range for this role is £100,000 - £124,000. Actual compensation ...

Remote Staff Software Engineer Identity and Access, Identity Squads UK Remote

Location
Calstock, Cornwall, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … experience. Experience working with OpenFGA and KiFederate IDP. Experience working in security critical environments. Experience contributing to or maintaining Open Source projects. Familiarity with observability tooling (e.g., Grafana, Prometheus, OpenTelemetry). Compensation & Rewards: In the UK, the Base compensation range for this role is £100,000 - £124,000. Actual compensation ...

Remote Staff Software Engineer Identity and Access, Identity Squads UK Remote

Location
Forfar, Angus, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … experience. Experience working with OpenFGA and KiFederate IDP. Experience working in security critical environments. Experience contributing to or maintaining Open Source projects. Familiarity with observability tooling (e.g., Grafana, Prometheus, OpenTelemetry). Compensation & Rewards: In the UK, the Base compensation range for this role is £100,000 - £124,000. Actual compensation ...

Remote Staff Software Engineer Identity and Access, Identity Squads UK Remote

Location
Bristol, Gloucestershire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … experience. Experience working with OpenFGA and KiFederate IDP. Experience working in security critical environments. Experience contributing to or maintaining Open Source projects. Familiarity with observability tooling (e.g., Grafana, Prometheus, OpenTelemetry). Compensation & Rewards: In the UK, the Base compensation range for this role is £100,000 - £124,000. Actual compensation ...

Remote Staff Software Engineer Identity and Access, Identity Squads UK Remote

Location
Kirkby, West Yorkshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … experience. Experience working with OpenFGA and KiFederate IDP. Experience working in security critical environments. Experience contributing to or maintaining Open Source projects. Familiarity with observability tooling (e.g., Grafana, Prometheus, OpenTelemetry). Compensation & Rewards: In the UK, the Base compensation range for this role is £100,000 - £124,000. Actual compensation ...

Lead AI Engineer

Hiring Organisation
Capco
Location
London, UK
Employment Type
Full-time
experience deploying LLMs and multi-modal models at scaleStrong engineering background in Python with proven backend and API development skillsSolid understanding of scalable MLOps, observability, and cloud-native AI deploymentExcellent communication, problem-solving, and project management skills in agile environmentsBonus Points ForExperience with agentic frameworks (e.g., LangChain, LlamaIndex)Experience … deep learning frameworks and front-end developmentFamiliarity with Langfuse, Langsmith, or other LLM observability toolsUnderstanding of Model Context Protocol and bias/hallucination mitigation techniquesPrevious success in integrating GenAI solutions into enterprise-scale systemsWhy Join CapcoDeliver high-impact technology solutions for Tier 1 financial institutionsWork in a collaborative, flat ...

Software Development Engineer II - BDP

Location
Welwyn Garden City, England, United Kingdom
backend event-driven platform using Java and Spring Boot• Pairing with more senior engineers to design, implement, test, and ship code• Learning to use observability tools like New Relic and Splunk to monitor live systems• Participating in planning sessions and team discussions to understand requirements and contribute ideas• Writing automated … feedback, and share what you're learning Nice to have:• Exposure to Spring Boot, NoSQL databases, or cloud services• Curiosity about performance, scalability, and observability in large-scale systems• Familiarity with Git, CI/CD pipelines, and containerisation• Some experience working with platforms like Kafka Whats ...

Site Reliability Engineer

Location
Leeds, England, United Kingdom
exceptional customer experiences through reliable, scalable and high-performing technology. Operating at the heart of our betting and gaming platforms, our SRE team combines observability, automation and engineering excellence to ensure our systems perform when it matters most. Working across operational, proactive and engineering initiatives, you'll collaborate with development … coordinate responses and restore services efficiently. Conduct detailed postmortem investigations, identifying root causes, challenging assumptions and driving meaningful improvements to prevent recurrence. Utilise observability tools such as Splunk, New Relic and CloudWatch to monitor platform health, investigate issues and uncover performance insights. Design, execute and analyse performance testing activities ...

Global Head of Applied AI & Decision Systems - Commodities & Global Markets

Location
Greater London, England, United Kingdom
will partner with business leaders to identify and prioritise high-value opportunities; work closely with Engineering and control partners to establish appropriate guardrails, observability, and risk management controls; and lead the design, deployment, and value realisation of AI use cases at enterprise scale. What you offer Significant experience in senior … modern software delivery practices. Hands-on experience designing and deploying AI and machine learning solutions in production, including large language model applications, evaluation frameworks, observability, and operational support models. Demonstrated ability to quantify business value, prioritise a portfolio of initiatives and track outcome realisation beyond the proof-of-concept stage. ...

Global Head of Applied AI & Decision Systems - Commodities & Global Markets

Location
United Kingdom
will partner with business leaders to identify and prioritise high-value opportunities; work closely with Engineering and control partners to establish appropriate guardrails, observability, and risk management controls; and lead the design, deployment, and value realisation of AI use cases at enterprise scale. What you offer Significant experience in senior … modern software delivery practices. Hands-on experience designing and deploying AI and machine learning solutions in production, including large language model applications, evaluation frameworks, observability, and operational support models. Demonstrated ability to quantify business value, prioritise a portfolio of initiatives and track outcome realisation beyond the proof-of-concept stage. ...

Cloud/Devops Engineer

Location
Leeds, England, United Kingdom
design Strong understanding of: Managed Identities Azure Key Vault RBAC secrets/certificate management Experience with Azure SQL connectivity and platform integration Monitoring/observability experience using Azure-native tooling Comfortable designing for: resilience failover high availability recovery secure production operation Able to produce and contribute to a detailed ...

Senior Fullstack Engineer, £500pd (Outside IR35)

Location
United Kingdom
TypeScript (Vue or React) REST APIs and Kafka event-driven systems AWS (including S3, Lambda, ECS, EKS, MSK) TDD, CI/CD and IaC Observability AI assisted engineering In return they’re offering a remote-first 6-month contract outside IR35 that’s paying up to £500pd. To be suitable ...

Team Leader - Data Engineering & Integration - Commodities Data

Hiring Organisation
Bloomberg
Location
London, UK
Employment Type
Full-time
direction, balancing modernization, production stability, partner needs, and business-as-usual delivery. Build team capability in data pipelines, integration patterns, Python, SQL, orchestration, automation, observability, controls, and production support. Partner with Product, Engineering, Data Modelling, Data Quality, Content Acquisition, Enablement, and regional teams to deliver business-aligned outcomes. Contribute … move teams toward scalable, automated, supportable, and well-controlled operating models. Strong working knowledge of Python, SQL, orchestration tools, workflow platforms, automation frameworks, observability, and production support practices. Experience embedding data quality controls, reconciliation, completeness checks, timeliness checks, and exception workflows into production processes. Ability to work closely with data ...