1,651 to 1,675 of 1,866 Observability Jobs

Automation Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Title: Automation and Observability Engineer Location: Glasgow (hybrid) Contract: 12-months (possible extensions) Pay rate: up to £500 p/d PAYE Are you an experienced Automation and Observability Engineer looking to make an impact in a global financial services environment? We are seeking a talented engineer to join … Enterprise Technology Services team, delivering high-quality automation and observability solutions to enhance data protection and system reliability. What you'll do: Work closely with internal teams to automate processes and improve alerting/observability solutions. Focus on enhancing the reliability of the data protection environment through automation or improved ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Wales, United Kingdom
working on Mondays and Fridays Take technical ownership of the platforms that support Centerprise services and customer operations. You will lead improvements in reliability, observability and automation while remaining closely involved in complex engineering, major incidents and service recovery. Role Summary As Principal Platform Engineer, you will … hours escalation when required. Identify and address technical debt, operational risk and platform weaknesses. Ensure services remain supportable, recoverable and operationally efficient. Observability and service health Own monitoring and observability tooling, standards and operational dashboards. Develop service health metrics that provide clear and useful operational insight. Improve the quality ...

Senior Software Engineer – Order Management Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technical direction of the Order Management platform, driving improvements in architecture, integration patterns, automation, reliability, and security. Drive operational excellence by designing systems with observability, monitoring, resilience, and supportability at their core. Use tools such as Dynatrace to improve platform visibility, monitoring, and alerting capabilities. Collaborate closely with Product Managers … fulfilment operations. Strong Java engineer with experience designing and developing enterprise services, APIs, and integrations across distributed systems. Experienced in CI/CD practices, observability, production support, and operational excellence. Comfortable designing systems with resilience, recoverability, monitoring, and supportability in mind. A collaborative engineer who enjoys working across disciplines, raising ...

Staff Backend Engineer - Grafana Second Horizon | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … remote opportunity, and we would be interested in applicants located in Spain, Sweden, UK, Ireland or Germany. The Opportunity: At Grafana Labs, we build observability tools that help users understand, respond to, and improve their systems – regardless of scale, complexity, or tech stack. We recently started a skunkworks initiative with ...

Site Reliability Engineer (SRE)

Hiring Organisation
Randstad Digital
Location
United Kingdom
Employment Type
Contract
Contract Rate
£55 - £60 per hour
Title: Site Reliability Engineer (SRE) - Observability Location: Remote About the Role We are looking for a Lead SRE to design, scale, and operate massive-scale observability systems that keep our global services online and performant. You will join an autonomous team of software engineers focused on solving complex data infrastructure … Thanos/Cortex, Kafka, the ELK stack, Ansible, or Consul . Comfortable diving into unfamiliar codebases and participating in an on-call rotation. Keywords: Observability, Monitoring, SRE, Site Reliability Engineering, DevOps, ElasticSearch, ELK, Prometheus, Kafka, Terraform, Linux, Bare Metal Randstad Technologies is acting as an Employment Business in relation ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
robust security, access controls, secrets management, encryption, and audit trails directly into the infrastructure foundations to meet strict regulatory standards from day one. Data & Observability: Own the underlying data and event backbone required for transactional correctness, alongside the observability stack needed for rapid, real‐time issue detection and resolution. Qualifications ...

Staff Software Engineer (Reliability & Platform)

Hiring Organisation
Jobleads-UK
Location
United Kingdom
rather than handing work off between silos. In this role, you will focus on improving production operability and reliability by debugging complex issues, strengthening observability, and eliminating recurring problems at the source. Responsibilities Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source … resolve (MTTR), and recurrence rates for issues Identify systemic issues and eliminate recurring problems through code fixes, architecture improvements, and better operational tooling Improve observability across services — logs, metrics, and alerting — for faster diagnosis and resolution Design and improve debugging workflows, runbooks, and internal tooling for engineers Reduce operational burden ...

SRE Engineer

Hiring Organisation
Pinnacle Technical Resources
Location
Jersey City, New Jersey, United States
Employment Type
Permanent
Salary
USD 65 Annual
business processes; i.e., the ability to understand 'the business' and the impact of technology solutions Experience working with thirdparty applications and integrations Familiarity with observability practices such as white and black box monitoring, service level objective alerting, and telemetry collection using tools like Grafana, Dynatrace, Prometheus, Datadog and Splunk. Recognize … eliminate toil through systems engineering or automation, and implement observability patterns to improve service level indicators, objectives monitoring, and alerting solutions. Note: Pay Range: $60 - $65 The specific compensation for this position will be determined by a number of factors, including the scope, complexity and location of the role ...

SRE Managing Consultant

Hiring Organisation
Akkodis
Location
City, London, United Kingdom
Employment Type
Permanent
Salary
GBP 90,000 - 100,000 Annual
include: Define and embed SRE engagement models aligned to modern engineering and traditional ITSM/ITIL practices Establish SLIs, SLOs, and Error Budgets Shape observability strategies using metrics, logs, and traces Design incident response models and post-incident learning loops Reduce toil through automation and engineering excellence Deliver SRE capability … Looking For Extensive experience in SRE, cloud operations, or DevOps Proven consulting or advisory background Experience with AWS, Azure, or GCP Strong observability and incident management expertise Ability to obtain UK SC clearance Modis International Ltd acts as an employment agency for permanent recruitment and an employment business ...

Senior Director Technology

Hiring Organisation
Jobleads-UK
Location
Langley, England, United Kingdom
buildsourproductionAWSplatform, and onboards development teams and products to this platform.You’llbe responsible fordriving the move toinfrastructure as code, account management, automated pipelines, observability, resilience, operating coverage and cost control. You’llalso help define how the platform supports AI, data engineering and high-scale product demand, including where AWS Bedrock … standards for infrastructure as code, CI/CD, AWS account management, platform guardrails and developer enablement. Improve the operational model for the platform, including observability, incident response, reliability and 24/7 support coverage. Partner with Product and Commercial teams to get ahead of major demand changes, customer commitments ...

Staff Security Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
engineer who enjoys solving complex security challenges at scale. You’ll work across engineering, data, AI and digital workplace teams to build security observability, automate control assurance and influence how security is embedded into products, platforms and processes. If you're passionate about turning security data into actionable insight, building … raising the security maturity of a fast-moving technology organisation, we'd love to hear from you. About the role Designing and building security observability capabilities that provide meaningful visibility across systems, infrastructure and applications Developing automated control monitoring, evidence collection and continuous testing solutions that strengthen security governance Partnering ...

Operations Engineer

Hiring Organisation
ASCENT PROFESSIONAL SERVICES LTD
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
continuity. Key Responsibilities Provide operational support for enterprise platforms, applications, integrations, and associated technologies. Monitor system health, availability, and performance using monitoring, alerting, and observability tools. Analyse, troubleshoot, and resolve incidents affecting services and platforms. Perform root cause analysis and contribute to implementing permanent solutions to prevent recurring issues. Coordinate … within IT operations, support engineering, or service management environments. Experience supporting business-critical production services and operational platforms. Knowledge of monitoring, logging, alerting, and observability practices. Experience working with incident, problem, change, and release management processes. Excellent communication skills with the ability to collaborate effectively across multiple technical and business ...

Platform Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Worcester, England, United Kingdom
direction for a portfolio of cloud platform services (for example landing zones, CI/CD enablement, IaC patterns, identity and access guardrails, runtime patterns, observability, vulnerability scanning, backup and DR), and guide the team in building and evolving them. Ensure services are self-service, discoverable, and opinionated, making the secure … production operation of platform services through the team: clear service ownership, SLOs and SLAs, error budgets, on‐call, and incident management. Build strong observability for platform services and connect technical signals to customer and product impact. Drive improvements in reliability, security, performance efficiency, and cost optimisation, including ...

Engineer, Storage Services

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
service improvement initiatives led by Storage Services leadership. Reliability, Monitoring, and Incident Response Monitor service health, performance, and capacity across the storage estate. Improve observability, alerting, and operational readiness for production storage services. Troubleshoot incidents, performance issues, and service problems using a structured approach. Support root cause analysis and help … storage portfolio. KPIs Service reliability and operational excellence Storage service performance and capacity health Incident resolution and root cause follow-through Progress on automation, observability, and service improvements About You Core Skills Experience as a Storage Engineer, Infrastructure Engineer, Systems Engineer, or Platform Engineer role, ideally in cloud, hosting, infrastructure ...

Integration Architect - SAP SAAS Products

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
4HANA Cloud, SuccessFactors, Ariba, Concur, Datasphere SAC) and SAP BTP. Define patterns, govern APIs/events, ensure secure, resilient data flows, and drive standardization, observability, and compliance. Core Responsibilities Architecture & Standards • Define canonical integration patterns (API-led, event-driven, batch/EDI) and reference architectures on SAP BTP. • Establish guidelines … SuccessFactors, Ariba, Concur, Datasphere, SAC, and S/4HANA Cloud integrations. • Strong security fundamentals (OAuth2, SAML, JWT, SCIM) and compliance awareness. • Hands-on with observability (Cloud ALM), performance tuning, and reliability engineering. • Experience with agile delivery, CI/CD (Git-based pipelines), and test automation for integrations. • Preferred Qualifications ...

Software Engineer, GenAI Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in. Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner … especially in Python and distributed systems. Experience building production services, APIs, data pipelines, or ML infrastructure at scale. Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization. Hands‐on experience with LLM inference and/or fine‐tuning of open‐weight models ...

Senior Cloud Infrastructure Engineer VMware

Hiring Organisation
100% IT Recruitment Ltd
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
infrastructure issues. Lead technical investigations and major incident recovery. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. Support … experience with Veeam Backup & Replication. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability ...

Data Engineer

Hiring Organisation
Global
Location
Greater London, United Kingdom
Employment Type
Full Time
/CD and infrastructure as code. Create reusable components and maintain clear technical documentation. • Quality & Governance (10%) : Implement robust data validation, testing, lineage and observability to ensure high-quality, trusted datasets. Support governance and privacy-conscious data handling. • Collaboration & Enablement (10%) : Partner with Data Science, MLOps, Product and commercial teams … cloud environments (preferably AWS) • Engineering Best Practice: Knowledge of CI/CD, testing, version control and infrastructure as code • Data Quality & Governance: Understanding of observability, validation and maintaining reliable data systems • Collaboration & Communication: Ability to translate business and data science needs into scalable solutions and communicate clearly with stakeholders • Mindset ...

Cloud SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
databases. Day to day, you will monitor service health, resolve incidents, execute controlled changes, maintain patching and hardening standards, and improve reliability through automation, observability, and clear operational documentation. This is a 6-month contract, to work remotely (outside IR35) Key Responsibilities Infrastructure discovery and assessment - Build and maintain … AIOps and self-healing to detect and resolve issues earlier. Cloud administration - Administer compute, storage, and OS-layer services in Azure; support application infrastructure. Observability & monitoring - Implement and tune monitoring, alerting, and dashboards; improve signal quality and reduce noise. Security & compliance - Apply hardening, patching, and access controls in line with ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
moment when the work is genuinely changing shape. Over the last year we've hardened the platform, reduced cost, and built serious observability into our highest-volume systems. The next year is about scaling that work, absorbing infrastructure from a recent acquisition, and being thoughtful about how AI shows … thrive. Here's what that looks like in practice: Month 1 : You're onboarded across our AWS estate, Terraform, and observability stack. You've completed your first on-call shift with support from the team, landed your first PR in the DevOps repo, and started working Claude Enterprise into your ...

IT Delivery - Lead Cloud and Platform Engineer - IT1

Hiring Organisation
Jobleads-UK
Location
Newcastle upon Tyne, England, United Kingdom
platform, leading how applications are built, deployed, and operated across Parkdean Resorts. From Azure architecture and AKS cluster management to CI/CD pipelines, observability, and Infrastructure-as-Code, you’ll bring it all together into one cohesive, scalable platform. What you will be doing... Owning and evolving our cloud … leading CI/CD across Azure DevOps – ensuring reliable, repeatable releases Driving consistency across environments – building stable, trusted platforms for all engineering teams Embedding observability and monitoring – with full visibility across production environments Implementing Infrastructure-as-Code – using Bicep and/or Terraform Strengthening security and governance – including secrets management ...

Lead AI Engineer

Hiring Organisation
Capco
Location
Borough of Tameside, United Kingdom
Employment Type
Full Time
LLMs and multi-modal models at scale Strong engineering background in Python with proven backend and API development skills Solid understanding of scalable MLOps, observability, and cloud-native AI deployment Excellent communication, problem-solving, and project management skills in agile environments Bonus Points For Experience with agentic frameworks (e.g., LangChain … LlamaIndex) Experience in deep learning frameworks and front-end development Familiarity with Langfuse, Langsmith, or other LLM observability tools Understanding of Model Context Protocol and bias/hallucination mitigation techniques Previous success in integrating GenAI solutions into enterprise-scale systems Why Join Capco Deliver high-impact technology solutions for Tier ...

ML Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
GPUs and cloud infrastructure. Develop internal tools and abstractions and agentic systems that reduce operational overhead for researchers and engineers. Drive improvements across observability, automation, reliability, and developer experience. Collaborate closely with researchers and product engineers to understand pain points and turn them into robust platform capabilities. Contribute to technical … model serving systems in production. Supporting research or data‐intensive workloads. Working with GPU‐based systems or other performance‐sensitive infrastructure. Experience with observability and debugging in distributed systems. Familiarity with Terraform, Datadog, GitHub Actions, or similar tools. Bonus points for Experience building agentic or LLM‐powered internal tools. Experience ...

Principal 5G Network Core Architect (UK)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security, key management, and trust boundaries Leading the cloud-native design and deployment of 5GC, OAM, and supporting control components (Kubernetes, CNFs/VNFs, observability, resilience) Defining and maintaining OAM data models and workflows aligned with standard management frameworks (O-RAN SMO/O1/O2, NETCONF/YANG) Contributing … security (3GPP security, IPsec/TLS, PKI, RBAC, logging/audit) Experience with cloud-native telco platforms (Kubernetes, CNFs/VNFs, Helm/Operators, observability stacks) Hands-on lab experience integrating 5GC + gNB + OAM using COTS components and standard interfaces (NG, F1/E1, O1/O2, NETCONF ...

Principal 5G Network Core Architect

Hiring Organisation
Jobleads-UK
Location
United Kingdom
security, key management, and trust boundaries Leading the cloud-native design and deployment of 5GC, OAM, and supporting control components (Kubernetes, CNFs/VNFs, observability, resilience) Defining and maintaining OAM data models and workflows aligned with standard management frameworks (O-RAN SMO/O1/O2, NETCONF/YANG) Contributing … security (3GPP security, IPsec/TLS, PKI, RBAC, logging/audit)Experience with cloud-native telco platforms (Kubernetes, CNFs/VNFs, Helm/Operators, observability stacks) Hands-on lab experience integrating 5GC + gNB + OAM using COTS components and standard interfaces (NG, F1/E1, O1/O2, NETCONF ...