1,576 to 1,600 of 1,767 Permanent Observability Jobs

Platform Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Worcester, England, United Kingdom
direction for a portfolio of cloud platform services (for example landing zones, CI/CD enablement, IaC patterns, identity and access guardrails, runtime patterns, observability, vulnerability scanning, backup and DR), and guide the team in building and evolving them. Ensure services are self-service, discoverable, and opinionated, making the secure … production operation of platform services through the team: clear service ownership, SLOs and SLAs, error budgets, on‐call, and incident management. Build strong observability for platform services and connect technical signals to customer and product impact. Drive improvements in reliability, security, performance efficiency, and cost optimisation, including ...

Engineer, Storage Services

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
service improvement initiatives led by Storage Services leadership. Reliability, Monitoring, and Incident Response Monitor service health, performance, and capacity across the storage estate. Improve observability, alerting, and operational readiness for production storage services. Troubleshoot incidents, performance issues, and service problems using a structured approach. Support root cause analysis and help … storage portfolio. KPIs Service reliability and operational excellence Storage service performance and capacity health Incident resolution and root cause follow-through Progress on automation, observability, and service improvements About You Core Skills Experience as a Storage Engineer, Infrastructure Engineer, Systems Engineer, or Platform Engineer role, ideally in cloud, hosting, infrastructure ...

Integration Architect - SAP SAAS Products

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
4HANA Cloud, SuccessFactors, Ariba, Concur, Datasphere SAC) and SAP BTP. Define patterns, govern APIs/events, ensure secure, resilient data flows, and drive standardization, observability, and compliance. Core Responsibilities Architecture & Standards • Define canonical integration patterns (API-led, event-driven, batch/EDI) and reference architectures on SAP BTP. • Establish guidelines … SuccessFactors, Ariba, Concur, Datasphere, SAC, and S/4HANA Cloud integrations. • Strong security fundamentals (OAuth2, SAML, JWT, SCIM) and compliance awareness. • Hands-on with observability (Cloud ALM), performance tuning, and reliability engineering. • Experience with agile delivery, CI/CD (Git-based pipelines), and test automation for integrations. • Preferred Qualifications ...

Software Engineer, GenAI Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in. Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner … especially in Python and distributed systems. Experience building production services, APIs, data pipelines, or ML infrastructure at scale. Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization. Hands‐on experience with LLM inference and/or fine‐tuning of open‐weight models ...

Senior Cloud Infrastructure Engineer VMware

Hiring Organisation
100% IT Recruitment Ltd
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
infrastructure issues. Lead technical investigations and major incident recovery. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. Support … experience with Veeam Backup & Replication. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability ...

Data Engineer

Hiring Organisation
Global
Location
Greater London, United Kingdom
Employment Type
Full Time
/CD and infrastructure as code. Create reusable components and maintain clear technical documentation. • Quality & Governance (10%) : Implement robust data validation, testing, lineage and observability to ensure high-quality, trusted datasets. Support governance and privacy-conscious data handling. • Collaboration & Enablement (10%) : Partner with Data Science, MLOps, Product and commercial teams … cloud environments (preferably AWS) • Engineering Best Practice: Knowledge of CI/CD, testing, version control and infrastructure as code • Data Quality & Governance: Understanding of observability, validation and maintaining reliable data systems • Collaboration & Communication: Ability to translate business and data science needs into scalable solutions and communicate clearly with stakeholders • Mindset ...

Cloud SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
databases. Day to day, you will monitor service health, resolve incidents, execute controlled changes, maintain patching and hardening standards, and improve reliability through automation, observability, and clear operational documentation. This is a 6-month contract, to work remotely (outside IR35) Key Responsibilities Infrastructure discovery and assessment - Build and maintain … AIOps and self-healing to detect and resolve issues earlier. Cloud administration - Administer compute, storage, and OS-layer services in Azure; support application infrastructure. Observability & monitoring - Implement and tune monitoring, alerting, and dashboards; improve signal quality and reduce noise. Security & compliance - Apply hardening, patching, and access controls in line with ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
moment when the work is genuinely changing shape. Over the last year we've hardened the platform, reduced cost, and built serious observability into our highest-volume systems. The next year is about scaling that work, absorbing infrastructure from a recent acquisition, and being thoughtful about how AI shows … thrive. Here's what that looks like in practice: Month 1 : You're onboarded across our AWS estate, Terraform, and observability stack. You've completed your first on-call shift with support from the team, landed your first PR in the DevOps repo, and started working Claude Enterprise into your ...

IT Delivery - Lead Cloud and Platform Engineer - IT1

Hiring Organisation
Jobleads-UK
Location
Newcastle upon Tyne, England, United Kingdom
platform, leading how applications are built, deployed, and operated across Parkdean Resorts. From Azure architecture and AKS cluster management to CI/CD pipelines, observability, and Infrastructure-as-Code, you’ll bring it all together into one cohesive, scalable platform. What you will be doing... Owning and evolving our cloud … leading CI/CD across Azure DevOps – ensuring reliable, repeatable releases Driving consistency across environments – building stable, trusted platforms for all engineering teams Embedding observability and monitoring – with full visibility across production environments Implementing Infrastructure-as-Code – using Bicep and/or Terraform Strengthening security and governance – including secrets management ...

Lead AI Engineer

Hiring Organisation
Capco
Location
Borough of Tameside, United Kingdom
Employment Type
Full Time
LLMs and multi-modal models at scale Strong engineering background in Python with proven backend and API development skills Solid understanding of scalable MLOps, observability, and cloud-native AI deployment Excellent communication, problem-solving, and project management skills in agile environments Bonus Points For Experience with agentic frameworks (e.g., LangChain … LlamaIndex) Experience in deep learning frameworks and front-end development Familiarity with Langfuse, Langsmith, or other LLM observability tools Understanding of Model Context Protocol and bias/hallucination mitigation techniques Previous success in integrating GenAI solutions into enterprise-scale systems Why Join Capco Deliver high-impact technology solutions for Tier ...

ML Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
GPUs and cloud infrastructure. Develop internal tools and abstractions and agentic systems that reduce operational overhead for researchers and engineers. Drive improvements across observability, automation, reliability, and developer experience. Collaborate closely with researchers and product engineers to understand pain points and turn them into robust platform capabilities. Contribute to technical … model serving systems in production. Supporting research or data‐intensive workloads. Working with GPU‐based systems or other performance‐sensitive infrastructure. Experience with observability and debugging in distributed systems. Familiarity with Terraform, Datadog, GitHub Actions, or similar tools. Bonus points for Experience building agentic or LLM‐powered internal tools. Experience ...

Principal 5G Network Core Architect (UK)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security, key management, and trust boundaries Leading the cloud-native design and deployment of 5GC, OAM, and supporting control components (Kubernetes, CNFs/VNFs, observability, resilience) Defining and maintaining OAM data models and workflows aligned with standard management frameworks (O-RAN SMO/O1/O2, NETCONF/YANG) Contributing … security (3GPP security, IPsec/TLS, PKI, RBAC, logging/audit) Experience with cloud-native telco platforms (Kubernetes, CNFs/VNFs, Helm/Operators, observability stacks) Hands-on lab experience integrating 5GC + gNB + OAM using COTS components and standard interfaces (NG, F1/E1, O1/O2, NETCONF ...

Principal 5G Network Core Architect

Hiring Organisation
Jobleads-UK
Location
United Kingdom
security, key management, and trust boundaries Leading the cloud-native design and deployment of 5GC, OAM, and supporting control components (Kubernetes, CNFs/VNFs, observability, resilience) Defining and maintaining OAM data models and workflows aligned with standard management frameworks (O-RAN SMO/O1/O2, NETCONF/YANG) Contributing … security (3GPP security, IPsec/TLS, PKI, RBAC, logging/audit)Experience with cloud-native telco platforms (Kubernetes, CNFs/VNFs, Helm/Operators, observability stacks) Hands-on lab experience integrating 5GC + gNB + OAM using COTS components and standard interfaces (NG, F1/E1, O1/O2, NETCONF ...

Senior Sales Engineer - Public Sector (UK)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscapeWho You Are:Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customerSomeone with strong written and oral communication … listed above may vary based on the country of your employment and the nature of your employment with Datadog.About Datadog:Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Senior Security Engineer: Observability, Automation & Risk

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
seeking a Staff Security Engineer to advance security observability, automate controls and strengthen governance across its global tech estate. This senior individual contributor role partners with engineering, data, AI and digital workplace teams to raise security maturity in a fast-moving environment. You will design scalable security observability, automate assurance ...

Senior AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
cutting‐edge techniques with pragmatism to deliver measurable impact. Apply strong software engineering principles, such as modularity, testing, code reviews, CI/CD and observability, to ensure AI systems are reliable, maintainable, production‐ready and can be readily adapted to future developments. Choose the right approach for the problem … engineers, and PMs, to scope and ship AI features iteratively. Ability to reason about system behavior end‐to‐end, including model performance, latency, and observability, and how these impact user experience. Clear, structured communicator, comfortable documenting and defending architectural decisions and engaging in thoughtful technical debate. Not required ...

Senior Data Analyst - Product Reliability

Hiring Organisation
Wise
Location
Greater London, United Kingdom
Employment Type
Full Time
Salary
60000 to 85000 GBP Annually
Grafana, Lightdash, or Superset Ability to build and manage data pipelines that are modularised and scalable, using tools like DBT and Airflow Familiarity with observability/reliability concepts (SLIs, SLOs, incidents) Some experience with Python/data transformation (DBT, etc.) is helpful This is not a data engineering role, depth … KPIs Scaling our infrastructure at Wise: how we make it work Wise's Tech Stack 2025 Measuring meaningful availability/uptime of Wise Why Observability is a must for product engineering teams Grafana Mimir Compaction: From Bottleneck to Savings For everyone, everywhere. We're people building money without borders — without ...

VP, Head of Data Science

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
effective. Scale practical AI adoption Design and implement AI workflows as skills, models and MCP tools in OpenWebUI, with cost control, provider routing and observability built in. Partner with other departments to turn AI capabilities into everyday workflows, reusable tools and knowledge layers that improve internal productivity and client delivery. … open‐source AI software, including LiteLLM and OpenWebUI, to maximise the internal value of AI and build durable institutional capability. Reliability, cost control, observability, instruction‐following, usability and adoption matter most. The team designs reusable workflows, skills, models and MCP tools that support Penta’s move towards an AI‐native ...

Lead Software Engineer - Proxy/SSE Network Security

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
resilience outcomes. Drive operational excellence at scale for perimeter, proxy, and SSE services in the US, including incident, change, and problem management rigor, observability and resiliency validation practices, automation to improve repeatability and evidence quality, reduction of client and partner impact, and execution of Technology Lifecycle Management (TLM) and modernization … design, exception frameworks, audit-ready traceability, and measurable risk reduction reporting. Experience with large-scale operations for externally facing or security enforcement services, including observability strategy, resilience testing, incident response alignment, and reduction of repeat incidents and client-impacting events. Experience designing and operating hybrid edge architectures and cloud interconnect ...

Senior Pre-Sales Solutions Architect (Observability)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Engineer (Pre-Sales) for their London office. This senior role involves delivering technical presentations and collaborating with sales executives to engage with customers on observability solutions. The ideal candidate will have at least 5 years in a customer-facing role, exceptional communication skills, and in-depth knowledge of observability technologies. ...

Software Development Engineer III

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About us Our vision at Tesco is to become every customer's favourite way to shop, whether they are at home or out on the move. Our core purpose is ‘Serving our customers, communities and ...

Senior Software Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
Impact and Responsibilities We are seeking a Senior Software Engineer who thrives on untangling complex systems and modernising core infrastructure without breaking production. This is an exciting opportunity to modernise core C#/SQL systems ...

Integration Engineer (Software Engineering + AI Product Integration)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Role Summary As an AI Engineer, you will design, implement, and productionise behavioural learning systems that integrate directly into our digital products and workflows. Your focus will be on turning advanced behavioural, sequential, and causal ...

Senior Software Engineer: Build Scalable, Secure Systems

Hiring Organisation
Jobleads-UK
Location
Wimbledon, England, United Kingdom
Jobtailor in Wimbledon, United Kingdom is seeking an experienced software engineer to contribute across the full SDLC with Engineering, DevOps, and Product. You will turn business requirements into robust, secure technical solutions and design scalable ...

Senior SRE: Cloud Reliability, Observability & Automation

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Solutions Ltd is looking for a Senior Site Reliability Engineer in the United Kingdom. The role involves operating and maintaining production clusters and developing observability solutions while collaborating with teams to enhance platform reliability. The ideal candidate will have experience with cloud technologies, container orchestration, and strong scripting skills. Benefits ...