1,726 to 1,750 of 1,901 Observability Jobs

Lead Engineer, Trading Platform Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
order routing, or FIX/exchange-native protocols).* Time Synchronization: Familiarity with PTP or other high-precision time synchronization for low-latency environments.* Observability: Experience using eBPF and tracing for observability in production.* Hardware Interfacing: Knowledge of RDMA, NIC offloads (TSO, LRO), or experience maintaining kernel modules and device ...

Observability Engineer: Data, Dashboards & Performance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Vercel is hiring a Software Engineer for Observability in London. This hybrid role focuses on designing and implementing observability features, enabling large-scale data ingestion, storage, and processing from distributed systems. You will develop cutting-edge visualization tools and ensure top-tier performance across the platform. You will collaborate with ...

Observability Software Engineer - Data & Dashboards

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Vercel is seeking a Software Engineer to join the Observability team in London. You will design, implement, and maintain observability features that help users monitor and understand their applications’ health and performance. The role offers in-office anchor days on Monday, Tuesday, and Friday for those within commuting distance; otherwise ...

Gen AI Engineer

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
based AI stack Experience with high volume document processing Familiarity with enterprise architecture security and compliance controls Exposure to monitoring model evaluation and AI observability tools Follow enterprise standards for security governance observability and performance We are a Disability Confident Employer : Capgemini is proud to be a Disability Confident Employer ...

Senior Software Engineer – Payment Capabilities

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Champion engineering excellence through clean, maintainable, and well-tested code, helping establish best practices across engineering teams. Drive operational excellence by designing systems with observability, monitoring, resilience, and supportability at their core. Utilise tools such as Dynatrace to improve platform visibility, monitoring, alerting, and service reliability. Requirements A senior backend … payment services, or highly transactional systems is advantageous. Passionate about software quality, testing, maintainability, and engineering best practices. Strong understanding of CI/CD, observability, production support, and operational excellence. Core Competencies Demonstrates expertise in designing and delivering scalable, resilient backend systems and APIs, with a strong focus on software ...

Test Lead / Test Architect

Hiring Organisation
Avensys Consulting UK
Location
London Area, United Kingdom
systems within a highly regulated banking environment. The successful candidate will drive automation architecture decisions, framework re-engineering, reporting modernization, CI/CD integration, observability enhancements, and structured knowledge transfer while working closely with Solution Architects, Product Owners, Development Teams, and Business Stakeholders. Experience Required 12-15+ years total … Provide technical guidance to SDETs and Automation Engineers Conduct code reviews and architecture reviews Support troubleshooting and defect resolution Drive continuous improvement initiatives Reporting & Observability Implement HTML and JSON reporting capabilities Develop unified reporting dashboards Improve automation observability and diagnostics Introduce structured logging and correlation IDs Delivery Governance Manage risks ...

Dynatrace/Observability Engineer £528 per day INSIDE

Hiring Organisation
17918
Location
Telford, Shropshire, United Kingdom
Title: Dynatrace/Observability Engineer Rate: £528 per day INSIDE Clearance Required: SC Eligible Duration: 6 months Location: Telford - 2 days min per month Mandatory Certifications: Dynatrace Associate Certification As a Dynatrace/Observability Engineer, you will be responsible for designing, implementing, and supporting monitoring solutions across a range ...

Network Support Engineer

Hiring Organisation
Pontoon
Location
Chester, Cheshire, England, United Kingdom
Employment Type
Contractor
Contract Rate
£500.00 - £550.00 per hour
availability risks. Ensure telemetry data supports operational decision-making , including incident triage, root cause analysis, and proactive issue identification. Partner with Monitoring and Observability architects to align device telemetry with broader network health and observability models . Validate data accuracy, completeness, and timeliness; identify telemetry gaps and work with stakeholders ...

platform engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Design, build, and operate core GenAI platform components, including an LLM routing gateway, vector search and RAG infrastructure, tool registry and MCP gateway, AI observability and evaluation tooling, and infrastructure for long-running agentic workflows; Own production-quality delivery of platform features from design through rollout, monitoring, and follow … Contribute to resilient system design with sensible APIs, failure handling, rate limiting, retries, idempotency, and safe change management; Improve reliability and observability through metrics, dashboards, alerting, incident follow-ups, and operational improvements; Partner with Applied AI Engineers and product teams to understand platform needs and help them build AI-powered ...

Lead Software Engineer - Agent Safety

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
enforce pre- and post-generation guardrails, managing the overarching governance of AI models operating within internal tools and platforms. Drive AI security and observability: Build out dedicated auth/permissions for internal AI agents, establish deep monitoring/observability pipelines, and define incident response protocols for AI‐specific anomalies. Verification … Inspect AI, Ragas, OpenAI Evals, NeMo Guardrails). Django experience and strong backend engineering patterns (security, performance, maintainability). Experience with Datadog for complex observability, tracing, and monitoring in AI environments. Familiarity with foundational AI engineering tooling like Pydantic AI, LiteLLM, or LangChain. Kraken is a certified Great Place ...

Senior Software Engineer - Login Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
teams to deliver a reliable and secure login experience. Contribute to architectural decisions and long-term technical strategy for the Login Platform. Improve tooling, observability, scalability, and performance across critical infrastructure services. Explore and champion new ideas to evolve our systems responsibly over time. Engineers on the team are involved … throughout the software lifecycle, including design, deployment, observability, incident response, and long-term platform evolution. Our technologies We work primarily in C++ and some Python, running mostly on Linux, with continued support for Solaris clients until full migration is complete. Because we sit deep in the infrastructure stack ...

AI Engineer

Hiring Organisation
McCabe & Barton
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £900 per day inside IR35
turn AI capability into real business impact. Key Responsibilities Responsibilities Design, build and maintain core AI platform components - LLM gateway, MCP connector layer, observability tooling, and privacy Proxy Develop and harden MCP connectors across data sources (M365, Salesforce, Moody's, internal systems) through PoC, pilot, and GA stages Build … production environments Experience building AI Solutions is essential Experience with agentic frameworks, tool use, and MCP or equivalent connector patterns desirable Familiarity with observability and evaluation tooling for AI systems Understanding of data platforms - Databricks or similar - and how to expose data safely to LLMs Security and governance awareness: prompt ...

Software Engineer (AI) Jobs UK 2026 – Visa Sponsorship Available

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Responsibilities As an AI Software Engineer, you will: Design and maintain highly available AI infrastructure and large-scale model-serving systems. Develop monitoring and observability tools to improve system reliability and performance. Define and manage Service Level Objectives (SLOs) for AI services. Lead incident response, root cause analysis, and system … with large-scale AI model training or inference infrastructure. Knowledge of GPU, TPU, or other AI accelerator technologies. Experience with cloud infrastructure, monitoring, and observability platforms. Understanding of networking technologies such as RDMA or InfiniBand. Strong problem-solving, collaboration, and communication skills. Experience contributing to open-source infrastructure ...

Gen AI Architect

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
production-grade AI systems using Amazon Bedrock, retrieval-augmented generation (RAG), agentic workflows, and cloud-native AWS services. Drive architecture standards, model orchestration, governance, observability, and operational excellence across the GenAI lifecycle while collaborating with engineering, security, compliance, and business stakeholders Hybrid working: The places that you work from … customization, prompt orchestration, retrieval pipelines, and agentic workflows Design agentic AI systems incorporating tool use, workflow orchestration, memory management, and autonomous decision flows Implement observability for prompts, model responses, vector retrieval quality, and agent execution workflows Integrate GenAI capabilities into enterprise applications, APIs, workflow platforms, and data ecosystems Work with ...

Lead Data Engineer

Hiring Organisation
Tenth Revolution Group
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £110,000 per annum
Python SQL Data Warehousing Data Lakes Data Pipelines APIs & Data Services Machine Learning Infrastructure AI & Machine Learning Systems Analytics Engineering Data Governance Data Observability Cloud-Native Architecture You'll Be Responsible For Designing and building scalable cloud-based data platforms. Creating robust ETL/ELT pipelines and data services. Developing … APIs and trusted data products for enterprise clients. Supporting machine learning and AI initiatives through high-quality data architecture. Establishing governance, lineage, quality and observability standards across the data estate. Working closely with Product, Engineering, Data Science and Leadership teams to influence strategic decisions. Helping define the long-term data ...

AI Engineer (Infrastructure SRE & Automation)

Hiring Organisation
Jobleads-UK
Location
Livingston, Scotland, United Kingdom
repetitive operational workflows across compute, storage, and networking layers. Integrate AI capabilities into CI/CD pipelines and infrastructure-as-code ecosystems. Data Engineering & Observability Build and maintain pipelines for high-volume telemetry data (metrics, logs, traces). Ensure data quality, labelling, and feature engineering for ML models. Leverage observability ...

Principal Cloud Engineer

Hiring Organisation
100 Percent
Location
Cardiff, South Glamorgan, United Kingdom
Employment Type
Permanent
Salary
£58000 - £60000/annum + Bonus + Benefits
infrastructure issues. Lead technical investigations and major incident recovery. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. Support … experience with Veeam Backup & Replication. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability ...

ITSM JiraSM Platform Engineer - Barclaycard Payments

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
ITSM change enablement workflows within Jira Service Management. Integrate change, incident, request, and CMDB/service management capabilities with monitoring, CI/CD, and observability platforms. Move change control from manual governance into platform‐enabled, policy‐driven guardrails that support speed, control, and regulatory confidence. Design, configure, and scale ITSM … Hands‐on engineering experience in JiraSM. Ability to design, configure, and scale ITSM workflows. Experience integrating ITSM platforms with monitoring, CI/CD, and observability tools. Experience automating service operations using scripting or workflow orchestration. Exposure to platform engineering, DevOps practices, and data models (CMDB) and service mapping in modern ...

Senior AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
agents and for storing agent memory. Azure API Management (APIM), particularly the secure exposure and management of large language models (LLMs). Azure observability, including Azure Monitor, Application Insights, and Azure Workbooks. Langfuse, including designing and conducting evaluations and assessing outcome quality and value relative to token usage and cost. … Datadog for logging, observability, and LLM analytics. RAG, vector databases, and advanced retrieval techniques. LiteLLM as an AI and Model Context Protocol (MCP) gateway. MCP gateways and their integration with MCP servers. Contemporary AI models-including, as of July 2026, Fable 5, GPT-5.6, and GLM-5-and an understanding ...

AI Engineer (Infrastructure SRE & Automation)

Hiring Organisation
Sky
Location
Livingston, West Lothian, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
Salary
GBP per hour
workflows across compute, storage, and networking layers . Integrate AI capabilities into CI/CD pipelines and infrastructure-as-code ecosystems . Data Engineering & Observability Build and maintain pipelines for high-volume telemetry data (metrics, logs, traces) . Ensure data quality, labelling, and feature engineering for ML models . Leverage … observability tools to inform AI systems . Reliability & Governance Define SLIs/SLOs for AI systems embedded in infrastructure workflows . Ensure robustness, explainability, and auditability of AI-driven decisions . Implement feedback loops to continuously improve model performance . AI-Driven SRE Enablement Develop ML/AI models ...

Technology Services Lead

Hiring Organisation
Bank of America
Location
Greater London, United Kingdom
Employment Type
Full Time
play a critical role in supporting high-frequency trading platforms across EMEA markets. This is an exciting opportunity to work with cutting-edge observability and service management technologies while contributing to the stability, resilience, and performance of mission-critical electronic trading systems. Role Description: As our Electronic Trading environment continues … measurable improvements in platform stability and reductions in incident volumes through continuous service improvement initiatives Champion automation, monitoring, and tooling enhancements to improve efficiency, observability, and operational control Collaborate with global teams, mentoring junior engineers and contributing to the development of overall engineering capability What we are looking ...

Integration Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Agent-ready API products, reusable integration assets, and governed API catalogs AI-enabled integration patterns across Salesforce, Data, ERP, customer, operational, and platform domains Observability, policy enforcement, auditability, and compliance for AI-driven integration flows The manager will oversee a broad portfolio of APIs, connectors, MCP-enabled tools, and integration … routing, policy enforcement, and agent-to-system orchestration. Ensure AI-enabled integrations follow strong standards for identity, consent, access control, data minimisation, rate limiting, observability, and auditability. Provide architectural guidance on complex use cases involving multi-system orchestration, agent workflows, human-in-the-loop controls, and event-driven automation. Ensure ...

Senior Software Engineer: Cloud Operating System

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
system evolution Collaborate closely with platform and infrastructure teams to deliver production-grade, secure, and scalable services Maintain high standards for code quality, observability, and maintainability through reviews and best practices What You Bring: Proven experience developing production systems in Golang Hands-on experience designing and deploying distributed systems … APIs (gRPC, REST) in production environments Solid understanding of Kubernetes internals (Scheduler, CNI, Operators, CSI, CAPI) Familiarity with CI/CD workflows, observability tooling, security, and automation best practices Collaborate across teams,contribute to architectural and development best practices A bias for ownership, accountability, and delivering high-quality software Preferred ...

Vice President, Build — Data, Engineering & AI

Hiring Organisation
Jobleads-UK
Location
Reigate and Banstead, England, United Kingdom
ready criteria through the Data Architect; run the Collibra dictionary, master & reference data operations and the governance council. Own data quality and observability: shift‐left checks, lineage and monitoring built in, not bolted on. Run data BAU — incidents, refreshes and access — baselined before any cost reduction is taken. Own data … access framework. Success measures Prioritized use cases delivered to production with named owner, evidence pack, run model and sunset criteria. Data quality, lineage and observability coverage across priority data products. Time from funded demand to production. Production reliability, incident rate and support performance. AI evaluation coverage, red‐team completion ...

Senior AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
agents and for storing agent memory. Azure API Management (APIM), particularly the secure exposure and management of large language models (LLMs). Azure observability, including Azure Monitor, Application Insights, and Azure Workbooks. Langfuse, including designing and conducting evaluations and assessing outcome quality and value relative to token usage and cost. … Datadog for logging, observability, and LLM analytics. RAG, vector databases, and advanced retrieval techniques LiteLLM as an AI and Model Context Protocol (MCP) gateway. MCP gateways and their integration with MCP servers Contemporary AI models—including, as of July 2026, Fable 5, GPT-5.6, and GLM-5—and an understanding ...