3,801 to 3,825 of 4,209 Observability Jobs

Senior Network Reliability Engineer - Azure Networking

Location
City Of London, England, United Kingdom
hybrid network estate across Azure and on-premises, with a hands-on, incident-driven focus. You will lead complex networking challenges, drive automation and observability, and contribute to a more Azure-native architecture from our Wimbledon office. You will partner with Cloud, Security and Engineering teams, mentor peers, and help ...

Azure SRE & Reliability Engineer – Multi-Region

Location
Belfast City District, Northern Ireland, United Kingdom
monitoring, and incident response across a multi-cloud, multi-region environment. You will work with senior engineers and cross-functional teams to improve reliability, observability, and best practices, while contributing to disaster recovery and on-call incident management. #J-18808-Ljbffr ...

Payments AI/ML Leader - GenAI, MLOps & Strategy

Location
Greater London, England, United Kingdom
will partner with product, operations, risk, and technology teams, mentor engineers and data scientists, and establish reusable patterns for scalable MLOps, governance, and observability in production. #J-18808-Ljbffr ...

Platform Engineer – Scale GPU Infra for AI Platform

Location
Greater London, England, United Kingdom
infrastructure powering large-scale GPU compute, ensuring reliability, perf and fast developer feedback. You’ll work across Kubernetes, containers, cloud providers and observability tools to keep the platform humming and accelerate experimentation. This is a hands-on role with real ownership and impact. #J-18808-Ljbffr ...

Platform SRE Lead: On-Call, Reliability & Cloud Ops

Location
Greater London, England, United Kingdom
platform. Open to mid-level or senior SREs, the role is ops-heavy with focus on health of live systems, Kubernetes, cloud infra, deployments, observability, and collaboration with product teams. Hybrid work with 3 days in the office. #J-18808-Ljbffr ...

Senior AI Engineer – GenAI, RAG & Agents (Hybrid)

Location
Greater London, England, United Kingdom
including Generative AI, RAG, and AI agents, collaborating with architecture, security, and infrastructure teams. You will lead the adoption of best practices, model evaluation, observability, and governance across SDLC while translating #J-18808-Ljbffr ...

Operations Team Lead — Production Reliability

Location
Chesterfield, England, United Kingdom
hands-on leadership role to scale reliability. You will shape incident response, runbooks, and on-call rotations, and lead a growing team to improve observability, MTTR, and availability. The ideal candidate has strong SRE/DevOps experience, calm communication under pressure, and a passion for preventing outages through systemic fixes. ...

Analytics Engineer: AI-Driven Data & Trust

Location
Greater London, England, United Kingdom
collaborate with senior engineers to raise our analytics depth. You'll work with dbt, Snowflake, Looker and modern data platforms, own data quality and observability, and enable self-service BI for the business. #J-18808-Ljbffr ...

Ops Lead Engineer – Big Data Platform

Location
Greater London, England, United Kingdom
efficiency across the platform. Oversee CI/CD pipelines and deployments, ensuring reliable, safe, and compliant delivery of data platform changes. Champion monitoring, observability, and automation to detect and resolve issues proactively while reducing manual intervention. Develop and maintain operational runbooks, escalation protocols, and incident playbooks to strengthen resilience. Partner … reduce manual intervention. FinOps Mindset: Experience in cost management, usage reporting, and running forums with business stakeholders to drive accountability and efficiency. Monitoring & Observability: Knowledge of modern monitoring, alerting, and data quality frameworks to ensure proactive platform health management. #J-18808-Ljbffr ...

Platform Engineer: AI Tools & Infra for Developers

Location
Rochdale, England, United Kingdom
Group is seeking an experienced engineer to build internal SDKs and gateways, and to own developer-facing tooling. You will contribute to shared evaluation, observability and cost-attribution tooling, and design multi-tenant data infrastructure with strong isolation. The role emphasizes platform engineering, distributed systems expertise, and ownership of tooling ...

Senior AI Engineer — Production LLMs (Remote Europe)

Location
United Kingdom
Senior AI Engineer to build AI agents, agentic workflows, and production-grade AI capabilities. You will own end-to-end implementation, evaluation, and observability, interfacing with product and platform teams to deliver value to 3M+ traders. Location: Europe (EU/UK); fully remote with overlap to Central European hours. Start ...

Senior AI Engineer — Architect of Agentic AI for Industry

Location
Greater London, England, United Kingdom
architect production‐grade infrastructure enabling product teams, FDEs, and customers to compose advanced AI workflows safely and reliably. You will lead platform decisions on observability, deployment, sandboxing, evaluations, and governance, ensuring deep tracing, cost tracking, and robust identity controls across the stack. #J-18808-Ljbffr ...

Senior Java Backend Engineer - Hybrid London

Location
City Of London, England, United Kingdom
this hybrid role, you will collaborate with business analysts and stakeholders, develop RESTful APIs, optimize data access, and contribute to system stability, performance, and observability #J-18808-Ljbffr ...

Full-Stack AI Engineer for Legal Tech Products

Location
Greater London, England, United Kingdom
deliver AI-enabled legal services. The AI Engineer will ship end-to-end AI workflows, integrate tools and data sources, and ensure security and observability across the product lifecycle. You will work with Python, TypeScript, and modern AI tooling in a hands-on, production-focused environment within a multidisciplinary team ...

MLOps Engineer — Real-Time AI in Production (Hybrid London)

Location
Greater London, England, United Kingdom
production systems, collaborating with research and engineering teams. The role emphasizes designing end-to-end ML pipelines, deploying and monitoring models in production, building observability and CI/CD, and enabling reproducible experimentation. #J-18808-Ljbffr ...

Senior ML Engineer - Conversational AI & GenAI | Hybrid

Location
Greater London, England, United Kingdom
full ML lifecycle from design to production, deploying models and building robust, scalable systems in collaboration with product managers. Strong focus on reliability, observability, and GenAI-enabled solutions. #J-18808-Ljbffr ...

24/7 HPC Infra SRE for AI & GPU Compute

Location
Gloucester, England, United Kingdom
work across network, storage, virtualization and orchestration with hands‐on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance benchmarking. This role champions observability, automation and on‐call reliability, shaping next‐gen HPC platforms within a globally distributed team. #J-18808-Ljbffr ...

Senior AWS Data Engineer: Build Secure Data Platforms

Location
Greater London, England, United Kingdom
solutions, establish engineering standards and patterns, and mentor other engineers. The role blends hands-on execution with technical leadership, enabling data quality, governance and observability while influencing #J-18808-Ljbffr ...

Staff Infrastructure Engineer - Cloud Platform & Security

Location
Greater London, England, United Kingdom
projects from inception to production, write RFCs and ADRs, and mentor teammates while shaping future capabilities of Engine — with a strong emphasis on automation, observability and best practices. #J-18808-Ljbffr ...

Investment Tech Engineer — AI & Data Pipelines (Hybrid)

Location
United Kingdom
platform. You will work on cloud-based tools (Streamlit/ReactJS), apply AI-assisted tooling, and help drive CI/CD, data quality, and observability across the lifecycle. Collaborating with Cloud Teams, Enterprise Architecture and Security, you will translate business requirements into robust technical solutions, build data pipelines for analytics ...

Senior IP Network Engineer & SRE Lead

Location
Ipswich, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...

SRE Associate: Build Reliable, Scalable Systems

Location
Birmingham, England, United Kingdom
Birmingham seeks an Associate Site Reliability Engineer to join Core Engineering. You will help build and operate high-performing, scalable systems, focusing on reliability, observability, and automation across critical services in a fast-paced financial environment. You will collaborate with engineering teams to reduce downtime, improve deployment speed, and implement ...

Senior DevOps Engineer — Remote, Global Cloud Ops

Location
Belfast City District, Northern Ireland, United Kingdom
users, and data, across multiple regions. You will influence architecture, reliability, automation, and cloud efficiency, collaborating with engineers worldwide. The position emphasizes cost optimization, observability, and scalable deployment practices, with occasional travel to Belfast or remote collaboration. #J-18808-Ljbffr ...

Operations Team Lead: Scale Reliability & Incident Mastery

Location
Stockport, England, United Kingdom
move from reactive firefighting to proactive reliability engineering. You will lead the Operations team, set standards, manage incidents, on-call rotations, and drive observability improvements. A strong SRE/DevOps background and calm, clear communication are essential. #J-18808-Ljbffr ...

Senior IP Network Engineer & SRE Lead

Location
Birmingham, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...