1,551 to 1,575 of 2,378 Observability Jobs in London

DevSecOps Architect - Director

Location
Greater London, England, United Kingdom
regulation, covering explainability, human oversight, drift monitoring and incident handling. Define the reference architecture for AI platform components: model gateway, retrieval layer, guardrails, observability and audit trail. Qualifications, Skills and Experience Significant experience architecting CI/CD and DevSecOps platforms at enterprise scale in a regulated environment. Demonstrable delivery … Veracode, Trivy, secrets management (HashiCorp Vault, cloud-native equivalents), OPA/policy-as-code, SBOM and SLSA concepts. Platform: Kubernetes, containers, Terraform, service mesh, observability stacks AI/ML: MLOps and LLMOps tooling, model registries, vector stores and retrieval architectures, evaluation frameworks, agent frameworks, AI gateway patterns Experience operating under ...

Full Stack Engineer

Hiring Organisation
Liberty Global
Location
London, UK
Employment Type
Full-time
features and platform capabilities end to end, from design through build, test and deployment into production. Contribute to production readiness: documentation, runbooks, monitoring and observability – ensuring nothing ships until it's genuinely ready. Maintain and improve existing platform components, addressing technical debt and performance issues with clear prioritisation. Standards & QualityChampion … with data pipelines, data quality processes and the infrastructure that supports AI/ML at scale. Experience with DevOps, CI/CD pipelines and observability tooling. Exposure to telecom or large-scale consumer-facing platforms. SKILLS & BEHAVIOURAL COMPETENCIESTechnical Depth & JudgementConsistently makes sound technical decisions, balancing short-term delivery with long ...

Senior AI Platform Engineer

Hiring Organisation
9fin
Location
London, UK
Employment Type
Full-time
developer tooling that enable self-service AI development across engineering teams. Design secure, scalable deployment pipelines for AI models and applications. Build AI observability capabilities including monitoring, tracing, evaluation, cost optimisation, and production quality measurement. Collaborate closely with AI Engineers, Backend Engineers and Engineering Leadership to define platform architecture … monitor, and operate AI services in production. AI Operations & ObservabilityHave experience implementing monitoring, tracing, evaluation, and cost optimisation for AI systems. Have experience with observability solutions such as Arize Phoenix, Langfuse, or LangsmithUnderstand the operational challenges of deploying LLM-powered applications, including latency, reliability, hallucination monitoring, and model quality evaluation. ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
enablement work Established objective measures of "done" and documented assumptions clearly, so both infrastructure and tooling initiatives had transparent success criteria Contributed to observability as code initiatives, both inside the team and across engineering Infrastructure & reliability Provided a highly available and performant platform, including our AWS, Kubernetes (EKS) and MongoDB … genuine interest in developer experience: you get satisfaction from removing friction for other engineers, not just from the infrastructure itself Worked with monitoring and observability tools such as Datadog General Proficiency in at least one of: Go, Python, Typescript, or strong shell scripting skills Strong communication skills and comfort influencing ...

Principal - Head of Compliance Architecture

Hiring Organisation
Northern Trust
Location
London, UK
Employment Type
Full-time
Drives enterprise resiliency, recoverability, and operational excellence initiatives including active-active architectures, hot-hot disaster recovery models, blue-green deployment strategies, high availability design, observability, fault tolerance, and business continuity readiness.* Establishes and improves engineering discipline and software delivery practices, including CI/CD pipelines, DevSecOps operating models, automated testing … warehouses, master data management, metadata management, AI enablement, and advanced analytics platforms. Enterprise resiliency, disaster recovery, active-active architectures, high availability, blue-green deployments, observability, and operational readiness frameworks. Modern DevSecOps and software engineering disciplines including CI/CD, automated testing, infrastructure-as-code, secure software delivery, and software lifecycle ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices. Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks. Linux & Automation: Strong Linux and Git fundamentals with Bash ...

Staff Software Engineer, Liquidity Management (C#/.NET) New London, UK

Location
Greater London, England, United Kingdom
work that distributes effectively across the team Establish coding standards, review practices, and testing strategies that improve overall code quality Champion engineering methodologies including observability, incident response, and production excellence Share your enthusiasm for tech trends, explore and learn new technologies, engage with tech communities, mentor fellow engineers, and lead … cloud platform (preferably Azure) Experience with containerization (Docker) and orchestration (Kubernetes) Understanding of infrastructure as code and CI/CD pipeline design Knowledge of observability: logging, metrics, tracing, and alerting strategies Experience with microservices deployment patterns and service mesh concepts Core Technologies: Deep expertise with C#/.NET Expert ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
features and platform capabilities end to end, from design through build, test and deployment into production. Contribute to production readiness: documentation, runbooks, monitoring and observability – ensuring nothing ships until it’s genuinely ready. Maintain and improve existing platform components, addressing technical debt and performance issues with clear prioritisation. Standards & Quality … with data pipelines, data quality processes and the infrastructure that supports AI/ML at scale. Experience with DevOps, CI/CD pipelines and observability tooling. Exposure to telecom or large-scale consumer-facing platforms. SKILLS & BEHAVIOURAL COMPETENCIES Technical Depth & Judgement Consistently makes sound technical decisions, balancing short-term delivery ...

Senior Software Engineer - Customer & Claims

Location
Greater London, England, United Kingdom
Create an environment where less experienced engineers can grow and do their best work. Continuous Improvement Drive improvements to engineering practices, testing, deployment and observability within the domain. Continuously grow your own capability and that of the engineers around you. Contribute to Homeprotect's wider engineering standards as they mature. … oriented language), and familiarity with a modern frontend framework. Deep, hands‐on understanding of modern software engineering practices: continuous integration and deployment, automated testing, observability and cloud‐native development. Experience operating production systems, in a public cloud environment such as Azure or GCP, that need to be reliable, secure ...

AI and Cloud Director - Technology Consulting (AI, data and analytics)

Hiring Organisation
Business Integration Partners
Location
London, UK
Employment Type
Full-time
AI and Cloud DirectorLocation: London/HybridBusiness: BIP UKPractice: xTech - AI, Data and CloudReporting to: UK xTech LeadershipEmployment type: Full timeAbout BIP UKFounded in 2003, BIP is an international consulting firm with more than 6 ...

Software Engineer (Backend)

Location
Greater London, England, United Kingdom
Software Engineer (backend) Location: London or New York - in office 4 days a week About Lorum Global payments are not broken. Incentives are. Clearing has been deprioritized inside balance sheet driven institutions whose models rely ...

AI & ML Engineer

Location
Greater London, England, United Kingdom
About the role The AI & ML Engineering team accelerates the adoption of AI across the business, championing innovation while ensuring our machine learning solutions are robust, scalable, and cost-efficient. We enable teams to solve ...

Senior Data Engineer

Hiring Organisation
McKesson
Location
London, UK
Employment Type
Full-time
Role OverviewThe Senior Data Engineer is the technical owner of the ClarusONE data platform. This is a hands-on engineering role with real ownership: alongside building and delivering data pipelines, you will raise the engineering ...

Senior Infrastructure Engineer

Location
Greater London, England, United Kingdom
mission is to abstract away multi-cloud complexity and reduce cognitive load for other developers. We treat our platform infrastructure which spans, orchestration, observability, and data. As a first-class product, enabling our internal teams and global clients to operate securely and reliably at massive scale. Duties Scale the Core … engineers and bank SREs. Design for Resiliency: Engineer robust, fault-tolerant architectures and Disaster Recovery/Business Continuity Planning (DR/BCP) systems. Advance Observability: Enhance our comprehensive observability suites that are shared across the entire Thought Machine product ecosystem, providing clients with deep, out-of-the-box monitoring ...

Senior Data Platform Engineer

Location
Greater London, England, United Kingdom
policy-as-code, and action logging. Build platform-, tool-, and agent-level kill switches, alongside dry-run/safe-testing modes for agent workflows. Observability & FinOps Implement AI system observability, including prompt logging, output monitoring, quality scoring, and drift detection. Establish operational monitoring for gateway usage, latency, error rates … plain English. And to really stand out from the crowd Experience with AI gateways, LLM proxies, or model-routing patterns. Exposure to AI observability tools, vector databases (Databricks AI Search), RAG pipelines, or non-deterministic system monitoring. Hands-on experience implementing policy-as-code, action logging, or AI kill switches. ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … call rota and provide on-call support for other SRE engineers.* Can write advanced automation scripts for incident response, including failovers and rollbacks.**Observability** * Has a deep technical understanding of observability techniques across the full stack and can bring clarity to complex incidents or performance issues.* Able to create templated ...

Lead Cloud Engineer

Location
Greater London, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£780 - £830 per day
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone … shape how thousands of infrastructure and application events are detected, correlated and actioned across EMEA, and you'll have genuine scope to modernise observability capability rather than simply keep the lights on. This is a hands-on technical leadership role with no direct reports, so your influence comes from your ...

Enterprise AI Deployment Architect

Location
Greater London, England, United Kingdom
commercial execution across product, engineering, and sales teams. You will translate business needs into deployment playbooks, drive measurable outcomes, and ensure governance, security, and observability across customer environments. #J-18808-Ljbffr ...

Site Reliability Engineer — Cloud & Live Ops (Hybrid)

Location
Uxbridge, England, United Kingdom
resilient, secure, and highly available platforms underpinning live, broadcast‐adjacent services. Based at Stockley Park in Uxbridge with hybrid options, the role involves improving observability, incident response, automation, disaster recovery, and collaborating with engineering, operations and project stakeholders. #J-18808-Ljbffr ...

SRE Engineer – Hybrid Cloud Reliability & Automation

Location
Greater London, England, United Kingdom
ensure ISO 27001 security alignment. This hands‐on role has significant influence across engineering, security and operations teams, with a focus on automation, observability and end‐to‐end service reliability. #J-18808-Ljbffr ...

Senior ML Platform & Ops Engineer - Hybrid (London)

Location
Greater London, England, United Kingdom
Preply is hiring a Senior ML Platform/Ops Engineer in London. You will help productionize ML systems with reliability, performance, and observability, working at the intersection of ML, data engineering, and cloud infrastructure. You’ll collaborate with ML Scientists, Backend and Data Engineers to shape the ML lifecycle foundations. ...

Lead Software Engineer, AI Safety & Evals

Location
Greater London, England, United Kingdom
evaluating models, safety guardrails, and governance across internal platforms, ensuring secure and deterministic AI behavior. You will lead continuous verification, red teaming, and observability initiatives, partnering with AI Foundations and engineering teams to scale safe AI across Kraken’s internal tools. #J-18808-Ljbffr ...

Senior ML Engineer, Real-Time Ranking & Personalization

Location
Greater London, England, United Kingdom
personalised rankings across search results and recommendations. You will collaborate with ML Scientists, Backend Engineers and MLOps to build scalable ranking models, improve latency, observability and real-time serving, and help harden the ML infrastructure for ranking systems. #J-18808-Ljbffr ...

Product-Led Engineering Manager

Location
Greater London, England, United Kingdom
Managers, and ensure the team ships reliably and safely with high standards of craft. You will own the technical roadmap, drive CI/CD, observability, and AI-assisted development, and align with engineering standards across the organisation. #J-18808-Ljbffr ...