1,651 to 1,675 of 2,378 Observability Jobs in London

Head of Engineering, POS Application Platform

Location
Greater London, England, United Kingdom
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end‐to‐end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

Interim Principal Platform Engineer

Location
Greater London, England, United Kingdom
reliability, and continuous improvement. The responsibilities will be diverse, ranging from maintaining AWS resources through Infrastructure as Code (we use OpenTofu) through to enhancing observability tools (we embrace OpenTelemetry) and supporting AI agentic workflows. You will be working with different software engineering groups to understand their problems and coming … best cloud resource support. Contribute to our CI/CD pipelines to driver further efficiency of our engineers – automate everything where possible. Provide observability tools to engineers so they can self-discover issues. Find gaps in developer productivity and design, implement solutions to solve them. Provide technical mentorship to others ...

Engineering lead

Location
Greater London, England, United Kingdom
security scanning, and performance engineering. Drive CI/CD adoption and DevSecOps practices. Monitor engineering metrics, technical debt, and platform stability. Improve system reliability, observability, and operational readiness. Stakeholder Management Act as the primary technology contact for business and delivery stakeholders. Present engineering updates, risks, dependencies, and mitigation plans. Work … Technical Skills Engineering & Architecture REST APIs Domain Driven Design (DDD) Docker OpenShift DevOps & Automation Jenkins Terraform Ansible TDD/BDD Secure Coding Performance Engineering Observability & Monitoring Tools JIRA Confluence Splunk Dynatrace SonarQube Banking Domain Knowledge Preferred experience in: Retail Banking Commercial Banking Digital Banking Identity & Authentication Customer Journeys Payments Lending ...

Engineering lead

Location
Greater London, England, United Kingdom
Championtest automation, code quality, security scanning, and performanceengineering. DriveCI/CD adoption and DevSecOps practices. Monitorengineering metrics, technical debt, and platform stability. Improvesystem reliability, observability, and operational readiness. Stakeholder Management Actas the primary technology contact for business and delivery stakeholders. Presentengineering updates, risks, dependencies, and mitigation plans. Workclosely with Product … initiatives. Required Technical Skills Engineering & Architecture RESTAPIs DomainDriven Design (DDD) Docker OpenShift DevOps & Automation Jenkins GitHub/Bitbucket Terraform Ansible TDD/BDD SecureCoding Observability& Monitoring Tools JIRA Confluence Splunk Dynatrace SonarQube Banking Domain Knowledge Preferred experience in: RetailBanking CommercialBanking DigitalBanking CustomerJourneys Payments LendingPlatforms OpenBanking Risk& Compliance SolutionDesign Reviews ReleaseReadiness ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
across our systems. Improve developer experience through automation, self-serve tooling and infrastructure-as-code practices that increase engineering velocity. Own and evolve our observability, incident response and reliability practices to minimise downtime and improve system performance. Partner closely with product and engineering teams to support new services, migrations … production environments. Strong experience operating and scaling infrastructure components such as Kafka, Redis and managed relational databases like RDS. Strong understanding of networking, security, observability and distributed systems fundamentals. Experience improving CI/CD pipelines and deployment automation in fast-moving engineering teams. Comfortable debugging production issues and participating ...

Intelligent Automation Engineering Lead

Location
Greater London, England, United Kingdom
Assistants, Copilots, or AI Enablement Platforms. Knowledge of MCP (Model Context Protocol) and enterprise AI integration frameworks. Experience with AI governance, evaluation frameworks, observability, and LLMOps. Experience with enterprise integration technologies, APIs, middleware, and event-driven architectures. Hands-on experience across both Azure and AWS cloud environments. Familiarity with GitLab …/CD or similar DevOps toolchains. Experience with observability and monitoring tools such as Splunk. Experience within Financial Services, FinTech, Payments, or other highly regulated environments. Understanding of security, risk, compliance, and governance frameworks supporting enterprise AI adoption. How We Work: Delta Capita is an equal opportunity employer. We positively ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations ...

Core AI Engineer

Location
Greater London, England, United Kingdom
running centralised MCP servers, ensuring secure, reliable access to enterprise tools and data Owning platform reliability, performance and scalability across Kubernetes‐based infrastructure, including observability, capacity planning and incident response Building self‐service tooling and APIs to enable teams to provision and consume AI infrastructure independently Integrating platform services with … services Familiarity with sandboxing and workload isolation technologies Experience in quantitative finance or low‐latency systems AWS experience particularly in hybrid environments Experience with observability tooling such as Prometheus, Grafana or OpenTelemetry Contributions to open‐source projects in relevant domains Why join us? Highly competitive compensation plus annual discretionary bonus ...

AI Engineers London

Location
Greater London, England, United Kingdom
Experience implementing, maintaining and evaluating retrieval systems (vector/graph databases, ingestion pipelines, chunking strategies, retrieval techniques such as HyDE) Implement feedback loops and observability to continuously improve system performance Craft effective prompts and optimize for latency, cost, and quality across different model providers and configurations Required Skills and Experience … fine‐tuning is preferable to prompt engineering or RAG Experience with real‐time streaming, multimodal models, or search technologies like Elasticsearch Familiarity with model observability tools (LangSmith, Weights & Biases) and cost optimization strategies Experience in specialized verticals (financial services, energy, healthcare, legal, retail) with understanding of compliance, security, and responsible ...

Senior Software Engineer I

Location
City Of London, England, United Kingdom
maintain generative AI services and reusable components using mostly Python and a little bit Java. Defining and promote best practices in engineering, including scalability, observability, testing, and CI/CD. Contributing to system designs spanning multiple services and modules, aligning with architectural best practices. Collaborating with product, platform, and research … work collaboratively across functions in an Agile or Kanban environment. Nice to Have Experience operationalizing LLMs or building internal AI platforms. Familiarity with observability practices (metrics, logging, alerts). Exposure to knowledge graphs or semantic search systems. U.S. National Base Pay Range U.S. National Base Pay Range ...

Senior Data Engineer

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
Engineers to operationalise machine learning models. Designing data models and architectures that enable reliable, trusted and accessible data across the business. Improving data quality, observability, lineage and monitoring capabilities across critical data products. Supporting the delivery of media measurement, customer value, pricing and personalisation initiatives. Optimising data processing workloads … enabling data science, analytics or machine learning workloads through robust data infrastructure. Strong understanding of CI/CD, Infrastructure as Code, automated testing and observability practices. Experience working with large-scale customer, marketing, commercial or behavioural datasets. Proven ability to lead the design and delivery of complex technical solutions from ...

Senior Machine Learning Engineer (ML Platform)

Location
Greater London, England, United Kingdom
infrastructure. Improve automation across the ML lifecycle, including model packaging, deployment, versioning, monitoring, and release processes. Maintain and improve the reliability and observability of the ML platform and model-serving services, including logging, metrics, and alerting. Participate in the team's on-call rotation, investigate production incidents, and contribute … experience provisioning and managing cloud infrastructure with Terraform. Experience with CI/CD pipelines, automated testing, and Git-based development workflows. Familiarity with observability practices, including logging, metrics, alerting, and production troubleshooting. Strong grasp of software engineering principles and best practices. Experience contributing to or leading data warehouse architecture ...

Data Lead

Hiring Organisation
UBS
Location
London, UK
Employment Type
Full-time
orchestration (Airflow, Prefect, or similar) and reliability patterns such as retries, backfills, and exactly-once semantics. Deep experience with Azure, including AKS, Helm, EventHub, observability tools (e.g. Log Analytics) and infrastructure-as-code (e.g. Terraform). Advanced experience operating Delta tables at scale, covering transaction semantics, schema evolution, partitioning … Marimo. Batch and streaming processing proficiency, including late data handling, watermarking, backfills, and explicit throughput vs. latency trade-offs. Production data reliability and observability mindset, covering lineage, data quality checks, SLAs, CI/CD for data pipelines, and close collaboration with quants, risk managers, and traders to deliver trustworthy data ...

Technical Lead (AWS Serverless & Microservices)

Location
Greater London, England, United Kingdom
through automation, testing practices and governance frameworks Promoting secure-by-design principles and embedding security best practice across teams and platforms Establishing and evolving observability capabilities including logging, metrics, tracing and alerting Championing metrics-led engineering to improve delivery performance, reliability and customer outcomes Providing technical guidance on complex engineering … Expertise in quality engineering, automated testing strategies and CI/CD practices Strong understanding of secure software development principles and security frameworks Experience implementing observability strategies using logging, metrics, tracing and alerting tools Proven ability to influence stakeholders and communicate complex technical concepts clearly Experience coaching, mentoring and developing high ...

Principal DevOps Engineer

Location
Greater London, England, United Kingdom
templates, shared steps, release promotion, rollback) Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable) Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards) Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning … Ansible (architecture, reusable components, secure operations) Strong deployment/release engineering experience with Octopus Deploy and GitHub (release governance, environment promotion, rollback) Monitoring/observability expertise with VictoriaMetrics and/or Prometheus (alerting strategy, metrics design, operational readiness) Production experience running Redis , RabbitMQ , Nginx (HA, tuning, troubleshooting) Strong understanding ...

Software Engineer, Enterprise

Hiring Organisation
Scale AI
Location
London, UK
Employment Type
Full-time
applications can operate at enterprise scale across hybrid and multi-cloud environments. Manage and evolve cloud infrastructure (AWS, Azure, or GCP), driving automation, observability, and security for large-scale AI deployments. Collaborate with ML and product teams to bring cutting-edge GenAI models into production through efficient APIs, model serving … experience with GenAI applications, model integration, or AI agent systems—understanding how to deploy, evaluate, and scale AI workloads in production. Strong understanding of observability, CI/CD, and security best practices for running services in enterprise or multi-tenant environments. Ability to balance rapid iteration with production-grade quality ...

Credit Principal Algo Data Lead

Hiring Organisation
UBS
Location
London, UK
Employment Type
Full-time
orchestration (Airflow, Prefect, or similar) and reliability patterns such as retries, backfills, and exactly-once semantics. Deep experience with Azure, including AKS, Helm, EventHub, observability tools (e.g. Log Analytics) and infrastructure-as-code (e.g. Terraform). Advanced experience operating Delta tables at scale, covering transaction semantics, schema evolution, partitioning … Marimo. Batch and streaming processing proficiency, including late data handling, watermarking, backfills, and explicit throughput vs. latency trade-offs. Production data reliability and observability mindset, covering lineage, data quality checks, SLAs, CI/CD for data pipelines, and close collaboration with quants, risk managers, and traders to deliver trustworthy data ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
create, maintain and debug reproducible multi-language CI pipelines, and optimize CI performance across large compute clusters. Build and maintain infrastructure observability, alerting, runbooks, and incident response workflows for Fractile's infrastructure. IaC TODO Scale and maintain Fractile's Bazel monorepo as we continue growing across Python, C++, Rust, SystemVerilog … high performance and extensibility, such as Bazel, Buck, Pants, Please, etc. Experience with infrastructure as code (Terraform, OpenTofu, or Pulumi) Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on‐call practice Strong proficiency ...

Senior Software Engineer: Agentic Development Enablement

Location
Greater London, England, United Kingdom
Claude Code and GitHub Copilot Design and implement practical guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Help define how controls should work consistently across local development environments and CI/CD pipelines Explore changes to the development environment … background as a software engineer Broad technical understanding across several of the following: developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end‐user application delivery Ability to work in ambiguous ...

Senior AI Engineer| London

Hiring Organisation
Infosys Technologies
Location
London, UK
Employment Type
Full-time
Agentic AI, classic ML and automation space Experience and good understanding of GenAI prompt engineering, RAG pipelines, Supervised/unsupervised ML and AI observability Experience in Enterprise-grade RAG-based solutions with LLMs (OpenAI, Hugging Face, LLaMA, etc.) and vector databases (Pinecone, Weaviate, FAISS). Experience in designing and scaling … Knowledge of AgentOps and OpenTelemetry Understanding of Network Security Concepts, Network Telemetry and Analytics Understanding of Cloud computing and Virtualization Exposure to APM/Observability tools (Dynatrace, AppDynamics, Datadog, Splunk etc) Exposure to onshore-offshore model working with professionals spread across the globePersonalBesides the professional qualifications, we respect and place ...

Head of Platform

Location
Greater London, England, United Kingdom
Experience with multi‐cloud environments (AWS and GCP). Experience with DataOps practices such as data contracts, data quality frameworks and lineage. Experience with observability tooling (metrics, logs, tracing) at scale Experience with FinOps and cloud cost optimisation. Experience using AI/LLM tooling to speed up engineering workflows. CAREER … ways for engineers to publish and use data. Build data quality, lineage and monitoring into the pipelines rather than bolting them on afterwards. Own observability, reliability and incident practice across the platform, including SLOs, alerting, runbooks and blameless post‐incident learning. Improve developer experience and measure it. Track outcomes such ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...

Senior Software Engineer, Full-Stack

Location
Greater London, England, United Kingdom
Write scripts and tooling to automate repetitive work, support data migrations, and keep the team moving quickly Instrument the systems you build with strong observability practices -logging, metrics, and tracing — to catch and resolve issues before they impact customers Benchmark and profile code to identify performance bottlenecks, and make evidence … ability to contribute there when needed Experience with, or solid understanding of, AWS Experience building or maintaining multi‐tenant, enterprise SaaS platforms Familiarity with observability tooling and practices (logging, metrics, tracing, alerting) Experience benchmarking and profiling code for performance Experience in the financial industry, particularly treasury, payments ...