1,801 to 1,825 of 4,038 Observability Jobs

Senior Data Engineer

Location
Maidenhead, England, United Kingdom
data solutions are reliable, scalable, performant, secure, and production‐ready Monitor, troubleshoot, and continuously improve pipeline performance, data quality, and platform stability Drive automation, observability, and supportability across data, analytics, and AI/ML solutions Our Ideal Candidate Strong data engineering experience with hands‐on delivery of scalable data pipelines … productionization, LLM‐based applications, or agentic AI patterns will be an added advantage Experience with DevOps and DataOps practices, including CI/CD, monitoring, observability, and incident support Maersk is committed to a diverse and inclusive workplace, and we embrace different styles of thinking. Maersk is an equal opportunities employer ...

Data Architect

Location
Greater London, England, United Kingdom
duties, isolation, and alignment to standards such as ISO 27001. Provide architectural direction for cloud data platforms, infrastructure as code, CI/CD, and observability, ensuring designs are cost‐aware and operable. Define the data lineage and controls that support responsible production AI — monitoring, confidence scoring, and human … approaches, and how each changes data architecture requirements, including context window economics and prompt caching. Familiarity with infrastructure as code, CI/CD, and observability for data platforms. Awareness of MCP and tool‐orchestration concepts and their implications for composable, data‐driven AI systems. Fluent use of AI‐assisted development ...

Client Service Delivery, Sr Manager

Location
Birmingham, England, United Kingdom
Service Delivery Management Own full lifecycle service delivery across infrastructure and cloud environments, ensuring alignment to SLAs, KPIs, scope, and cost. Leverage AIOps and observability tools (e.g.Dynatrace, Datadog, New Relic, Elastic) to proactivelymonitorservice health and performance. Utilisepredictive alerting and anomaly detection to prevent incidents andoptimisedelivery priorities. Coordinate across internal teams … infrastructure and cloud environments Strong understanding of IT Managed Services frameworks Hands-on experience with AIOps tools such as Dynatrace and ServiceNow Familiarity with observability tools (e.g.Datadog, New Relic, Elastic) Knowledge of event analytics tools such as Splunk IT Service Intelligence andMoogsoft Experience in stakeholder and client management Financial management ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
access patterns. Support hybrid connectivity models (site‐to‐site VPN, client VPN, ExpressRoute, Direct Connect, SD‐WAN). Monitor network performance and reliability using observability and telemetry tools; proactively address capacity and performance issues. Troubleshoot complex network and cross‐domain infrastructure issues spanning network, compute, and cloud layers. Develop … constructs (VPC/VNet design, routing, security groups, load balancers). Experience with SD‐WAN architectures and implementations. Familiarity with network monitoring, logging, and observability tools (e.g., SNMP, NetFlow, Syslog, modern NPM tools). Working knowledge of compute platforms and operating systems (Windows, Linux, virtualization such as VMware/Hyper ...

AI Platform Engineer

Location
Greater London, England, United Kingdom
patterns that support the safe and scalable adoption of AI-assisted development.Improve developer experience through streamlined workflows, tooling integration and self-service capabilities. Develop observability and measurement capabilities to provide insights into engineering productivity, quality and platform adoption. Collaborate with Technical Leads and AI Software Engineers to identify recurring engineering … cost optimisation, including token monitoring, caching strategies and model selection considerations. Knowledge of approaches for managing and reducing token consumption costs.Deep understanding of observability, automated testing and software delivery tooling. Knowledge of platform security, governance and operational controls. Strong programming, automation and problem-solving capabilities. Passionate about improving developer experience ...

Principal Software Engineer - Squad Lead Engineer

Location
Greater London, England, United Kingdom
complete complex bug fixes and performance improvements* Define and uphold Definition of Ready/Done including code quality, automated test coverage, security checks, and observability* Establish/maintain CI/CD pipelines, quality gates, and sensible branching/release strategies* Drive a pragmatic quality strategy: test pyramid balance, contract tests … Windows* Experience with relational and non-relational data stores, performance tuning, and data modelling* Knowledge of CI/CD platforms, containers, cloud technologies, observability, and monitoring practices* Understanding of secure coding, performance optimisation, reliability engineering, and incident response **Work in a Way That Works for You**We promote a healthy ...

Senior Lead Software Engineer - Mobile Engineering

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
drive measurable improvements in stability and release confidence. Own mobile build/release and operational maturity: CI/CD pipelines, distribution, feature flags, observability, crash/performance monitoring, and incident response. Mentor and coach engineers; support team growth through feedback, technical guidance, and strong engineering culture. Communicate clearly with senior … patterns, accessibility standards, component libraries).Experience with CI/CD for mobile (e.g., build automation, signing, distribution, feature flags, release trains).Experience with observability and production support practices: crash analytics, performance monitoring, logging, alerting, and operational readiness. Experience leading multiple engineers/teams (people leadership or strong matrix leadership), including ...

Infrastructure Lead

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
other teams, with documentation and examples. In addition, you should have experience working with Kubernetes in production environments at scale, and be familiar with observability tools such as Prometheus and Grafana. Strong Linux Server Administration and Configuration Management skills, as well as some networking experience, are also required. The ideal ...

Senior Platform Manager (Server Infrastructure)

Location
Gloucester, England, United Kingdom
current and future organisational needs. This is an exciting opportunity to play a key role in modernising infrastructure services and driving adoption of automation, observability, and platform reliability best practices. What you would be doing You will be responsible for the operational management, maintenance, and continual improvement of enterprise server … looking for Proven experience in managing enterprise server infrastructure in a complex environment. Experience in Leading Technical Teams. Knowledge of infrastructure monitoring, alerting, and observability tooling. Strong troubleshooting and problem-solving skills with the ability to manage competing priorities effectively. Excellent communication and stakeholder engagement skills, with the ability ...

Principal Engineer - Integration Services

Location
Greater London, England, United Kingdom
duplication and improve interoperability. Providing technical leadership on significant integration decisions spanning multiple platforms, teams and domains. Ensuring integration approaches consider security, resilience, scalability, observability, operability and long‐term maintainability. Working with Architecture and senior engineering leaders to align integration direction with wider FT technology strategy. Establishing a clear view … progress, risks, trade‐offs and technical constraints. Operational Excellence & Reliability Owning the operational performance of shared integration services. Establishing appropriate approaches to monitoring, observability, incident management, problem management and operational readiness. Ensuring critical integrations have clear ownership, appropriate resilience and effective recovery mechanisms. Leading the response to significant incidents affecting ...

DevOps Engineer - Active SC, AWS

Hiring Organisation
Investigo Change Solutions
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP 535 Daily
development best practices Experience with monitoring, logging and alerting tools including Elastic and Kibana, with a focus on improving platform resilience, security and observability Proven capability supporting live production environments, leading third-line incident investigation and resolution, and automating environment provisioning and operational processes Experience delivering DevOps solutions within secure ...

Applied AI Engineer

Hiring Organisation
Nufin
Location
London, UK
Employment Type
Full-time
business impact. Create representative test datasets, evaluation criteria, regression tests, and human-review processes. Measure and improve accuracy, latency, cost, and user experience. Establish observability and feedback loops that make agent behavior understandable and continuously improvable. Develop the AI Application ArchitectureApply context-engineering techniques such as RAG, MCP and knowledge … LlamaIndex, or comparable frameworksContext engineering: RAG, MCP, knowledge graphs, tool use, memory, and retrieval systemsAI evaluation: offline and online evaluations, test datasets, regression testing, observability, and human reviewLanguage models: Gemini, OpenAI, Anthropic, Llama, Mistral, or similarBackend engineering: Python or Java, REST APIs, Kafka, microservices, and distributed systemsData systems: SQL, PostgreSQL ...

Principal Software Engineer - Customer Platforms

Location
Greater London, England, United Kingdom
customers to a desired outcome, without prescribing it Authoritative skills at cloud computing (network, security, serverless, Kubernetes etc) and automation Experience with implementation of Observability and Reliability using market technologies (e.g.: New Relic) Good experience with Performance Engineering (load testing, derivations, tuning, core web vitals, page speed etc.) Expertise … organisation(s) Tech Stack M&S uses a variety of technologies including; Java, Spring, SpringBOOT, Micronaut React, Next.js, Typescript, Angular Azure Cloud, Kubernetes, Dynatrace (observability) SQL Server, MongoDB Ignite, Redis What’s In It For You Working at M&S means being part of something bigger - helping to deliver quality ...

Senior Site Reliability Engineer

Hiring Organisation
Pathfinder Business Solutions Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
experience. An Site Reliability Engineering (SRE) background with responsibility for running live services. Experience with microservices, containers and API management. Full stack monitoring and observability experience. Experience with DevOps, pipelines, releases and automation. Experience in a large organisation with established processes. Confidence working with stakeholders, leading improvements and supporting colleagues. ...

Staff Engineer, iOS

Location
United Kingdom
ready mobile systems. Raise the bar for reliability and execution: Shape how mobile work is designed, reviewed, tested, shipped, and operated, including the automation, observability, rollout practices, and AI‐enabled validation needed to move quickly with confidence. Shape mobile system contracts: Bring enough Android and backend context to align teams … driven testing systems to improve engineering velocity, quality, and regression confidence. Strong reliability and release‐quality mindset, including experience with testing strategy, automation, observability, CI/CD, and production mobile quality practices. Track record of leading significant technical initiatives, setting engineering standards, mentoring engineers, and raising the technical bar across ...

Director of Software Engineering

Hiring Organisation
Spire Global
Location
Glasgow, UK
Employment Type
Full-time
infrastructureStay hands-on: review code, prototype solutions, and get into the details when it mattersEstablish engineering standards across code quality, system design, testing, and observability, and hold the team to themBe the person engineers come to when the problem is genuinely hardTeam Building & Culture Recruit, develop, and retain a team … monitoring problemsExperience writing performance software in RustBackground in space systems, aerospace, or highly constrained real-time environmentsExperience building data lakes, telemetry platforms, or observability infrastructure at scaleA history of leading teams through technical transformations and not just maintaining the status quoSpire operates a hybrid work model, and this position will ...

Senior AI Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £700 to £900 per day
predictable failure behaviour. Deploy AI systems across cloud and on-premises environments, with an understanding of the constraints associated with each. Build evaluation and observability capabilities to measure model performance, agent behaviour and system reliability. Take end-to-end ownership of technical workstreams, from architecture and implementation through to deployment. … would be advantageous: Multimodal AI and reasoning. Edge or offline AI deployments. Kubernetes, particularly EKS or OpenShift. MLOps, including model evaluation, monitoring and reproducibility. Observability for agentic AI systems, including model performance, agent behaviour and drift. Agent orchestration and inter-agent communication protocols such as A2A. Model Context Protocol ...

Staff Software Engineer

Location
Greater London, England, United Kingdom
from problem framing and design review through to production rollout, migration and decommissioning of what came before Drive operational excellence by improving observability, alerting, capacity planning, load and chaos testing, and incident response, and by closing the loop on the root causes you find Raise the engineering bar across teams … development (we use AWS & Azure), containers, orchestration and infrastructure as code Experience owning production systems on call, with a strong sense of ownership for observability, monitoring and incident response Experience leading complex migrations or re-architectures of live, business-critical systems with no loss of availability or data integrity Excellent ...

Staff Software Engineer - AI Agents (Satori)

Location
Belfast City District, Northern Ireland, United Kingdom
ready" means at Proofpoint: the eval suites, safety checks, telemetry, and AuthX that gate every rollout. The primitives you set — from prompt management to observability — will shape how mission‐critical agentic software ships across the company. You will partner with product tech leads to turn one‐off integrations into reusable … similar) and keep up with the field. Prior experience taking agentic or LLM‐based systems to production — not just prototypes — including the eval, observability, and rollback story. Deep experience with Python; TypeScript/Node a plus. Hands‐on work with agent frameworks and eval tooling. Track record of debugging distributed ...

Principal Software Engineer - Squad Lead Engineer

Location
Greater London, England, United Kingdom
complete complex bug fixes and performance improvements Define and uphold Definition of Ready/Done including code quality, automated test coverage, security checks, and observability Establish/maintain CI/CD pipelines, quality gates, and sensible branching/release strategies Drive a pragmatic quality strategy: test pyramid balance, contract tests … Windows Experience with relational and non‐relational data stores, performance tuning, and data modelling Knowledge of CI/CD platforms, containers, cloud technologies, observability, and monitoring practices Understanding of secure coding, performance optimisation, reliability engineering, and incident response Work in a Way That Works for You We promote a healthy ...

Principal Engineer

Location
Milton Keynes, England, United Kingdom
Azure services, containers, application services, messaging, API management, relational and NoSQL data platforms. Improve continuous integration, continuous delivery, automated quality controls, infrastructure management, observability and release safety. Promote cost‐effective architecture and operational practices, including capacity planning, service ownership and FinOps awareness. Lead the adoption of reusable global components … queues, API Management, Azure SQL, Cosmos DB, containers and monitoring services. Strong understanding of CI/CD, automated testing, code quality controls, observability, incident diagnosis, reliability engineering and DevSecOps practices. Knowledge of application security, identity and access management, OWASP guidance, threat modelling, privacy by design and relevant compliance requirements. Experience ...

Senior Cloud Platform Engineer

Location
Milton Keynes, England, United Kingdom
available cloud platform solutions Leading the implementation of cloud infrastructure, automation and Infrastructure as Code Enhancing CI/CD pipelines and deployment capabilities Driving observability, monitoring and platform security best practices Supporting cloud transformation and platform modernisation initiatives Providing technical leadership during major incidents and service restoration activities Mentoring engineers ...

ML Data & Platform Engineer

Location
United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

ML Data & Platform Engineer

Location
Cambridge, England, United Kingdom
models efficiently and reliably in production Optimising infrastructure for both iteration speed and production reliability, including GPU utilisation, job scheduling, and training efficiency Implementing observability (monitoring, logging, alerting) across data pipelines and ML systems to catch issues early and keep things running smoothly Troubleshooting complex issues across distributed systems, spanning … lifecycle, from data through to model training, evaluation, and serving Experience with data quality practices (validation, cleaning, normalisation) and/or production-grade observability Ability to design resilient, scalable architectures, and comfort operating and troubleshooting distributed systems MLOps experience, for example model serving, experiment tracking, GPU/distributed training optimisation ...

Cloud Advisory Architecture Associate Manager

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
where GenAI and Agentic play a role. Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps. Manage & Mentor: Lead teams of architects and engineers, providing technical coaching, career counselling, performance management, and coaching ...