2,426 to 2,450 of 5,534 Observability Jobs

AI Engineer

Location
Greater London, England, United Kingdom
full-stack AI engineering role, with applied AI product delivery at its core. You will build AI workflows, tools and integrations, evaluation and observability capabilities, together with the APIs, services and user interfaces needed to ship them reliably. This is not a research-only role. You will apply agreed enterprise … feedback loops. Manage prompts, model configuration, tool schemas and routing as tested product assets, including fallbacks and cost/latency trade‐offs. Implement AI observability through traces, logs, metrics and evaluation results, so quality, reliability, failure modes, latency and cost are visible and can be improved. Build clear APIs ...

Agent Engineer

Location
Greater London, England, United Kingdom
concept through to production. Take part in technical design reviews, planning sessions, and code reviews to continuously improve system quality. Contribute to infrastructure and observability practices alongside the Engineering team — you won't own this alone, but you'll be expected to care about how your services run in production. … design. Nice to Have Experience with Node.js frameworks like NestJS or Express. Hands-on experience with Terraform, or infrastructure-as-code tooling. Experience with observability platforms like Datadog (metrics, tracing, alerting). Exposure to DynamoDB or other NoSQL databases at scale. Experience with distributed or event-driven architectures ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
vulnerability management and secure provisioning. Edge & Cloud Integration Integrate devices with cloud IoT platforms and backend services. Define and improve telemetry, health monitoring, and observability for deployed devices. Support reliable operation of connected devices in the field, including debugging fleet issues and improving resilience. Compliance & Testing Support testing and validation …/CD workflows for embedded software. Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers, or observability platforms. Strong problem-solving skills and the ability to work across hardware and software boundaries. Nice to Have Experience with low-power wireless technologies such ...

Backend Engineers (Ruby)

Location
Greater London, England, United Kingdom
technical insight into upcoming work and helping pull the team together to ship it Deliver your work using agile methodologies and tools like tests, observability, A/B tests, and feature flags Mentor colleagues to help them grow as engineers, and actively support their development Contribute to cross-cutting concerns … Rails, building data models, APIs, and business logic services in a production monolith Comfortable delivering with agile methodologies and practices like automated testing, observability, A/B testing, and feature flags A collaborative approach — you work well with product, design, and analytics partners, and enjoy shaping work beyond just ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
across firewalls, NAT, VPN, security policies, and multi-tenant segmentation. Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure. Reliability, Observability & Operations Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues. Establish measurable SLOs and operational standards for network services. … technical direction for network observability, telemetry, monitoring, and alerting. Ensure clear visibility into fabric health, traffic patterns, performance, and capacity. Develop runbooks, automation, and engineering improvements that systematically reduce operational toil. Act as a senior 3rd/4th line escalation point for complex networking issues. Network Data & Configuration Management Ensure ...

GenAI Engineer

Location
Greater London, England, United Kingdom
indexing, retrieval policies, grounding, guardrails) and agent frameworks. Take basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost tuning. Participate in on‐call for your area and drive root‐cause analysis with crisp follow‐ups. 15% Collaborate Pair with back … series analysis (forecasting, change‐point, drift). Cloud & ops: Basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost control. Communication: You explain results clearly, align stakeholders, and write crisp docs. Bonus points DevOps wizardry; GPU/accelerator experience. Multimodal pipelines (text ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Data Platform DevOps Analyst

Location
Reading, England, United Kingdom
journey supporting the migration from Teradata to Databricks, validating production readiness, improving monitoring and automation capabilities and helping embed modern DataOps practices across deployment, observability and support processes. What You’ll Bring Here at Primark, we want everyone to feel valued – so please bring your authentic self to work … platforms and DataOps practices with exposure to Azure DevOps, Databricks, dbt Cloud, Azure data services, SQL, Python, ETL/ELT processes, job scheduling, monitoring, observability and deployment automation. Experience supporting production environments and service operations including monitoring, alerting, incident management, ticketing systems, release management, environment management, change governance and enterprise ...

Senior Cloud Infrastructure and Network Operations Solutions Architect

Location
Reading, England, United Kingdom
network OS internalsand host networkingand OS‐level security, together with an understandingof the storage and compute traffic patternsof HPC/AI clusters. Automation, GitOps & Observability: Proficiency in Pythonand Bash scripting, configuration management and Infrastructure‐as‐Code tools (e.g. Ansible, Terraform), GitOps‐based network configuration and firmware/upgrade management … large fleets, and observability stacks (Grafana, Loki, Prometheus) applied to fabric telemetry. Fleet Reliability& Customer Engagement: Demonstrated abilityto measure and improve MTBI and goodput on large GPU clusters—fault detection, drain and remediation workflows, SLO/error‐budget definition and post‐incident review–combining strong consultative background leading architectural reviews ...

Technical Support Engineer

Location
Paignton, England, United Kingdom
major escalations and transformation initiatives. Lead or contribute to the design and deployment of scalable support workflows, including automated log parsing, triage tooling, and observability pipelines. Identify trends in support data to drive systemic product or process improvements. Collaborate cross-functionally to evolve support readiness standards for new product releases … lead technical investigations, influence engineering teams, and manage high-pressure escalations independently. Experience with support and engineering tools such as Salesforce, Jira, Git, and observability stacks. Desirable Skills & Knowledge Knowledge of GNSS simulation systems, timing/sync architectures, or satellite navigation protocols. Familiarity with DevOps, CI/CD, or site ...

Senior Developer Experience Security Engineer

Location
Greater London, England, United Kingdom
experience across Motorway.. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster … experience across Motorway.. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Location
Greater London, England, United Kingdom
with product, marketing, and partners to translate requirements into well-designed technical solutions Improve engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practices Ensure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controls Required Qualifications, Capabilities … with campaign management and attribution systems Background in AI-enabled content generation, experimentation, or workflow automation Skills in test automation, CI/CD, and observability tools Experience optimizing performance and scalability Familiarity with consent management and data minimization practices Ability to drive innovation in personalization and measurement Experience working ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
service outcomes. Key responsibilities You will lead the development and maturity of Made Tech’s operational capabilities incident, problem, and change management; monitoring and observability; automation and AIOps; governance; operational playbooks; runbooks; service health metrics; and 24/7/365 support patterns. You will ensure our teams have … ITIL practices blended with modern DevOps, SRE, Agile and platform‐engineering approaches. Broad technical awareness across cloud platforms, application architectures, data platforms, networks, observability tooling, security‐by‐design, and automation. Ability to create and evolve operational standards, playbooks, governance models, templates, and frameworks that drive consistency, stability, and efficiency. Skilled ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Insight teams to maintain reliable, accurate and meaningful data models as our products evolve Establish and promote good engineering practices across data pipelines, modelling, observability, performance and reliability Ensure data quality and integrity throughout our pipelines, identifying and addressing issues before they impact downstream users Provide technical guidance to engineers … data transformation skills, particularly using technologies such as Python, Spark and SQL Experience working with cloud-based data platforms Experience designing systems for reliability, observability, performance and data quality Strong understanding of data modelling and the principles behind well-structured analytical data Experience making technical design decisions and taking ownership ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
automationCollaborate with product, marketing, and partners to translate requirements into well-designed technical solutionsImprove engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practicesEnsure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controlsRequired Qualifications, Capabilities, and Skills … Skills: Experience with campaign management and attribution systemsBackground in AI-enabled content generation, experimentation, or workflow automationSkills in test automation, CI/CD, and observability toolsExperience optimizing performance and scalabilityFamiliarity with consent management and data minimization practicesAbility to drive innovation in personalization and measurementExperience working in cross-functional teamsJ.P. Morgan ...

Senior Data Engineer

Location
Maidenhead, England, United Kingdom
data solutions are reliable, scalable, performant, secure, and production‐ready Monitor, troubleshoot, and continuously improve pipeline performance, data quality, and platform stability Drive automation, observability, and supportability across data, analytics, and AI/ML solutions Our Ideal Candidate Strong data engineering experience with hands‐on delivery of scalable data pipelines … productionization, LLM‐based applications, or agentic AI patterns will be an added advantage Experience with DevOps and DataOps practices, including CI/CD, monitoring, observability, and incident support Maersk is committed to a diverse and inclusive workplace, and we embrace different styles of thinking. Maersk is an equal opportunities employer ...

Data Architect

Location
Greater London, England, United Kingdom
duties, isolation, and alignment to standards such as ISO 27001. Provide architectural direction for cloud data platforms, infrastructure as code, CI/CD, and observability, ensuring designs are cost‐aware and operable. Define the data lineage and controls that support responsible production AI — monitoring, confidence scoring, and human … approaches, and how each changes data architecture requirements, including context window economics and prompt caching. Familiarity with infrastructure as code, CI/CD, and observability for data platforms. Awareness of MCP and tool‐orchestration concepts and their implications for composable, data‐driven AI systems. Fluent use of AI‐assisted development ...

Client Service Delivery, Sr Manager

Location
Birmingham, England, United Kingdom
Service Delivery Management Own full lifecycle service delivery across infrastructure and cloud environments, ensuring alignment to SLAs, KPIs, scope, and cost. Leverage AIOps and observability tools (e.g.Dynatrace, Datadog, New Relic, Elastic) to proactivelymonitorservice health and performance. Utilisepredictive alerting and anomaly detection to prevent incidents andoptimisedelivery priorities. Coordinate across internal teams … infrastructure and cloud environments Strong understanding of IT Managed Services frameworks Hands-on experience with AIOps tools such as Dynatrace and ServiceNow Familiarity with observability tools (e.g.Datadog, New Relic, Elastic) Knowledge of event analytics tools such as Splunk IT Service Intelligence andMoogsoft Experience in stakeholder and client management Financial management ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
access patterns. Support hybrid connectivity models (site‐to‐site VPN, client VPN, ExpressRoute, Direct Connect, SD‐WAN). Monitor network performance and reliability using observability and telemetry tools; proactively address capacity and performance issues. Troubleshoot complex network and cross‐domain infrastructure issues spanning network, compute, and cloud layers. Develop … constructs (VPC/VNet design, routing, security groups, load balancers). Experience with SD‐WAN architectures and implementations. Familiarity with network monitoring, logging, and observability tools (e.g., SNMP, NetFlow, Syslog, modern NPM tools). Working knowledge of compute platforms and operating systems (Windows, Linux, virtualization such as VMware/Hyper ...

Principal Software Engineer - Squad Lead Engineer

Location
Greater London, England, United Kingdom
complete complex bug fixes and performance improvements* Define and uphold Definition of Ready/Done including code quality, automated test coverage, security checks, and observability* Establish/maintain CI/CD pipelines, quality gates, and sensible branching/release strategies* Drive a pragmatic quality strategy: test pyramid balance, contract tests … Windows* Experience with relational and non-relational data stores, performance tuning, and data modelling* Knowledge of CI/CD platforms, containers, cloud technologies, observability, and monitoring practices* Understanding of secure coding, performance optimisation, reliability engineering, and incident response **Work in a Way That Works for You**We promote a healthy ...

AI Platform Engineer

Location
Greater London, England, United Kingdom
patterns that support the safe and scalable adoption of AI-assisted development.Improve developer experience through streamlined workflows, tooling integration and self-service capabilities. Develop observability and measurement capabilities to provide insights into engineering productivity, quality and platform adoption. Collaborate with Technical Leads and AI Software Engineers to identify recurring engineering … cost optimisation, including token monitoring, caching strategies and model selection considerations. Knowledge of approaches for managing and reducing token consumption costs.Deep understanding of observability, automated testing and software delivery tooling. Knowledge of platform security, governance and operational controls. Strong programming, automation and problem-solving capabilities. Passionate about improving developer experience ...

Senior Lead Software Engineer - Mobile Engineering

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
drive measurable improvements in stability and release confidence. Own mobile build/release and operational maturity: CI/CD pipelines, distribution, feature flags, observability, crash/performance monitoring, and incident response. Mentor and coach engineers; support team growth through feedback, technical guidance, and strong engineering culture. Communicate clearly with senior … patterns, accessibility standards, component libraries).Experience with CI/CD for mobile (e.g., build automation, signing, distribution, feature flags, release trains).Experience with observability and production support practices: crash analytics, performance monitoring, logging, alerting, and operational readiness. Experience leading multiple engineers/teams (people leadership or strong matrix leadership), including ...

Infrastructure Lead

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
other teams, with documentation and examples. In addition, you should have experience working with Kubernetes in production environments at scale, and be familiar with observability tools such as Prometheus and Grafana. Strong Linux Server Administration and Configuration Management skills, as well as some networking experience, are also required. The ideal ...

Senior Platform Manager (Server Infrastructure)

Location
Gloucester, England, United Kingdom
current and future organisational needs. This is an exciting opportunity to play a key role in modernising infrastructure services and driving adoption of automation, observability, and platform reliability best practices. What you would be doing You will be responsible for the operational management, maintenance, and continual improvement of enterprise server … looking for Proven experience in managing enterprise server infrastructure in a complex environment. Experience in Leading Technical Teams. Knowledge of infrastructure monitoring, alerting, and observability tooling. Strong troubleshooting and problem-solving skills with the ability to manage competing priorities effectively. Excellent communication and stakeholder engagement skills, with the ability ...

Principal Engineer - Integration Services

Location
Greater London, England, United Kingdom
duplication and improve interoperability. Providing technical leadership on significant integration decisions spanning multiple platforms, teams and domains. Ensuring integration approaches consider security, resilience, scalability, observability, operability and long‐term maintainability. Working with Architecture and senior engineering leaders to align integration direction with wider FT technology strategy. Establishing a clear view … progress, risks, trade‐offs and technical constraints. Operational Excellence & Reliability Owning the operational performance of shared integration services. Establishing appropriate approaches to monitoring, observability, incident management, problem management and operational readiness. Ensuring critical integrations have clear ownership, appropriate resilience and effective recovery mechanisms. Leading the response to significant incidents affecting ...