1,051 to 1,075 of 5,047 Observability Jobs

Software Engineer

Hiring Organisation
IBM SIXworks Limited
Location
Taunton, Somerset, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Work on Technology That Protects What Matters At SiXworks , we build secure digital solutions that support Defence and National Security missions . Our teams work on complex problems where reliability, security, and speed of innovation ...

Senior Software Engineer - iCloud Platform - Observability

Location
Greater London, England, United Kingdom
services platform and infrastructure. You will be working on foundational systems that power iCloud services, including distributed data platforms, storage systems, and a unified observability infrastructure that standardizes monitoring and telemetry approaches across all iCloud services serving billions of customers.iCloud manages data and services at massive scale! Our unified observability … health of all iCloud services, providing comprehensive telemetry collection, real-time processing, and sophisticated analysis capabilities that span billions of active Apple customers. This observability ecosystem is purpose-built to deliver highly scalable and performant solutions, that maintain the highest standards of user privacy and security.We are a world-class ...

Solution Engineer (Pre-Sales)

Location
Greater London, England, United Kingdom
## Solution Engineer (Pre-Sales)London, UK · Full-time · Senior#### About The PositionCoralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs … metrics, trace and security events with features such as APM, RUM, SIEM, Kubernetes monitoring and more, all enhancing operational efficiency and reducing observability spend by up to 70%.Solution Architects in Coralogix are key in meeting our customers’ expectations and helping them utilize their observability and security data. ...

SRE

Location
Hove, England, United Kingdom
No. of Positions: 1 We are seeking an experienced Site Reliability Engineer (SRE) to drive the modernization of IT operations through the implementation of observability practices, automation, and reliability engineering principles. The role requires a strategic thinker with strong hands‐on expertise who can enhance system reliability, scalability, and operational … practices, automate operational workflows, and establish robust monitoring and incident management frameworks. Key Responsibilities Collaborate with engineering teams to modernize IT operations by improving observability, automation, and operational efficiency. Design and implement observability platforms to effectively monitor system health, performance, and reliability. Develop strategies for AI-driven alerting and proactive ...

Senior Platform Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£65,000
Senior Platform Engineer Deliver and support cloud platform engineering solutions across client environments. Design services with a reliability mindset, using SLIs, SLOs, and observability practices. Implement and maintain Infrastructure as Code using Terraform across environments. Support incident management, problem management, and continuous improvement of production platforms. Contribute to observability solutions … Infrastructure as Code across non-production and production environments. Understanding of SRE principles including SLIs, SLOs, error budgets, resilience, and reliability. Experience with observability and monitoring tools such as Dynatrace or similar. Experience supporting production platforms including incident and problem management. Exposure to AIOps practices and automation for proactive issue ...

Site Reliability Engineer

Location
Cardiff, Wales, United Kingdom
week in Cardiff office) About the Role This company provides managed AI operations for technology businesses. The company operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software. This is a chance to take real ownership in a Cardiff business on the front line … Datadog Advanced Partner (UK) and holds the accolade of being the world's first accredited MSP powered by Datadog. Datadog is a Nasdaq-listed observability platform. The company's technical focus is on observability and LLM observability specifically. Key Responsibilities Managed Service Delivery - Support customers across a mix of traditional ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...

Core Platform Developer

Location
Greater London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment.**Key Responsibilities*** Design and build shared backend services, frameworks … developer tooling that support internal application and service development.* Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities.* Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams.* Help ...

AI Native SW Eng

Location
United Kingdom
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … open-source models with fallback routing, token, cost, and latency management Implement LLMOps in production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions workshops, proofs of concept ...

Software Architect (Java or C#)

Location
United Kingdom
Engineering Partnership Work actively with Engineering teams during design, development, and production-readiness reviews. Advise and challenge teams on service architecture, fault tolerance, scalability, observability, deployment safety, and operational readiness, helping them to make pragmatic trade-offs. Support teams in diagnosing complex performance, latency, throughput, and resource-utilisation issues. Help … establish engineering standards and reusable patterns for reliable, maintainable services. Performance & Observability Lead investigations into performance bottlenecks across applications, infrastructure, databases, queues, networks, and third-party dependencies. Improve observability through metrics, logs, traces, dashboards, alerting, and service-level indicators. Help teams design meaningful alerts that identify user-impacting issues while ...

Software Engineers

Location
Douglas, Isle of Man, United Kingdom
integrations between enterprise systems and third-party platforms. Supporting and optimising cloud infrastructure and platform services within Azure environments. Driving operational excellence through monitoring, observability, automation and service reliability practices. Developing and maintaining automated testing frameworks and quality engineering standards. Supporting DevOps, CI/CD and infrastructure automation initiatives. Contributing … automation, quality engineering and automated testing frameworks. API development, application integration and distributed systems. DevOps tooling, CI/CD pipelines and automation practices. Monitoring, observability, incident management and operational resilience. Agile software delivery and modern engineering methodologies. Strong analytical, troubleshooting and problem-solving skills. Excellent communication and stakeholder management capabilities. ...

SC Cleared AWS DevOps & Platform Engineer (Remote)

Location
England, United Kingdom
cloud infrastructure in AWS, build automated pipelines with GitLab CI/CD and ArgoCD, and manage Kubernetes, Docker, Terraform and related tools while ensuring observability with Grafana/Prometheus. Strong hands-on AWS skills, container orchestration, IaC and CI/CD expertise are essential #J-18808-Ljbffr ...

AI / Machine Learning Engineer – Agentic LLM Systems (Contract)

Location
Greater London, England, United Kingdom
services and APIs to support AI-driven applications and workflowsBuild and run experiments to improve reliability, latency, cost, and success ratesContribute to evaluation frameworks, observability, and monitoring of LLM/agent performance What We’re Looking For: 4+ years’ experience in ML/AI engineering (LLMs, recommender systems, optimisation … codeComfortable debugging complex, distributed AI/ML systemsExperience running and analysing large-scale experiments and performance metrics (latency, accuracy, cost)Exposure to monitoring/observability tools for production systems Nice to Have: Experience with multi-agent systems or distributed AI architecturesExperience optimising LLM usage across multiple providers (cost/performance ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
pipelines Supporting the deployment of applications into production Working with Azure Kubernetes Service (AKS) and containerised applications Improving cloud networking, security, resilience and observability Working with Entra ID and identity/access management Partnering closely with Software Engineers to make deployments simpler, safer and more reliable Helping shape platform standards … Cloud networking Deploying and supporting applications in production Experience with Kubernetes/AKS would be particularly useful. Azure DevOps, Entra ID, cloud security, observability, cost optimisation or relevant Azure/Kubernetes certifications would all add value, but you don't need to tick every box. More important is that ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … call rota and provide on-call support for other SRE engineers.* Can write advanced automation scripts for incident response, including failovers and rollbacks.**Observability** * Has a deep technical understanding of observability techniques across the full stack and can bring clarity to complex incidents or performance issues.* Able to create templated ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
Role Profile:**We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a **Technical Lead SRE**, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. … projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design.Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.Design and evolve monitoring and alerting solutions that ...

Senior Platform Engineer (Azure)

Location
Greater London, England, United Kingdom
improve CI/CD pipelines and deployment processes Work closely with engineering, product and QA teams to support reliable delivery Improve platform reliability, security, observability and performance Help drive DevSecOps, SRE and automation best practices across the platform What they're looking for: Strong cloud-native Azure experience Production experience … haves... AZ-305 or CKA certifications Helm Azure DevOps pipelines Experience within a regulated environment such as financial services, insurance or healthcare Experience with observability, SRE and DevSecOps principles This role would suit a Senior Platform Engineer who still enjoys being hands on technically and wants real ownership over ...

Principal AI Platform Engineer

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
golden paths, and standard service templates to simplify service provisioning and operations. Contribute to cloud-native platform architecture, including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Integrate AI-assisted engineering workflows to support faster delivery, improved code quality, automation, and data-driven operational decisions. … Establish governance, security, and compliance guardrails through policy-as-code and auditable platform patterns. Improve platform reliability using SLOs, observability practices, resilience engineering, and insights from incidents. Collaborate with product, engineering, security, and architecture teams to align platform capabilities with business priorities and user needs. Drive efficiency and sustainability through ...

Enterprise Architect

Location
Greater London, England, United Kingdom
observable platforms. The role requires strong hands‐on architecture depth across application modernization, integration architecture, microservices, APIs, event streaming, container platforms, database modernization, DevSecOps, observability, resiliency, and governance. You will collaborate with business, application, data, infrastructure, security, DevOps, operations, and delivery teams to define transition roadmaps, manage architectural risks, establish … Application Architecture Microservices Architecture Event-Driven Architecture API & Integration Architecture Cloud-Agnostic Architecture Domain-Driven Design Strangler Pattern Migration DevSecOps CI/CD & IaC Observability DataBricks Data Governance & Lineage Cloud Portability Kubernetes Kafka .NET MSSQL/PostgreSQL Nice to Have Financial Services/Banking Domain experience Key Competencies Proven enterprise ...

Lead DevOps Engineer Software Engineering · Stallingborough HQ ·

Location
Grimsby, England, United Kingdom
Role Summary A hands‐on Lead DevOps Engineer to lead our DevOps team and own the day‐to‐day evolution of myenergi's cloud, observability and data platforms. The team runs the platform around the clock, from hosting and CI/CD to observability, the data platform and security controls. … GitHub Actions) so teams ship safely and often; lead cost reviews. Operate Postgres, Redshift, Redis and OpenSearch with clear SLAs and backups. Own the observability platform (ClickHouse, OpenTelemetry): logs, metrics, traces, synthetics, SLOs and actionable alerting. Own the data platform: Glue pipelines, Apache Iceberg lakehouse, Redshift and CDC; serve ...

Principal AI Platform Engineer

Location
Greater London, England, United Kingdom
golden paths, and standard service templates to simplify service provisioning and operations. Contribute to cloud-native platform architecture, including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Integrate AI-assisted engineering workflows to support faster delivery, improved code quality, automation, and data-driven operational decisions. … Establish governance, security, and compliance guardrails through policy-as-code and auditable platform patterns. Improve platform reliability using SLOs, observability practices, resilience engineering, and insights from incidents. Collaborate with product, engineering, security, and architecture teams to align platform capabilities with business priorities and user needs. Drive efficiency and sustainability through ...

Senior SRE (AWS)

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
Strong hands-on experience with both AWS, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior SRE (AWS)

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent
Salary
£75,000
Strong hands-on experience with both AWS, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Cloud Network & Security Engineer , Vice President

Hiring Organisation
State Street Bank
Location
Edinburgh, Midlothian, United Kingdom
Salary
£ 70 K
platform engineering capabilities that improve consistency, reliability, and speed of delivery. Develop AI-enabled operational capabilities that support network investigation, incident response, change validation, observability, and workflow automation. Define architecture standards, engineering patterns, and technology roadmaps for enterprise-wide adoption. Partner with cloud, security, infrastructure, application, and operations teams globally. … platforms. Exposure to platform engineering, GitOps, event-driven automation, or self-service infrastructure models. Interest or experience in AI-enabled operations, AI agent workflows, observability platforms, knowledge systems, or intelligent automation. Experience working in a regulated, global enterprise environment. What We Value Engineering ownership and accountability. A builder’s mindset ...

Senior Cloud Network & Security Engineer , Vice President

Location
Greater London, England, United Kingdom
platform engineering capabilities that improve consistency, reliability, and speed of delivery. Develop AI-enabled operational capabilities that support network investigation, incident response, change validation, observability, and workflow automation. Define architecture standards, engineering patterns, and technology roadmaps for enterprise-wide adoption. Partner with cloud, security, infrastructure, application, and operations teams globally. … platforms. Exposure to platform engineering, GitOps, event-driven automation, or self-service infrastructure models. Interest or experience in AI-enabled operations, AI agent workflows, observability platforms, knowledge systems, or intelligent automation. Experience working in a regulated, global enterprise environment. What We Value Engineering ownership and accountability. A builder’s mindset ...