1,176 to 1,200 of 5,621 Observability Jobs

Software Engineer (Backend)

Location
Greater London, England, United Kingdom
Software Engineer (backend) Location: London or New York - in office 4 days a week About Lorum Global payments are not broken. Incentives are. Clearing has been deprioritized inside balance sheet driven institutions whose models rely ...

Senior Software Engineer, Git Systems

Location
United Kingdom
About GitHub GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 ...

Solution Architect

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
The team you'll be working with: We are looking for a Solution Architect to help our clients design and deliver modern, scalable and secure digital solutions, primarily on Microsoft technologies. You will combine architecture ...

AI and Cloud Director - Technology Consulting (AI, data and analytics)

Hiring Organisation
Business Integration Partners
Location
London, UK
Employment Type
Full-time
AI and Cloud DirectorLocation: London/HybridBusiness: BIP UKPractice: xTech - AI, Data and CloudReporting to: UK xTech LeadershipEmployment type: Full timeAbout BIP UKFounded in 2003, BIP is an international consulting firm with more than 6 ...

Senior Data Engineer

Hiring Organisation
McKesson
Location
London, UK
Employment Type
Full-time
Role OverviewThe Senior Data Engineer is the technical owner of the ClarusONE data platform. This is a hands-on engineering role with real ownership: alongside building and delivering data pipelines, you will raise the engineering ...

Full Stack Engineer (AI Startup)

Hiring Organisation
Trismik
Location
United Kingdom, UK
Why join us? At Trismik, we’re a team of tech geeks from Cambridge University, Salesforce, and Amazon. With over 38 years of research behind us, we’re not just building tools - we’re sparking ...

Cloud Architect (DV Security Cleared)

Location
Greater London, England, United Kingdom
The Cloud Architect is responsible for analysing existing on-premise systems and designing, planning, and supporting their migration into a secure, scalable, and cloud-native environment. The role focuses on modernising architectures, leveraging managed services ...

AI & ML Engineer

Location
Greater London, England, United Kingdom
About the role The AI & ML Engineering team accelerates the adoption of AI across the business, championing innovation while ensuring our machine learning solutions are robust, scalable, and cost-efficient. We enable teams to solve ...

Software Engineer

Hiring Organisation
IBM SIXworks Limited
Location
Taunton, Somerset, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Work on Technology That Protects What Matters At SiXworks , we build secure digital solutions that support Defence and National Security missions . Our teams work on complex problems where reliability, security, and speed of innovation ...

Senior Software Engineer - iCloud Platform - Observability

Location
Greater London, England, United Kingdom
services platform and infrastructure. You will be working on foundational systems that power iCloud services, including distributed data platforms, storage systems, and a unified observability infrastructure that standardizes monitoring and telemetry approaches across all iCloud services serving billions of customers.iCloud manages data and services at massive scale! Our unified observability … health of all iCloud services, providing comprehensive telemetry collection, real-time processing, and sophisticated analysis capabilities that span billions of active Apple customers. This observability ecosystem is purpose-built to deliver highly scalable and performant solutions, that maintain the highest standards of user privacy and security.We are a world-class ...

Solution Engineer

Location
Greater London, England, United Kingdom
## Solution EngineerLondon, UK · Full-time · Senior#### About The PositionCoralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs, metrics, trace … security events with features such as APM, RUM, SIEM, Kubernetes monitoring and more, all enhancing operational efficiency and reducing observability spend by up to 70%.Solution Architects in Coralogix are key in meeting our customers’ expectations and helping them utilize their observability and security data. We are looking for hard ...

SRE

Location
Hove, England, United Kingdom
No. of Positions: 1 We are seeking an experienced Site Reliability Engineer (SRE) to drive the modernization of IT operations through the implementation of observability practices, automation, and reliability engineering principles. The role requires a strategic thinker with strong hands‐on expertise who can enhance system reliability, scalability, and operational … practices, automate operational workflows, and establish robust monitoring and incident management frameworks. Key Responsibilities Collaborate with engineering teams to modernize IT operations by improving observability, automation, and operational efficiency. Design and implement observability platforms to effectively monitor system health, performance, and reliability. Develop strategies for AI-driven alerting and proactive ...

Senior Platform Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£65,000
Senior Platform Engineer Deliver and support cloud platform engineering solutions across client environments. Design services with a reliability mindset, using SLIs, SLOs, and observability practices. Implement and maintain Infrastructure as Code using Terraform across environments. Support incident management, problem management, and continuous improvement of production platforms. Contribute to observability solutions … Infrastructure as Code across non-production and production environments. Understanding of SRE principles including SLIs, SLOs, error budgets, resilience, and reliability. Experience with observability and monitoring tools such as Dynatrace or similar. Experience supporting production platforms including incident and problem management. Exposure to AIOps practices and automation for proactive issue ...

Site Reliability Engineer

Location
Cardiff, Wales, United Kingdom
week in Cardiff office) About the Role This company provides managed AI operations for technology businesses. The company operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software. This is a chance to take real ownership in a Cardiff business on the front line … Datadog Advanced Partner (UK) and holds the accolade of being the world's first accredited MSP powered by Datadog. Datadog is a Nasdaq-listed observability platform. The company's technical focus is on observability and LLM observability specifically. Key Responsibilities Managed Service Delivery - Support customers across a mix of traditional ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...

AI Native SW Eng

Location
Greater London, England, United Kingdom
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … open-source models — with fallback routing, token, cost, and latency management Implement LLMOps in production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions — workshops, proofs of concept ...

Core Platform Developer

Location
Greater London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment.**Key Responsibilities*** Design and build shared backend services, frameworks … developer tooling that support internal application and service development.* Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities.* Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams.* Help ...

Software Architect (Java or C#)

Location
United Kingdom
Engineering Partnership Work actively with Engineering teams during design, development, and production-readiness reviews. Advise and challenge teams on service architecture, fault tolerance, scalability, observability, deployment safety, and operational readiness, helping them to make pragmatic trade-offs. Support teams in diagnosing complex performance, latency, throughput, and resource-utilisation issues. Help … establish engineering standards and reusable patterns for reliable, maintainable services. Performance & Observability Lead investigations into performance bottlenecks across applications, infrastructure, databases, queues, networks, and third-party dependencies. Improve observability through metrics, logs, traces, dashboards, alerting, and service-level indicators. Help teams design meaningful alerts that identify user-impacting issues while ...

Software Engineers

Location
Douglas, Isle of Man, United Kingdom
integrations between enterprise systems and third-party platforms. Supporting and optimising cloud infrastructure and platform services within Azure environments. Driving operational excellence through monitoring, observability, automation and service reliability practices. Developing and maintaining automated testing frameworks and quality engineering standards. Supporting DevOps, CI/CD and infrastructure automation initiatives. Contributing … automation, quality engineering and automated testing frameworks. API development, application integration and distributed systems. DevOps tooling, CI/CD pipelines and automation practices. Monitoring, observability, incident management and operational resilience. Agile software delivery and modern engineering methodologies. Strong analytical, troubleshooting and problem-solving skills. Excellent communication and stakeholder management capabilities. ...

SC Cleared AWS DevOps & Platform Engineer (Remote)

Location
England, United Kingdom
cloud infrastructure in AWS, build automated pipelines with GitLab CI/CD and ArgoCD, and manage Kubernetes, Docker, Terraform and related tools while ensuring observability with Grafana/Prometheus. Strong hands-on AWS skills, container orchestration, IaC and CI/CD expertise are essential #J-18808-Ljbffr ...

AI / Machine Learning Engineer – Agentic LLM Systems (Contract)

Location
Greater London, England, United Kingdom
services and APIs to support AI-driven applications and workflowsBuild and run experiments to improve reliability, latency, cost, and success ratesContribute to evaluation frameworks, observability, and monitoring of LLM/agent performance What We’re Looking For: 4+ years’ experience in ML/AI engineering (LLMs, recommender systems, optimisation … codeComfortable debugging complex, distributed AI/ML systemsExperience running and analysing large-scale experiments and performance metrics (latency, accuracy, cost)Exposure to monitoring/observability tools for production systems Nice to Have: Experience with multi-agent systems or distributed AI architecturesExperience optimising LLM usage across multiple providers (cost/performance ...

AI Native SW Engineering

Location
Greater London, England, United Kingdom
production‐grade agentic systems at enterprise scale: multi‐agent orchestration across complex environments, RAG pipelines, policy‐based routing, memory management, and programme‐level lifecycle observability Define RAG pipeline standards across engagements: establish chunking and embedding strategies, set quality benchmarks, and ensure metric‐backed tradeoff decisions are documented and transferable … standard design practice across providers including OpenAI, Anthropic, Vertex AI, and open‐source models Own LLMOps at programme scale: eval strategy, prompt governance, observability tooling standards, safety monitoring and cost controls across multiple concurrent systems Lead client engineering engagements at senior level — facilitate architecture design sessions, lead proof‐of‐concept ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
pipelines Supporting the deployment of applications into production Working with Azure Kubernetes Service (AKS) and containerised applications Improving cloud networking, security, resilience and observability Working with Entra ID and identity/access management Partnering closely with Software Engineers to make deployments simpler, safer and more reliable Helping shape platform standards … Cloud networking Deploying and supporting applications in production Experience with Kubernetes/AKS would be particularly useful. Azure DevOps, Entra ID, cloud security, observability, cost optimisation or relevant Azure/Kubernetes certifications would all add value, but you don't need to tick every box. More important is that ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … call rota and provide on-call support for other SRE engineers.* Can write advanced automation scripts for incident response, including failovers and rollbacks.**Observability** * Has a deep technical understanding of observability techniques across the full stack and can bring clarity to complex incidents or performance issues.* Able to create templated ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
Role Profile:**We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a **Technical Lead SRE**, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. … projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design.Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.Design and evolve monitoring and alerting solutions that ...