1,051 to 1,075 of 5,035 Observability Jobs

Machine Learning Engineering Lead

Location
United Kingdom
maintainability, reliability, and production support. Strong understanding of ML engineering and MLOps practices, including model lifecycle management, CI/CD, testing, monitoring, release management, observability, and operational support. Practical experience with LLM-based capabilities, including retrieval-augmented generation, semantic search, embeddings, prompt design, evaluation, guardrails, and observability. Experience with agentic ...

Machine Learning Engineering Lead

Hiring Organisation
Hackajob Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
maintainability, reliability, and production support. Strong understanding of ML engineering and MLOps practices, including model lifecycle management, CI/CD, testing, monitoring, release management, observability, and operational support. Practical experience with LLM-based capabilities, including retrieval-augmented generation, semantic search, embeddings, prompt design, evaluation, guardrails, and observability. Experience with agentic ...

Agentforce Operations FDE

Location
Greater London, England, United Kingdom
tools (Cursor, Claude, Vibes) embedded directly into your daily developer workflow to move from architecture sketch to deployable code in days, not months. Telemetry & Observability: Model data, ship ETL/ELT pipelines, and build real‐time agent performance dashboards to track execution health, fallback triggers, and business outcomes. Field ...

Machine Learning Engineering Lead

Location
Farringdon, England, United Kingdom
maintainability, reliability, and production support. Strong understanding of ML engineering and MLOps practices, including model lifecycle management, CI/CD, testing, monitoring, release management, observability, and operational support. Practical experience with LLM‐based capabilities, including retrieval‐augmented generation, semantic search, embeddings, prompt design, evaluation, guardrails, and observability. Experience with agentic ...

Software Engineer (Backend)

Location
Greater London, England, United Kingdom
Software Engineer (backend) Location: London or New York - in office 4 days a week About Lorum Global payments are not broken. Incentives are. Clearing has been deprioritized inside balance sheet driven institutions whose models rely ...

Senior Software Engineer, Git Systems

Location
United Kingdom
About GitHub GitHub is the world’s leading platform for agentic software development — powered by Copilot to build, scale, and deliver secure software. Over 180 million developers, including more than 90% of the Fortune 100 ...

Solution Architect

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
The team you'll be working with: We are looking for a Solution Architect to help our clients design and deliver modern, scalable and secure digital solutions, primarily on Microsoft technologies. You will combine architecture ...

AI and Cloud Director - Technology Consulting (AI, data and analytics)

Hiring Organisation
Business Integration Partners
Location
London, UK
Employment Type
Full-time
AI and Cloud DirectorLocation: London/HybridBusiness: BIP UKPractice: xTech - AI, Data and CloudReporting to: UK xTech LeadershipEmployment type: Full timeAbout BIP UKFounded in 2003, BIP is an international consulting firm with more than 6 ...

Senior Data Engineer

Hiring Organisation
McKesson
Location
London, UK
Employment Type
Full-time
Role OverviewThe Senior Data Engineer is the technical owner of the ClarusONE data platform. This is a hands-on engineering role with real ownership: alongside building and delivering data pipelines, you will raise the engineering ...

Cloud Architect (DV Security Cleared)

Location
Greater London, England, United Kingdom
The Cloud Architect is responsible for analysing existing on-premise systems and designing, planning, and supporting their migration into a secure, scalable, and cloud-native environment. The role focuses on modernising architectures, leveraging managed services ...

Full Stack Engineer (AI Startup)

Hiring Organisation
Trismik
Location
United kingdom
Why join us? At Trismik, we’re a team of tech geeks from Cambridge University, Salesforce, and Amazon. With over 38 years of research behind us, we’re not just building tools - we’re sparking ...

AI & ML Engineer

Location
Greater London, England, United Kingdom
About the role The AI & ML Engineering team accelerates the adoption of AI across the business, championing innovation while ensuring our machine learning solutions are robust, scalable, and cost-efficient. We enable teams to solve ...

Software Engineer

Hiring Organisation
IBM SIXworks Limited
Location
Taunton, Somerset, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Work on Technology That Protects What Matters At SiXworks , we build secure digital solutions that support Defence and National Security missions . Our teams work on complex problems where reliability, security, and speed of innovation ...

Senior Software Engineer - iCloud Platform - Observability

Location
Greater London, England, United Kingdom
services platform and infrastructure. You will be working on foundational systems that power iCloud services, including distributed data platforms, storage systems, and a unified observability infrastructure that standardizes monitoring and telemetry approaches across all iCloud services serving billions of customers.iCloud manages data and services at massive scale! Our unified observability … health of all iCloud services, providing comprehensive telemetry collection, real-time processing, and sophisticated analysis capabilities that span billions of active Apple customers. This observability ecosystem is purpose-built to deliver highly scalable and performant solutions, that maintain the highest standards of user privacy and security.We are a world-class ...

Solution Engineer

Location
Greater London, England, United Kingdom
## Solution EngineerLondon, UK · Full-time · Senior#### About The PositionCoralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs, metrics, trace … security events with features such as APM, RUM, SIEM, Kubernetes monitoring and more, all enhancing operational efficiency and reducing observability spend by up to 70%.Solution Architects in Coralogix are key in meeting our customers’ expectations and helping them utilize their observability and security data. We are looking for hard ...

SRE

Location
Hove, England, United Kingdom
No. of Positions: 1 We are seeking an experienced Site Reliability Engineer (SRE) to drive the modernization of IT operations through the implementation of observability practices, automation, and reliability engineering principles. The role requires a strategic thinker with strong hands‐on expertise who can enhance system reliability, scalability, and operational … practices, automate operational workflows, and establish robust monitoring and incident management frameworks. Key Responsibilities Collaborate with engineering teams to modernize IT operations by improving observability, automation, and operational efficiency. Design and implement observability platforms to effectively monitor system health, performance, and reliability. Develop strategies for AI-driven alerting and proactive ...

Senior Platform Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£65,000
Senior Platform Engineer Deliver and support cloud platform engineering solutions across client environments. Design services with a reliability mindset, using SLIs, SLOs, and observability practices. Implement and maintain Infrastructure as Code using Terraform across environments. Support incident management, problem management, and continuous improvement of production platforms. Contribute to observability solutions … Infrastructure as Code across non-production and production environments. Understanding of SRE principles including SLIs, SLOs, error budgets, resilience, and reliability. Experience with observability and monitoring tools such as Dynatrace or similar. Experience supporting production platforms including incident and problem management. Exposure to AIOps practices and automation for proactive issue ...

Site Reliability Engineer

Location
Cardiff, Wales, United Kingdom
week in Cardiff office) About the Role This company provides managed AI operations for technology businesses. The company operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software. This is a chance to take real ownership in a Cardiff business on the front line … Datadog Advanced Partner (UK) and holds the accolade of being the world's first accredited MSP powered by Datadog. Datadog is a Nasdaq-listed observability platform. The company's technical focus is on observability and LLM observability specifically. Key Responsibilities Managed Service Delivery - Support customers across a mix of traditional ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...

Core Platform Developer

Location
Greater London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment.**Key Responsibilities*** Design and build shared backend services, frameworks … developer tooling that support internal application and service development.* Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities.* Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams.* Help ...

AI Native SW Eng

Location
United Kingdom
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … open-source models with fallback routing, token, cost, and latency management Implement LLMOps in production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions workshops, proofs of concept ...

Software Architect (Java or C#)

Location
United Kingdom
Engineering Partnership Work actively with Engineering teams during design, development, and production-readiness reviews. Advise and challenge teams on service architecture, fault tolerance, scalability, observability, deployment safety, and operational readiness, helping them to make pragmatic trade-offs. Support teams in diagnosing complex performance, latency, throughput, and resource-utilisation issues. Help … establish engineering standards and reusable patterns for reliable, maintainable services. Performance & Observability Lead investigations into performance bottlenecks across applications, infrastructure, databases, queues, networks, and third-party dependencies. Improve observability through metrics, logs, traces, dashboards, alerting, and service-level indicators. Help teams design meaningful alerts that identify user-impacting issues while ...

Software Engineers

Location
Douglas, Isle of Man, United Kingdom
integrations between enterprise systems and third-party platforms. Supporting and optimising cloud infrastructure and platform services within Azure environments. Driving operational excellence through monitoring, observability, automation and service reliability practices. Developing and maintaining automated testing frameworks and quality engineering standards. Supporting DevOps, CI/CD and infrastructure automation initiatives. Contributing … automation, quality engineering and automated testing frameworks. API development, application integration and distributed systems. DevOps tooling, CI/CD pipelines and automation practices. Monitoring, observability, incident management and operational resilience. Agile software delivery and modern engineering methodologies. Strong analytical, troubleshooting and problem-solving skills. Excellent communication and stakeholder management capabilities. ...

SC Cleared AWS DevOps & Platform Engineer (Remote)

Location
England, United Kingdom
cloud infrastructure in AWS, build automated pipelines with GitLab CI/CD and ArgoCD, and manage Kubernetes, Docker, Terraform and related tools while ensuring observability with Grafana/Prometheus. Strong hands-on AWS skills, container orchestration, IaC and CI/CD expertise are essential #J-18808-Ljbffr ...

AI / Machine Learning Engineer – Agentic LLM Systems (Contract)

Location
Greater London, England, United Kingdom
services and APIs to support AI-driven applications and workflowsBuild and run experiments to improve reliability, latency, cost, and success ratesContribute to evaluation frameworks, observability, and monitoring of LLM/agent performance What We’re Looking For: 4+ years’ experience in ML/AI engineering (LLMs, recommender systems, optimisation … codeComfortable debugging complex, distributed AI/ML systemsExperience running and analysing large-scale experiments and performance metrics (latency, accuracy, cost)Exposure to monitoring/observability tools for production systems Nice to Have: Experience with multi-agent systems or distributed AI architecturesExperience optimising LLM usage across multiple providers (cost/performance ...