551 to 575 of 4,012 Observability Jobs

Applied AI ML Lead - Python & Agentic AI

Location
Glasgow, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...

Platform / DevOps Engineer

Location
City of Westminster, England, United Kingdom
Manager, EventBridge provisioned as code, sized sensibly, and cost‐aware (you'll make the calls on things like NAT vs VPC endpoints). Own observability and reliability. CloudWatch alarms and dashboards, SNS alerting, data freshness and quality signals, and automated recovery for the pipelines that need it. When something breaks … deliberate applies, and easy rollbacks. Everything is code. Infrastructure, pipelines, access, and policy all live in version control and ship through review. Observability first. If we run it, we can see it - and we get told before our users do. You own what you ship. Strong ownership, low ceremony. ...

Software Engineer III - Python

Location
Glasgow, Scotland, United Kingdom
infrastructure-as-code using Terraform within established team patterns across modules, environments, and state management Improve operability of services by adding and using observability tooling including logs, metrics, traces, dashboards, and alerts, and participate in incident response and root-cause analysis Leverage enterprise-authorized AI coding assist tools within … implementing application logic and APIs on top of relational data Experience building APIs and microservices using REST or gRPC, including contracts, security basics, and observability Practical experience delivering LLM-based features as part of software systems, with familiarity with agentic patterns Working knowledge of delivery and operations including CI/ ...

Technology Lead (Remote - UK)

Hiring Organisation
Reonomy
Location
London, United Kingdom
Salary
£ 70 K
reliability, scalability, and long-term sustainabilityManage and reduce technical debt strategicallyDesign cost-efficient, cloud-native solutionsEngineering ExcellenceSet and uphold standards for code quality, testing, observability, security, and documentationChampion automated testing and “shift-left” quality practicesPromote DevSecOps principles to embed security and compliance into developmentLead through thorough code reviews, design documentation … Deep understanding of system design, distributed systems, scalability, APIs, and data modelingExperience with Infrastructure as Code (Terraform or similar)Strong CI/CD and observability experienceExperience implementing robust testing strategiesComfortable working cross-functionally in Agile environmentsStrong communication skills and ability to operate effectively in ambiguous situationsUnlock your Altus Experience ...

Sr. Software Engineer

Hiring Organisation
Meltwater Group
Location
London, UK
Employment Type
Full-time
while architecting systems that handle high-throughput data pipelines. Participate in building robust export pipelines, streaming architectures, webhook integrations and MCP servers. Maintain high observability and reliability standards using tools like Coralogix, CloudWatch, and Grafana. Participate in on-call rotation and incident response for owned services. What You'll Bring4+ … Swagger, static site generators).Familiarity with authentication, API gateways, and rate limiting strategies. Experience in compliance standards for APIs and data handling. Experience with observability tools and practices. Our Tech Stack: Languages: Golang (primary) with some TypeScriptInfrastructure: AWS, S3, Lambda, SQS, SNS, CloudFront, Kubernetes (Helm), Kong API GatewayDatabases: Postgres, Redis ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks … model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents) LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request) Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails Any other ...

SC Cleated DevOps Engineer

Location
England, United Kingdom
well as working with JFrog within software development and deployment environments. Python: Comfortable using Python for scripting, automation and supporting platform engineering activities. Monitoring & Observability: Experience working with tools such as Grafana, Prometheus and OpenTelemetry to monitor, troubleshoot and improve the reliability of cloud and platform environments. Development Practices: Familiarity ...

Integration Architect

Location
United Kingdom
Terraform or Bicep . Design integration patterns using REST APIs, microservices, Azure Service Bus, Event Grid and event-driven architectures . Establish monitoring and observability using Azure Monitor, Log Analytics and Application Insights . Present solution designs through architecture governance and review processes. Provide technical leadership and guidance to engineering … implementing APIOps/GitOps approaches for API lifecycle management. Knowledge of Azure API Center or similar API cataloguing/governance platforms. Experience with Azure observability tooling including Application Insights, Azure Monitor and Log Analytics . Knowledge of emerging AI integration patterns such as Model Context Protocol (MCP) or AI Gateway ...

Site Reliability Engineer

Location
Newcastle upon Tyne, England, United Kingdom
Makes This Role Great: Develop and maintain scalable infrastructure as code (IaC) using Terraform to ensure reliable and scalable cloud environments. Implement and enhance observability solutions using tools like New Relic, DataDog, Sumologic and Splunk for monitoring, logging, and alerting. Perform code deployments and manage CI/CD pipelines using … management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized SRE observability experience with New Relic or DataDog. Familiarity with OpenTelemetry, AIOps, MLOps, or SecOps. Location: Newcastle, UK - In-Office (at least 4 days per week ...

Platform Engineer III - (Pipelines & Developer Experience)

Location
Leeds, England, United Kingdom
least privilege, policy‐as‐code, scanning). Improve developer experience: faster feedback loops, quality gates, ephemeral/preview environments, and great documentation. Instrument pipeline observability (Datadog or equivalent) and define SLOs (queue time, lead time, change fail rate, MTTR) to drive reliability. Automate IaC workflows (Terraform/Terragrunt) and integrate … implement and maintain CI/CD release pipelines. Experience with scripting and programming (.NET preferred; familiarity with Go, Python, PowerShell beneficial). Knowledge of observability tooling, chaos testing, and incident management. Strong analytical and problem‐solving abilities, with the capability to closely collaborate with engineering teams. Highly outcome‐oriented, pragmatic ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity ...

Lead SRE - Chase UK

Location
Greater London, England, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational ExcellenceDesign and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry).Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skillsFamiliarity with infrastructure ...

Senior MLOps Engineer

Hiring Organisation
AECOM
Location
London, United Kingdom
Salary
£ 70 K
/CD, automation, and model delivery workflows used across the engineering teamDevelop Python services, APIs, and platform tooling that support production ML workloadsOwn reliability, observability, performance, availability, and security across production ML systemsTroubleshoot complex production issues across software, infrastructure, and ML workflowsDrive architecture decisions and partner with ML engineers, data … with Kubernetes and infrastructure as code, such as TerraformExperience with ML lifecycle tooling such as MLflow, Weights & Biases, or similar platformsKnowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, or OpenTelemetryExperience with distributed ML workloads, model serving, GPU infrastructure, or performance optimizationBackground in reinforcement learning, optimization, or generative ...

Senior Full Stack Engineer

Location
Greater London, England, United Kingdom
technical designs and code produced by delivery partners, identifying risks and opportunities for improvement Champion engineering excellence across code quality, testing, continuous delivery, security, observability and platform reliability Collaborate closely with AI, Machine Learning, Product, Infrastructure and DevOps teams to deliver integrated platform capabilities and share knowledge across the engineering … Familiarity with AI orchestration frameworks, intelligent workflows and emerging AI technologies Experience with data platforms and infrastructure that support AI workloads Knowledge of monitoring, observability and continuous delivery practices Experience within telecommunications or other large scale customer facing digital platforms What's in it for you? Competitive salary and bonus ...

Devops Engineer

Location
Greater London, England, United Kingdom
deployment, software repositories, databases, and web servers Own the patching and update lifecycle for managed systems Monitoring & Reliability Implement and maintain monitoring, alerting, and observability across both the existing VM estate and the new container environment Proactively identify risks, bottlenecks, and failure patterns before they impact users Define and track … e.g. Terraform, Ansible, Puppet, or similar) Solid scripting ability in Bash and at least one higher‐level language (Python preferred) Experience with monitoring and observability tooling (e.g. Prometheus, Grafana, Datadog, or similar) Strong incident diagnosis skills—able to work from vague symptoms to root cause using logs, metrics, and reasoning ...

Intermediate Data Engineer

Location
Cheadle, England, United Kingdom
maintain automated deployment, testing and operational workflows using GitLab pipelines, version control, Terraform/OpenTofu and scripting. Contribute to data quality, lineage, governance and observability practices to ensure data is trusted, traceable, secure and supportable throughout its lifecycle. Investigate data quality, platform reliability, pipeline performance and operational issues, identifying practical … willingness to learn, AI-assisted engineering tools such as Kiro, GitHub Copilot or similar technologies to improve engineering effectiveness and delivery outcomes. Familiarity with observability and troubleshooting tooling such as Dynatrace, OpenSearch/Elasticsearch or CloudWatch logs. Ability to communicate clearly, analyse problems, work collaboratively and explain technical concepts ...

Site Reliability Engineer (SRE) – Cloud Engineer

Location
Glasgow, Scotland, United Kingdom
potential extension) Start Date: Immediate Key Responsibilities Design, build and maintain highly available cloud infrastructure across multi-cloud environments. Drive SRE best practices including observability, resilience, reliability engineering and automation. Develop and maintain Infrastructure as Code (IaC) solutions using Terraform. Build and enhance automation tooling using Python. Support cloud platform ...

Python Developer - AI / LLM Platform

Hiring Organisation
Ncounter
Location
Reading, Berkshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £100,000 per annum
while improving existing components through structured refactoring • Integrate LLM APIs and develop practical AI capabilities across live platforms • Take ownership of production issues, testing, observability and ongoing platform reliability • Strengthen application security, deployment practices and technical documentation We are looking for: • 5+ years of commercial Python development experience, ideally including ...

Solutions Architect

Location
Leeds, England, United Kingdom
Strong problem-solving, communication, and leadership skills. Experience with GraphQL, gRPC, and Service Mesh (Istio/Linkerd). Familiarity with AI/ML integrations, observability tools (Dynatrace, Grafana, ELK). Experience or good understanding in Agentic AI #J-18808-Ljbffr ...

Staff Platform Site Reliability Engineer

Hiring Organisation
Index Exchange
Location
London, UK
Employment Type
Full-time
about its architecture. Must Have8+ years in platform engineering, SRE, infrastructure engineering, or DevOps. Deep experience with Linux internals: kernel tuning, network stack, system observability, security. Strong Kubernetes expertise: cluster lifecycle, networking, storage, RBAC, multi-cluster—across bare-metal and cloud (EKS, GKE).Infrastructure-as-code at scale: Terraform, Ansible … driving technical strategy across teams—not just executing within one. Valuable ExperienceDistributed storage systems (e.g. Ceph)Big data infrastructure: Hadoop, Spark, HBase, Kafka. Observability stack design: Prometheus, Grafana, ELK, Mimir, Loki, Tempo. Secrets management (Vault), certificate management, access control at scale. Hybrid cloud architectures: federating public cloud (AWS, GCP) with ...

Senior DevOps

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£80,000
including Kubernetes and Docker Infrastructure as Code using Terraform or similar tools CI/CD pipeline development with GitLab, Jenkins or equivalent Monitoring and observability tools such as Prometheus, Grafana or Elastic Secure infrastructure, automation and platform engineering Agile software delivery within cross-functional teams Building resilient, scalable and highly ...

Senior DevOps

Hiring Organisation
Anson Mccade
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£80,000
including Kubernetes and Docker Infrastructure as Code using Terraform or similar tools CI/CD pipeline development with GitLab, Jenkins or equivalent Monitoring and observability tools such as Prometheus, Grafana or Elastic Secure infrastructure, automation and platform engineering Agile software delivery within cross-functional teams Building resilient, scalable and highly ...

Azure DevOps Engineer

Hiring Organisation
CBSbutler Holdings Limited
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£600 - £645 per day
Code using Terraform Work across Kubernetes/Docker/AWS EKS Implement secrets, identity and access management using technologies such as HashiCorp Vault Establish observability, monitoring, logging and audit controls Develop secure-by-design engineering practices aligned to MOD security requirements Coordinate DevSecOps engineers, developers, testers and infrastructure teams Create ...

DevOps Lead Engineer

Hiring Organisation
Sanderson Government and Defence
Location
London, United Kingdom
Employment Type
Contract
Kubernetes, AKS, EKS, Azure Virtual Desktop, and cloud-native platforms. Strong understanding of cloud security, identity management, and regulated environments. Experience with monitoring and observability tools such as Splunk and Azure Monitor. Excellent stakeholder management and communication skills. This is an excellent opportunity to join a critical defence programme where ...