2,326 to 2,350 of 5,534 Observability Jobs

Senior DevOps Engineer - Cloud Automation & AI Tools

Location
Greater London, England, United Kingdom
operate RX Cloud Platforms across AWS and Azure from Richmond, London. You will partner with software engineers to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes automation-first priorities, AI-assisted tooling, and collaboration with agile teams to drive incident resolution and continuous improvement. #J ...

Platform Engineer: Kubernetes, Cloud & Automation

Location
Greater London, England, United Kingdom
harden infrastructure, apply IaC with Terraform/Ansible/Helm, and ensure secure, scalable operations across GCP, AWS and Azure, while contributing to observability and performance tuning. #J-18808-Ljbffr ...

VP DevOps/SRE: AI/ML Infra, CI/CD & Reliability

Location
Glasgow, Scotland, United Kingdom
DevOps/SRE Engineer - Vice President to own automation, reliability, and production operations for AI/ML services. You will build CI/CD, observability, and incident-management practices across international markets. Leverage Terraform, Kubernetes, and cloud-native tooling to scale release automation and reliability, while mentoring engineers and enforcing ...

AI Platform Engineer - Azure, Kubernetes & CI/CD

Location
Manchester, England, United Kingdom
responsibilities include building AKS-based deployments, IaC with Terraform, and robust CI/CD pipelines using Azure DevOps, with a focus on reliability, observability, and secure practices. #J-18808-Ljbffr ...

Senior Cloud DevOps Engineer — AWS/Azure, IaC & CI/CD

Location
Richmond, England, United Kingdom
from Richmond, London. You’ll collaborate with software teams to improve infrastructure, deployment automation and platform reliability. You will implement security best practices, maintain observability, and drive continuous improvement using AI-assisted tooling and modern CI/CD processes. Flexible working patterns are available. #J-18808-Ljbffr ...

Cloud DevOps Engineer — AI Tools, Flexible Hours

Location
Richmond, England, United Kingdom
Cloud Platforms across AWS and Microsoft Azure. You will work closely with software engineering and cloud teams to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes Infrastructure as Code, Terraform/CloudFormation, GitHub Actions, and cloud security, with scripting #J-18808-Ljbffr ...

Senior Cloud SRE: Azure, Terraform & Kubernetes

Location
Newcastle upon Tyne, England, United Kingdom
Trimble is seeking a Site Reliability Engineer to own and scale production infrastructure on cloud environments. You will implement IaC with Terraform, enhance observability using multiple monitoring tools, and drive CI/CD pipelines with Jenkins and GitHub. The role involves leading incident response and collaborating across teams to ensure ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

DevOps and Machine Learning Operations Engineer

Location
Manchester, England, United Kingdom
runtime platform the project depends on, as well as the model-serving path. The role covers infrastructure as code, continuous integration and deployment, observability, cost control, and production support. Applicants must have experience running systems in production and being accountable for their reliability. Main Responsibilities Infrastructure and Environments: Define … gates that fail closed. Automate database migration and rollback for reversible releases. Support progressive delivery, including staged rollout and fast rollback, with deployment tracking. Observability and Operations: Instrument services with structured logging, metrics, and tracing, and define user-focused alerts. Establish service level objectives and report against them. Run incident ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Principal Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Drive automation across infrastructure, application delivery, operational processes, and platform management Establish platform standards, engineering patterns, and best practices Improve platform reliability, scalability, performance, observability, and operational efficiency Reduce engineering friction and accelerate software delivery Establish engineering principles and guardrails for security, reliability, and governance Lead complex platform initiatives from … automated software delivery Experience automating operational processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives ...

Senior AI Platform Engineer - GenAI & LLM Architect

Location
Sheffield, England, United Kingdom
oversee LLM, RAG, and Agentic AI integrations across Azure, AWS, and GCP, collaborate with architects to ensure secure, maintainable solutions, and set standards for observability, governance, and #J-18808-Ljbffr ...

Software Engineer – Manchester (SC Clearance)

Location
Manchester, England, United Kingdom
cloud‐first applications in an agile setting. You’ll work on traditional and serverless services, using containers and CI/CD while ensuring robust observability and security considerations. Required are several years of experience across Java, Typescript, Python, Go or C#, plus strong Cloud Native and IaC tooling. Some travel ...

Cloud DevOps Engineer — Flexible Hours & AI Automation

Location
Greater London, England, United Kingdom
will build, automate, secure, and support RX Cloud Platforms on AWS and Azure, collaborating with software and cloud teams to boost deployment automation, observability, and platform reliability. You will implement IaC with Terraform, contribute to CI/CD pipelines using GitHub Actions, and apply security best practices while improving monitoring ...

Python Developer – Security – Cambridge/London

Location
Cambridge, England, United Kingdom
Collaborate closely with product managers and researchers to translate ideas into working demos and production features. Help troubleshoot, debug, and optimise systems (performance, reliability, observability). Contribute to CI/CD, containerisation (Docker) workflows, and cloud deployment patterns as required. Communicate trade-offs and design decisions clearly; adapt to changing … demonstrable ability considered). Experience with Kubernetes or container orchestration. Exposure to ML/AI projects, data pipelines, or research-driven engineering. Experience with observability tooling, CI/CD, and testing best practices. Degree in Computer Science, Engineering, or related technical field (or demonstrable equivalent). *Rates depend on experience ...

Product Associate - SRE Team

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Product Associate - SRE Team

Location
Westminster, West End, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Senior Cloud Software Engineer – AWS, AI‐Driven Dev

Location
Christchurch, England, United Kingdom
scalable software in an agile setting. The role demands hands-on AWS development, proficiency in modern languages, and experience with Kubernetes, IaC, and observability tools. You will use AI-assisted coding tools and participate in design reviews to strengthen engineering practices. #J-18808-Ljbffr ...

Senior Software Engineer, GenAI Platform

Location
Greater London, England, United Kingdom
technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance … hours and cutting inference cost by multiples - while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in. Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner ...

Remote Developer Support Engineer (London)

Hiring Organisation
Braintrust
Location
Derbyshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Hiring Organisation
Braintrust
Location
Lancashire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Hiring Organisation
Braintrust
Location
Worcestershire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...