2,326 to 2,350 of 5,499 Permanent Observability Jobs

Agent Engineer

Location
Greater London, England, United Kingdom
users to create user-friendly solutions for complex business processes Participate in technical design reviews, planning sessions, and code reviews Contribute to infrastructure and observability practices with the Engineering team Continuously improve quality, reliability, and usability across internal platforms Support Elliptic's mission to make crypto markets safer, more transparent … framework experience such as NestJS or Express is nice to have Terraform or infrastructure-as-code experience is nice to have Datadog or similar observability platform experience is nice to have DynamoDB or other NoSQL database experience at scale is nice to have Distributed or event-driven architecture experience, including ...

Principal Software Engineer - Full Stack - AI

Location
York and North Yorkshire, England, United Kingdom
full stack, guiding the development of responsive frontend applications (React/TypeScript) and robust, scalable backend services (Python, Java, Kotlin, or Node.js). LLM Observability & Reliability: Establish robust LLM observability, evaluations, and caching, implementing latency optimisations and comprehensive monitoring (logging, usage tracking, agent behaviour). Operational Excellence: Champion high availability … performance optimisation, and observability across frontends and backend microservices, focusing on practices that maintain platform reliability and optimise MTTD and MTTR. Mentorship & Collaboration: Elevate the engineering organisation by mentoring senior and junior engineers, conducting rigorous code and system design reviews, and partnering with product managers to translate product visions into ...

Azure Systems Engineer

Location
Greater London, England, United Kingdom
integrations and production environments Implementing Infrastructure as Code (IaC) and automation solutions Supporting identity, access management, security and governance across Azure Improving platform monitoring, observability, resilience and performance Supporting SQL Server and Azure SQL environments, including access, backups and troubleshooting Contributing to vulnerability management, incident response and operational security Strong … Entra ID, IAM, RBAC and privileged access Understanding of secure application delivery, configuration and secrets management Experience with Azure Monitor, Application Insights or equivalent observability tooling Knowledge of SQL Server/Azure SQL , including access management, backups and troubleshooting Strong understanding of TCP/IP, DNS, routing, firewalls and private ...

Senior DevOps Engineer - Cloud Automation & AI Tools

Location
Greater London, England, United Kingdom
operate RX Cloud Platforms across AWS and Azure from Richmond, London. You will partner with software engineers to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes automation-first priorities, AI-assisted tooling, and collaboration with agile teams to drive incident resolution and continuous improvement. #J ...

Platform Engineer: Kubernetes, Cloud & Automation

Location
Greater London, England, United Kingdom
harden infrastructure, apply IaC with Terraform/Ansible/Helm, and ensure secure, scalable operations across GCP, AWS and Azure, while contributing to observability and performance tuning. #J-18808-Ljbffr ...

VP DevOps/SRE: AI/ML Infra, CI/CD & Reliability

Location
Glasgow, Scotland, United Kingdom
DevOps/SRE Engineer - Vice President to own automation, reliability, and production operations for AI/ML services. You will build CI/CD, observability, and incident-management practices across international markets. Leverage Terraform, Kubernetes, and cloud-native tooling to scale release automation and reliability, while mentoring engineers and enforcing ...

AI Platform Engineer - Azure, Kubernetes & CI/CD

Location
Manchester, England, United Kingdom
responsibilities include building AKS-based deployments, IaC with Terraform, and robust CI/CD pipelines using Azure DevOps, with a focus on reliability, observability, and secure practices. #J-18808-Ljbffr ...

Senior Cloud DevOps Engineer — AWS/Azure, IaC & CI/CD

Location
Richmond, England, United Kingdom
from Richmond, London. You’ll collaborate with software teams to improve infrastructure, deployment automation and platform reliability. You will implement security best practices, maintain observability, and drive continuous improvement using AI-assisted tooling and modern CI/CD processes. Flexible working patterns are available. #J-18808-Ljbffr ...

Cloud DevOps Engineer — AI Tools, Flexible Hours

Location
Richmond, England, United Kingdom
Cloud Platforms across AWS and Microsoft Azure. You will work closely with software engineering and cloud teams to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes Infrastructure as Code, Terraform/CloudFormation, GitHub Actions, and cloud security, with scripting #J-18808-Ljbffr ...

Senior Cloud SRE: Azure, Terraform & Kubernetes

Location
Newcastle upon Tyne, England, United Kingdom
Trimble is seeking a Site Reliability Engineer to own and scale production infrastructure on cloud environments. You will implement IaC with Terraform, enhance observability using multiple monitoring tools, and drive CI/CD pipelines with Jenkins and GitHub. The role involves leading incident response and collaborating across teams to ensure ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

DevOps and Machine Learning Operations Engineer

Location
Manchester, England, United Kingdom
runtime platform the project depends on, as well as the model-serving path. The role covers infrastructure as code, continuous integration and deployment, observability, cost control, and production support. Applicants must have experience running systems in production and being accountable for their reliability. Main Responsibilities Infrastructure and Environments: Define … gates that fail closed. Automate database migration and rollback for reversible releases. Support progressive delivery, including staged rollout and fast rollback, with deployment tracking. Observability and Operations: Instrument services with structured logging, metrics, and tracing, and define user-focused alerts. Establish service level objectives and report against them. Run incident ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Principal Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Drive automation across infrastructure, application delivery, operational processes, and platform management Establish platform standards, engineering patterns, and best practices Improve platform reliability, scalability, performance, observability, and operational efficiency Reduce engineering friction and accelerate software delivery Establish engineering principles and guardrails for security, reliability, and governance Lead complex platform initiatives from … automated software delivery Experience automating operational processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives ...

Senior AI Platform Engineer - GenAI & LLM Architect

Location
Sheffield, England, United Kingdom
oversee LLM, RAG, and Agentic AI integrations across Azure, AWS, and GCP, collaborate with architects to ensure secure, maintainable solutions, and set standards for observability, governance, and #J-18808-Ljbffr ...

Software Engineer – Manchester (SC Clearance)

Location
Manchester, England, United Kingdom
cloud‐first applications in an agile setting. You’ll work on traditional and serverless services, using containers and CI/CD while ensuring robust observability and security considerations. Required are several years of experience across Java, Typescript, Python, Go or C#, plus strong Cloud Native and IaC tooling. Some travel ...

Cloud DevOps Engineer — Flexible Hours & AI Automation

Location
Greater London, England, United Kingdom
will build, automate, secure, and support RX Cloud Platforms on AWS and Azure, collaborating with software and cloud teams to boost deployment automation, observability, and platform reliability. You will implement IaC with Terraform, contribute to CI/CD pipelines using GitHub Actions, and apply security best practices while improving monitoring ...

Python Developer – Security – Cambridge/London

Location
Cambridge, England, United Kingdom
Collaborate closely with product managers and researchers to translate ideas into working demos and production features. Help troubleshoot, debug, and optimise systems (performance, reliability, observability). Contribute to CI/CD, containerisation (Docker) workflows, and cloud deployment patterns as required. Communicate trade-offs and design decisions clearly; adapt to changing … demonstrable ability considered). Experience with Kubernetes or container orchestration. Exposure to ML/AI projects, data pipelines, or research-driven engineering. Experience with observability tooling, CI/CD, and testing best practices. Degree in Computer Science, Engineering, or related technical field (or demonstrable equivalent). *Rates depend on experience ...

Product Associate - SRE Team

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Product Associate - SRE Team

Location
Westminster, West End, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Senior Cloud Software Engineer – AWS, AI‐Driven Dev

Location
Christchurch, England, United Kingdom
scalable software in an agile setting. The role demands hands-on AWS development, proficiency in modern languages, and experience with Kubernetes, IaC, and observability tools. You will use AI-assisted coding tools and participate in design reviews to strengthen engineering practices. #J-18808-Ljbffr ...

Senior Software Engineer, GenAI Platform

Location
Greater London, England, United Kingdom
technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance … hours and cutting inference cost by multiples - while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in. Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner ...