1,326 to 1,350 of 4,709 Permanent Observability Jobs

AI Native SW Engineering

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … Vertex AI, and open-source models with fallback routing, token, cost, and latency management ImplementLLMOpsin production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions workshops, proofs ...

AI Native SW Eng

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … Vertex AI, and open-source models with fallback routing, token, cost, and latency management ImplementLLMOpsin production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions workshops, proofs ...

Software Engineer - Cloud Compute Platform

Location
Hursley, England, United Kingdom
Contribute To Include Workload orchestration across Kubernetes clusters Platform APIs and Kubernetes operators Cloud platform integrations Multi-tenant workload isolation and security Observability, health checks, and operational tooling Workload scheduling, disruption management, and rolling updates What You Will Do Design and implement services, APIs, and Kubernetes controllers using Go. Contribute … engineers to understand requirements and deliver reliable solutions. Participate in technical design discussions and help evaluate implementation trade-offs. Improve the reliability, scalability, observability, and maintainability of existing systems. Write automated tests, documentation, and operational runbooks. Participate in code reviews and provide constructive feedback to teammates. Help investigate and resolve ...

AI Native SW Engineering

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
govern production-grade agentic systems at enterprise scale: multi-agent orchestration across complex environments, RAG pipelines, policy-based routing, memory management, andprogramme-level lifecycle observability Define RAG pipeline standards across engagements:establishchunking and embedding strategies, set quality benchmarks, and ensure metric-backed tradeoff decisions are documented and transferable Set multi … default, fallbackroutingand cost governance as standard design practice across providers including OpenAI, Anthropic, Vertex AI, and open-source models OwnLLMOpsatprogrammescale: eval strategy, prompt governance, observability tooling standards, safetymonitoringand cost controls across multiple concurrent systems Lead client engineering engagements at senior level facilitatearchitecture design sessions, lead proof-of-concept delivery ...

Platform Engineer

Location
Milton Keynes, England, United Kingdom
Terraform, CloudFormation, or CDK), CI/CD pipelines, API Gateway, Lambda, and Aurora PostgreSQL — and confident applying fundamentals such as high availability, fault tolerance, observability, and cost control in a live environment. Payments or fintech exposure is an advantage but not required; what matters more is a practical, first-principles … rollback processes Support deployment of AI generated applications and tooling Support change control and release management alongside Engineering and the Information Security Officer Reliability, Observability & Data Infrastructure Design and maintain systems for high availability, fault tolerance, and resilience by default Implement logging, monitoring, and alerting (for example CloudWatch) across services ...

AI Platform & Site Reliability Engineering Managing Consultant

Location
United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services. You will work with technology, engineering, operations … operational requirements. AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. Reliability Engineering & SRE: Establish SRE practices including ...

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Workday Integrations Product Engineer

Location
Hook, England, United Kingdom
ecosystem supporting our global workforce. This is a hands‐on engineering role with a strong focus on Workday integration development, architecture, automation, reliability, security, observability, and operational excellence. Your Responsibilities: Design, develop, test, deploy, and support robust integrations between Workday and enterprise applications using Workday Studio, EIB, Core Connectors, Cloud … technology ecosystem. Apply strong software engineering principles including version control, peer reviews, automated testing, CI/CD, and structured release management. Build integrations with observability, operational resilience, proactive monitoring, alerting, logging, automated recovery, and self‐healing error handling by design. Troubleshoot complex production issues, conduct root‐cause analysis, and continuously ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services.You will work with technology, engineering, operations … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices including ...

Full Stack Engineer - Specialist

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
stack with a strong Java backend emphasis, working within agile teams that own the full product lifecycle from design and build through to deployment, observability, and iteration. We are particularly interested in candidates holding aMasters in Computer Science or Artificial Intelligence, ideally with some industrial placement or internship experience … ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into ...

Software Engineer - Cloud Compute Platform

Location
Greater London, England, United Kingdom
contribute to include: Workload orchestration across Kubernetes clusters Platform APIs and Kubernetes operators Cloud platform integrations Multi-tenant workload isolation and security Observability, health checks, and operational tooling Workload scheduling, disruption management, and rolling updates What You Will Do: Design and implement services, APIs, and Kubernetes controllers using Go. Contribute … engineers to understand requirements and deliver reliable solutions. Participate in technical design discussions and help evaluate implementation trade-offs. Improve the reliability, scalability, observability, and maintainability of existing systems. Write automated tests, documentation, and operational runbooks. Participate in code reviews and provide constructive feedback to teammates. Help investigate and resolve ...

Oracle OSS Stack Lead

Location
Newbury, England, United Kingdom
monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk, and enterprise monitoring platforms. Act as the senior technical escalation point for business stakeholders, support teams … Experience with Unix/Linux administration and Oracle SQL performance troubleshooting. Knowledge of cloud technologies, containerisation, and Kubernetes environments is advantageous. Familiarity with monitoring, observability, and enterprise operations tooling. Understanding of telecommunications network integrations and Oracle Fusion Middleware integration concepts is beneficial. Strong knowledge of Incident, Problem, Change, and Release ...

Backend Engineer

Location
Manchester, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
London, United Kingdom
Employment Type
Full-Time
Salary
£65,000 - £70,000 per annum
automate deployments using Terraform and Ansible, and build CI/CD pipelines using GitHub Actions. You'll also work across cloud security, networking and observability, while having the opportunity to develop your technical expertise through training and professional certifications. This role would suit an experienced DevOps Engineer looking to work … private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join ...

Network Automation & OSS Designer

Location
Greater London, England, United Kingdom
designs that span multiple network domains. Key Responsibilities/Job Description Design end-to-end network automation solution architectures spanning orchestration, inventory, compliance, and observability layers. Develop reference architectures and detailed solution blueprints for multi-domain network automation (IP/MPLS, optical, mobile RAN/Core, SD-WAN, cloud). … Terraform, Ansible, Helm) for network resource provisioning. Design Kafka-based event streaming and messaging architectures for real‐time network telemetry and automation triggers. Define observability strategies covering metrics, logs, traces, and network telemetry pipelines. Architect AIOps capabilities including closed‐loop automation, anomaly detection, and predictive analytics for network operations. Integrate ...

Backend Engineer

Location
West of England, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

Backend Engineer

Location
City of Edinburgh, Scotland, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

Backend Engineer 1861

Location
Leeds, England, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
Camden, London, Camden Town, United Kingdom
Employment Type
Permanent
Salary
£65000 - £70000/annum + Remote + Progression
automate deployments using Terraform and Ansible, and build CI/CD pipelines using GitHub Actions. You'll also work across cloud security, networking and observability, while having the opportunity to develop your technical expertise through training and professional certifications. This role would suit an experienced DevOps Engineer looking to work … private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
capabilities, reusable golden paths, and standardised service templates for efficient software delivery. Provide technical direction across platform domains including compute, Kubernetes, networking, secrets management, observability, CI/CD, and infrastructure as code. Enable AI‐assisted engineering workflows to enhance productivity, automate routine activities, and improve decision‐making. Embed governance, security … compliance controls through policy‐as‐code and auditable platform practices. Improve platform reliability through SLOs, observability frameworks, resilience engineering, and incident‐driven improvements. Collaborate with product, engineering, security, and architecture teams to translate business needs into scalable platform solutions. Promote engineering excellence through automation, standardisation, and efficiency‐driven practices, including ...

platform engineer for legal AI

Location
Greater London, England, United Kingdom
environments; Implement and evolve Infrastructure-as-Code using Terraform; Collaborate with developers to improve CI/CD pipelines, deployment strategies, and developer experience; Enhance observability and reliability through alerting, monitoring, and incident response; Collaborate on cloud optimization projects to improve performance, cost efficiency, and security posture; Mentor and guide team … management; Experience with CI/CD systems such as Buildkite and GitHub Actions; Experience with relational or non-relational datastores in production; Familiarity with observability platforms such as Datadog, Prometheus, and Grafana; Familiarity with Linux systems administration, networking, and troubleshooting; Excellent communication and documentation abilities, with a focus on knowledge ...

Principal DevSecOps Engineer

Hiring Organisation
83zero Limited
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
workflows * Establish secure-by-design engineering practices and enforce security and technical standards * Lead Infrastructure as Code (IaC) practices across teams and environments * Drive observability, monitoring, logging and audit controls * Support incident response, patching, compliance reporting and technical debt remediation * Partner with developers and delivery teams to improve engineering quality … Security & compliance - Trivy, vulnerability management, HashiCorp Vault, cert-manager * Containers & cloud - Docker, AWS EKS, AWS IAM, S3 and network policies * Infrastructure as Code - Terraform * Observability - Grafana, Loki * Automation - Python and Bash * Experience delivering within the UK Government Digital Service (GDS) lifecycle on a public sector engagement Why join ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
scalability, and performance. Develop and maintain automation workflows using Ansible, Salt, and related frameworks to reduce operational toil. Build and operate monitoring, alerting, and observability dashboards using tools such as Grafana and Splunk. Proactively identify network bottlenecks, performance issues, and reliability risks, implementing long‐term fixes rather than reactive solutions. … network protocols (BGP, OSPF, EIGRP, STP, VXLAN, VPNs, QoS, MPLS, etc.). Strong experience with infrastructure automation using Ansible and Salt. Proficiency with observability tooling such as Grafana, Splunk, or equivalent. Solid understanding of SRE practices including SLIs, SLOs, error budgets, and proactive reliability. Strong troubleshooting, analytical, and performance optimization ...

Oracle OSS Stack Lead

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk, and enterprise monitoring platforms. Act as the senior technical escalation point for business stakeholders, support teams … Experience with Unix/Linux administration and Oracle SQL performance troubleshooting. Knowledge of cloud technologies, containerisation, and Kubernetes environments is advantageous. Familiarity with monitoring, observability, and enterprise operations tooling. Understanding of telecommunications network integrations and Oracle Fusion Middleware integration concepts is beneficial. Strong knowledge of Incident, Problem, Change, and Release ...

AWS Data Engineer

Location
Greater London, England, United Kingdom
version control, CI/CD pipelines, automated testing, and modular code principles. Participate in pair programming, code reviews, and architectural design sessions. Data Quality & Observability Build monitoring and observability into data flows. Implement basic data quality checks and contribute to continuous improvements. Stakeholder Collaboration Work closely with business teams ...