576 to 600 of 4,890 Permanent Observability Jobs

Remote Head of DevOps

Hiring Organisation
1inch
Location
Tranent, East Lothian, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Crieff, Perth & Kinross, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Ballyclare, Co. Antrim, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Enniskillen, Co. Fermanagh, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
South Shields, Tyne and Wear, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Software Engineer Backend

Location
West of England, England, United Kingdom
other runtime conditions. Design appropriate controls for authentication, authorization, permissions, and secure access to tools, models, services, and data. Design for scalability, reliability, performance, observability, and fault tolerance across services. Write high-quality, readable, testable code that follows established engineering standards and best practices. Participate in code and design reviews … strong engineering and quality practices across the team. Troubleshoot and resolve complex issues across backend services, workflows and integrations. Build monitoring, logging, testing, and observability into production services to support reliable agent execution and product operations. Collaborate with DevOps, and infrastructure teams to support reliable deployment and operation. Stay current ...

Remote Head of DevOps

Hiring Organisation
grabjobs
Location
Ashton-Under-Lyne, Greater Manchester, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Analytics Services Platform Engineer

Location
Greater London, England, United Kingdom
developer experience Collaborating with research, data and engineering teams to accelerate time‐to‐insight through modern analytics solutions Driving improvements in automation, observability and resilience across analytics services Evaluating and adopting emerging technologies such as AI assistants, data mesh and cloud‐native analytics solutions Defining SLAs, KPIs and monitoring strategies … using Terraform or Ansible Deep understanding of AWS analytics technologies including EMR, MSK, Athena, Redshift, Glue and MWAA Experience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetry Strong problem‐solving skills and a systematic approach to diagnosing and resolving issues Highly Desirable Skills ...

Senior Tech Consultant

Location
Leeds, England, United Kingdom
direction, identifying opportunities to improve a client’s product or technology beyond the immediate brief. Implement and improve CI/CD, cloud infrastructure and observability, primarily using GitHub Actions, AWS and Azure. Promote strong engineering practices around code quality, automated testing, peer review and security, while mentoring and supporting other … building production‐grade AI‐native products, integrating frontier models and designing reliable agentic systems, with a strong understanding of evaluation, context engineering, tool use, observability, security, latency, cost and non‐deterministic behaviour. Strong problem‐solving skills, with the ability to take complex or ambiguous problems and turn them into well ...

Senior DevOps & Cloud SRE Lead (AWS / K8s)

Location
Greater London, England, United Kingdom
production EKS/Kubernetes clusters, ingress controllers, and service meshes Build automated CI/CD pipelines using GitHub Actions, Docker, and Helm Set up observability stack (Prometheus, Grafana, Datadog) and incident management Requirements 4+ years in DevOps, Site Reliability Engineering, or Cloud Architecture Deep expertise in AWS, Kubernetes, Terraform, Docker ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer – Associate - London London · United Kingdom · Associate

Location
Greater London, England, United Kingdom
build, run and continuously improve this business-critical service. The role combines software engineering, systems engineering and production expertise to improve the reliability, scalability, observability, incident response and operational efficiency of the CTL platform. Required Skills/Experience Strong programming ability in one or more modern languages – Java … building maintainable automation beyond simple scripts. Good understanding of networking, messaging, distributed systems, data structures, algorithms and software design fundamentals. Hands-on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes ...

Analytics Services Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
usability, scalability and the developer experienceCollaborating with research, data and engineering teams to accelerate time-to-insight through modern analytics solutionsDriving improvements in automation, observability and resilience across analytics servicesEvaluating and adopting emerging technologies such as AI assistants, data mesh and cloud-native analytics solutionsDefining SLAs, KPIs and monitoring strategies … code, using Terraform or AnsibleDeep understanding of AWS analytics technologies including EMR, MSK, Athena, Redshift, Glue and MWAAExperience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetryStrong problem-solving skills and a systematic approach to diagnosing and resolving issuesHighly desirable skillsExperience with streaming frameworks ...

Applied AI ML Lead - Python & Agentic AI

Location
Auchentibber, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...

Software Engineer

Location
Greater London, England, United Kingdom
Python, C/C++, Go, Typescript, or Java. CI/CD and DevOps: Knowledge of building pipelines, automated testing and deployment strategies. Monitoring and observability: Familiarity with tools like Grafana, Prometheus or other observability platforms. Containerisation and orchestration: Docker, Kubernetes. We are open to a wide range of backgrounds ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Senior Platform Engineer

Location
Welwyn Garden City, England, United Kingdom
GitOps workflows to enable safe, fast, and repeatable delivery. Championing DevSecOps principles, embedding security and compliance into the software delivery lifecycle. Establishing and improving observability, monitoring, and incident response practices, including vulnerability management and remediation. Mentoring engineers and contributing to a strong engineering culture through knowledge sharing, documentation, and technical … would be great if you have the following Experience with Helm, Kustomize, and Kubernetes ecosystem tooling. Familiarity with Azure and Azure DevOps. Experience with observability platforms and Kubernetes policy enforcement tools. Proficiency in scripting or programming (e.g. Bash, Python, PowerShell, C#). Experience designing multi‐region or highly available systems. ...

Full Stack Engineer - Must be DV Cleared

Hiring Organisation
VIQU IT Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
platform engineers so models, pipelines and infrastructure land as one product rather than three. Improving the day to day: CI/CD, observability, security hardening, documentation. Sitting in front of customers and partners and turning what they need into something buildable. What You Will Need Python. Serious commercial experience building … secrets management. Putting AI or ML models into production, LLM features included. Infrastructure as code with Terraform, Ansible or similar, plus configuration management. Observability and incident response: metrics, logging, tracing. Message queues, event-driven architectures or data pipelines. Also Welcome Kubernetes. Deploying, running or debugging workloads, including K3s, RKE2 ...

Lead GCP Engineer

Location
City Of London, England, United Kingdom
practices across the platform. Drive adoption and maturity of CI/CD pipelines, Infrastructure as Code and automated deployment processes. Implement and improve platform observability, monitoring, alerting and operational support capabilities. Develop and maintain reusable automation patterns and engineering standards. Leverage Terraform and Infrastructure as Code principles to ensure repeatable … Infrastructure as Code expertise. Strong experience with CI/CD tooling and deployment automation. Experience building and operating resilient cloud platforms. Strong understanding of observability, monitoring, incident management and operational excellence. Data Engineering & Integration Experience designing secure data integration solutions. Knowledge of APIs, databases, data pipelines and cloud‐based data ...

Senior Platform Engineering Manager at Prolific

Location
United Kingdom
operational excellence, and innovation. Champion SRE Culture: Own availability and embed SRE principles across the organization, including defining SLOs, SLAs, error budgets, and enhancing observability and incident remediation. Platform & Developer Experience: Own the developer-facing platform — golden paths, self-service infrastructure, and internal tooling — so teams can provision, deploy … Infrastructure-as-Code (Terraform/Terragrunt, Crossplane), with GitOps workflows using tools like ArgoCD. Reliability & Architecture: Solid understanding of complex infrastructure and application architecture, observability principles, and incident management. Stability & Velocity: Experience balancing the need for platform stability and reliability with the goal of increasing developer productivity and velocity. Bridge ...

Principal Software Architect (UK)

Location
Greater London, England, United Kingdom
containerization frameworks (Kubernetes, Docker, Helm). Experience with infrastructure-as-code, configuration management (Ansible, Terraform), and network automation protocols (NetConf, YANG). Distributed Systems & Observability: Deep familiarity with modern distributed systems tooling, including RESTful APIs, gRPC, message brokers (Kafka, NATS), and data serialization formats (JSON, YAML). Demonstrated expertise … system observability, telemetry, and monitoring stacks (Prometheus, Grafana). High-Performance Networking: Knowledge of high‐performance packet processing techniques (e.g., DPDK, eBPF) and low‐latency systems tuning. Proficiency in network protocols and architecture, including TCP/IP, UDP, routing, and network performance optimization. Preferred Qualifications Master's degree in computer ...

AI Engineer Python

Location
Bracknell, England, United Kingdom
Terraform. Establish strong testing, code quality, documentation-as-code and secure software development practices. Improve application performance, scalability and reliability through profiling, benchmarking, observability and architectural improvements. Implement monitoring, alerting and operational standards for platform services. Conduct thorough code reviews, mentor engineers and raise engineering standards across the team. Partner … including automated testing, code reviews, CI/CD and release engineering. Practical experience with Docker, Kubernetes and Infrastructure as Code, ideally Terraform. Experience with observability, monitoring, alerting, performance profiling and benchmarking. Solid understanding of secure coding practices, vulnerability management and security considerations in software delivery. Applied understanding of Generative ...

Senior Manager- Software Engineering

Location
Greater London, England, United Kingdom
microservices, including REST and event‐driven patterns, and enforce best practices for versioning, contracts, and backward compatibility. Advance operational excellence by defining SLOs, improving observability and alerting, hardening on‐call procedures and runbooks, and leading incident response and post‐mortems. Solve complex distributed system challenges (such as throughput, latency, consistency … with microservices, API design, and event‐driven systems, as well as experience with containerization and orchestration (Docker, Kubernetes). Strong operational mindset: expert in observability, monitoring, incident response, performance engineering, and adherence to security best practices. Familiarity with CI/CD pipelines, automated testing strategies (unit, integration, e2e), and modern ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
grabjobs
Location
Bristol, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Glasgow, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
grabjobs
Location
Oakley, Hampshire, UK
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...