1,076 to 1,100 of 4,503 Permanent Observability Jobs

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Senior Software Engineer: Agentic Development Enablement

Location
Greater London, England, United Kingdom
Claude Code and GitHub Copilot Design and implement practical guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Help define how controls should work consistently across local development environments and CI/CD pipelines Explore changes to the development environment … background as a software engineer Broad technical understanding across several of the following: developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end‐user application delivery Ability to work in ambiguous ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Senior Software Engineer - Backend

Location
Manchester, England, United Kingdom
caching and data-access strategies Building reliable transactional workflows Developing applications and services within AWS Deploying and operating containerised applications Improving monitoring, alerting and observability Exploring AI-assisted engineering and modern development tooling At Senior level, you’ll be expected to understand the wider system rather than only the individual … traffic or data volumes Performance optimisation, caching and latency reduction Designing for resilience and failure Containerisation and orchestration technologies such as Docker and Kubernetes Observability, monitoring and operating production systems Automated testing and modern engineering practices We don’t expect candidates to have worked with every technology in our stack. ...

Senior Platform Engineer

Location
Warminster, England, United Kingdom
Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This is not a pure cloud or greenfield platform role. … failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. Understanding of Distributed Systems in production. ...

Software Engineering Specialist

Location
Belfast City District, Northern Ireland, United Kingdom
identify dependency, migration, and simplification opportunities. Cloud Platform, Reliability & Operations · Ensure billing services are deployed and operated effectively in AWS and EKS, with strong observability, logging, alerting, and operational controls. · Lead production readiness, resilience planning, performance tuning, and root‐cause analysis for customer and revenue‐impacting incidents. · Improve CI/… microservices, and API‐first architectures. Cloud Platforms (AWS/EKS): Proven expertise deploying and operating cloud‐native applications in AWS and EKS, including containerisation, observability, resilience, and CI/CD practices. Technical Leadership & Delivery: Experience leading end‐to‐end engineering delivery, making architectural decisions, mentoring engineers, and driving operational excellence ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Dunfermline, Fife, UK
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Remote Data Engineering Manager

Location
Southampton, Hampshire, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Security & Network Engineer - 12 months Fixed term

Location
Maidenhead, England, United Kingdom
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

Senior Software Engineer, Full-Stack

Location
Greater London, England, United Kingdom
Write scripts and tooling to automate repetitive work, support data migrations, and keep the team moving quickly Instrument the systems you build with strong observability practices -logging, metrics, and tracing — to catch and resolve issues before they impact customers Benchmark and profile code to identify performance bottlenecks, and make evidence … ability to contribute there when needed Experience with, or solid understanding of, AWS Experience building or maintaining multi‐tenant, enterprise SaaS platforms Familiarity with observability tooling and practices (logging, metrics, tracing, alerting) Experience benchmarking and profiling code for performance Experience in the financial industry, particularly treasury, payments ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
whole, with a clear path to shaping platform-level decisions at Staff level. Embed DataOps best practices across the team: CI/CD, testing, observability, data drift monitoring, and incident response using tools like Grafana and Datadog. Collaborate with ML engineers and data scientists to productionise models and agentic tooling … systems and an eagerness to build expertise in stream processing technologies such as Apache Flink, Kafka, or Spark. (Nice to have: ClickHouse.) Familiarity with observability tooling such as Grafana or Datadog, and a good instinct for keeping systems healthy and well-monitored. Active engagement with the evolving AI/ ...

Senior DevOps / Platform Engineer - Autonomous Vulnerability Research (Harness Engineering)

Location
Greater London, England, United Kingdom
reach sanctioned targets.* Build CI/CD pipelines with integrated supply-chain security: SBOMs, image signing, artefact provenance, and automated policy gates.* Deliver comprehensive observability - logs, metrics, distributed traces, and cost telemetry - across long-running, non-deterministic agent workloads.* Build evidence-capture pipelines: immutable audit trails, artefact retention, and reproducible … egress control, network policy, and segmentation in cloud-native environments.* Strong CI/CD engineering skills and experience embedding security controls into delivery pipelines.* Observability expertise across logging, tracing, and metrics, including designing for auditability and evidence retention.* A security-first mindset with the judgement to balance researcher velocity against ...

Senior DevOps / Platform Engineer - Autonomous Vulnerability Research (Harness Engineering)

Location
City of Edinburgh, Scotland, United Kingdom
reach sanctioned targets.* Build CI/CD pipelines with integrated supply-chain security: SBOMs, image signing, artefact provenance, and automated policy gates.* Deliver comprehensive observability - logs, metrics, distributed traces, and cost telemetry - across long-running, non-deterministic agent workloads.* Build evidence-capture pipelines: immutable audit trails, artefact retention, and reproducible … egress control, network policy, and segmentation in cloud-native environments.* Strong CI/CD engineering skills and experience embedding security controls into delivery pipelines.* Observability expertise across logging, tracing, and metrics, including designing for auditability and evidence retention.* A security-first mindset with the judgement to balance researcher velocity against ...

DevOps Engineer - SC Cleared - Hybrid - Inside IR35

Hiring Organisation
VIQU IT Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£700.00 - £800.00 per day
Experience of working within a production environment, with strong on-prem Kubernetes/OpenShift deployments IAC using Terraform with CI/CD pipelines (Jenkins) Observability tools to include design and operate end to end logging, metrics, tracing, dashboards to include alerting systems using ELK, Splunk, Grafana Cloud platforms to include ...

AI Engineer ( Manager )

Location
Greater London, England, United Kingdom
where appropriate. Write evaluations, instrument agent behaviour, and ensure failure modes are visible and recoverable, using automated functional and non-functional tests, trace-level observability, red-team scenarios and measures for quality, groundedness, latency and cost. Build the interfaces through which finance users review, approve, override and evidence agent actions … agent frameworks, including embeddings, vector or graph retrieval, prompt and configuration management, and integration with enterprise data or services. Practical experience with evaluation, observability and LLMOps or AgentOps, including versioning, monitoring, release and rollback, and analysis of quality, latency and cost. Understanding of secure enterprise integration and sensitive-data handling ...

Senior DevOps / Platform Engineer - Autonomous Vulnerability Research (Harness Engineering)

Location
Greater Manchester, England, United Kingdom
reach sanctioned targets. Build CI/CD pipelines with integrated supply-chain security: SBOMs, image signing, artefact provenance, and automated policy gates. Deliver comprehensive observability - logs, metrics, distributed traces, and cost telemetry - across long-running, non-deterministic agent workloads. Build evidence-capture pipelines: immutable audit trails, artefact retention, and reproducible … egress control, network policy, and segmentation in cloud-native environments. Strong CI/CD engineering skills and experience embedding security controls into delivery pipelines. Observability expertise across logging, tracing, and metrics, including designing for auditability and evidence retention. A security-first mindset with the judgement to balance researcher velocity against ...

Lead Cloud Engineer

Location
Manchester, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Lead Cloud Engineer

Location
Leeds, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Senior DevOps / Platform Engineer - Autonomous Vulnerability Research (Harness Engineering)

Hiring Organisation
Lloyds Banking Group
Location
Edinburgh, UK
Employment Type
Full-time
reach sanctioned targets. Build CI/CD pipelines with integrated supply-chain security: SBOMs, image signing, artefact provenance, and automated policy gates. Deliver comprehensive observability - logs, metrics, distributed traces, and cost telemetry - across long-running, non-deterministic agent workloads. Build evidence-capture pipelines: immutable audit trails, artefact retention, and reproducible … egress control, network policy, and segmentation in cloud-native environments. Strong CI/CD engineering skills and experience embedding security controls into delivery pipelines. Observability expertise across logging, tracing, and metrics, including designing for auditability and evidence retention. A security-first mindset with the judgement to balance researcher velocity against ...

Lead Cloud Engineer

Location
Greater London, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Agentic AI & Commerce Architect

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
across LLM platforms, orchestration frameworks, vector stores, and agentic workbenches; design multi-model routing strategies to avoid single-provider lock-in Embed AI governance, observability, and guardrail design as first-class architectural concerns — audit logging, human-in-the-loop escalation, adversarial testing, and compliance controls — not retrofitted after deployment Commerce … retrieval architectures and embedding pipelines in production (semantic search, RAG) Experience designing AI governance frameworks: auditability, HITL, red teaming, compliance controls AgentOps/LLMOps: observability, tracing, and model evaluation in production Commerce stack integration: OMS, PIM, DAM, CRM, payments, and fulfilment systems Proficiency in Python, agent frameworks, vector databases, APIs ...

Integration Developer

Hiring Organisation
Robert Half Limited
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£70,000
/or AWS CDK Experience with RDS or DynamoDB Knowledge of CI/CD, Git-based workflows and DevOps practices Experience with monitoring and observability tools such as Grafana, Prometheus and CloudWatch Exposure to unit and integration testing Robert Half Ltd acts as an employment business for temporary positions ...