1,301 to 1,325 of 1,896 Observability Jobs

AI Platform/ DevOps Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, Castle Baynard, United Kingdom
Employment Type
Permanent
Salary
£70000 - £100000/annum + Benefits
Bedrock Knowledge Bases) and embedding pipelines Build and maintain CI/CD pipelines for inference services, retrievers, ingestion workflows, and RAG components Implement observability across AI workloads using CloudWatch, MLflow, and OpenTelemetry - covering latency, throughput, cost, and system health Apply secure-by-design principles including IAM, encryption, network controls … Terraform experience for infrastructure-as-code, provisioning and managing cloud infrastructure at scale Experience operating containerised services, managing CI/CD pipelines, and owning observability and reliability Familiarity with vector databases or search infrastructure (OpenSearch, Algolia) is a strong advantage Python proficiency for scripting, automation, and deploying production services Solid ...

Senior Tech Lead - FinTech

Hiring Organisation
Carousel Consultancy Ltd
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Shaping the platform architecture Working closely with the tech team to evolve the platform Designing scalable backend systems and services Improving reliability, performance and observability Helping modernise legacy parts of the platform Enhancing the platform and DevOps - collaborating with AWS infrastructure and cloud-native services, refining CI/CD pipelines … Native Infrastructure: AWS (ECS, EKS, RDS, S3, Lambda) Containers and Orchestration: Docker, Kubernetes CI/CD: Jenkins, GitHub Actions Databases: MySQL, PostgreSQL Monitoring and Observability: Sentry, CloudWatch, Grafana Skills and experience required: Solid software engineering experience (c8+ years), working as a Senior, Staff, Principal Engineer or Tech Lead FinTech ...

Corporate KYC : Principle Software Engineer - Executive Director

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Site Reliability Engineer (DV Security Clearance)

Hiring Organisation
CGI
Location
Manchester, United Kingdom
Employment Type
Full Time
Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production environments. The successful candidate will work closely with engineering, platform … services across cloud and containerised environments. - Manage and support Kubernetes clusters and Helm-based deployments across multiple environments. - Enhance monitoring, alerting, logging, and observability solutions to improve operational visibility and system reliability. - Investigate incidents, analyse logs, identify root causes, and drive timely resolution of production issues. - Participate in incident response ...

Platform Engineer

Hiring Organisation
CGI
Location
Newry Mourne and Down, United Kingdom
Employment Type
Full Time
take ownership of designing and evolving resilient digital platforms that accelerate delivery and power mission-critical services. Working across cloud infrastructure, automation, security, and observability, you will help set engineering standards and drive measurable improvements in reliability, performance, and cost efficiency. Within a collaborative, multidisciplinary environment, you will have … Code standards and reusable modules Build & Optimise CI/CD pipelines for efficient, secure delivery Enable & Support containerised workloads and orchestration platforms Implement & Improve observability, logging, metrics, and alerting standards Secure & Govern platforms in line with enterprise security frameworks Troubleshoot & Resolve critical incidents, strengthening resilience and diagnostics Optimise & Control cloud ...

Technical Release Manager

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
functional teams throughout the development lifecycle, maintaining the release calendar, championing engineering and DevOps best practice, and continuously raising the bar on deployment quality, observability and release velocity across both AWS and Azure. A little about you... Proven experience managing software releases within a SaaS environment, combined with strong project … services and configuration management. Familiarity with containers and orchestration and modern deployment patterns to support zero‐down‐time releases. Understanding of monitoring, logging and observability tooling and their role in validating release health. Awareness of change management, release governance and security/compliance considerations within a SaaS delivery context. Proactive ...

Head of Infrastructure & Security

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automation strategy, partnering with Product and Engineering to design delivery pipelines that drive growth and feature velocity. Measure system performance using clear KPIs and observability dashboards. Engineering Enablement & Reliability (approx. 15%): Partner with tech leaders to streamline deployment processes across the development lifecycle. Drive incident response planning and Service Level …/GCP ecosystems and Kubernetes orchestration. You must be an expert systems administrator with hands‐on experience in Infrastructure as Code (Terraform) and modern observability tools (Datadog, Prometheus). Problem‐Solving: You are a data‐driven strategist who can devise high‐level strategy and is equally comfortable rolling up your ...

DevOps Engineer

Hiring Organisation
Big Red Recruitment Midlands Limited
Location
Coalville, Stanton under Bardon, Leicestershire, United Kingdom
Employment Type
Permanent
Salary
£59999 - £65000/annum £60,000 - £65,000
/CD pipelines to enable faster, safer software delivery. Embedding DevSecOps principles and security tooling throughout the development lifecycle. Improving platform resilience, monitoring, observability and operational performance. Driving automation and adopting AI-assisted engineering to improve efficiency. Working closely with software engineers, infrastructure specialists and support teams to continuously improve … advantageous) CI/CD pipelines and deployment automation Docker, containers and Linux DevSecOps, SAST/DAST and cloud security best practices Monitoring, logging and observability platforms Agile software delivery environments Why join? Build technology that has a genuine positive impact on society. Join a business investing heavily in cloud, automation ...

Systems Engineer, Production

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
scalable solutions to thousands of customers every day. Our mission is to make deploying and operating software effortless and safe. We focus on automation, observability, and reliability, ensuring every engineering team at the company can move faster and with confidence. You’ll be part of a globally distributed team, collaborating … Code using Terraform, ensuring reproducibility and compliance. Collaborate with developers to improve CI/CD pipelines, deployment strategies, and overall developer experience. Enhance observability and reliability, refining alerting, monitoring, and incident response. Collaborate on cloud optimization projects, improving performance, cost efficiency, and security posture. Mentor and guide team members, fostering ...

Senior Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
health: tech debt, refactoring, security, and governance Take ownership of features from start to finish using agile methodologies, delivering with feature flags, tests, and observability Champion high‐quality technical communications: proposals, specs, testing reports, and release planning Drive AI‐first ways of working within the squad — embedding AI tooling into … relates to the Manage domain Contribute to CI/CD pipeline improvements and progressive delivery practices across squads Drive reliability monitoring and observability within Manage (Prometheus, Grafana, Sentry) Contribute to security posture improvements: vulnerability scanning, pen testing coordination, and enforcement of standards Cross‐Squad & Leadership Collaboration Work closely with ...

DV Cleared DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Malvern, England, United Kingdom
environment provisioning, configuration management, and deployment processes. Collaborate with cross-disciplinary teams to ensure system reliability and security. Support incident investigations and contribute to observability strategies involving metrics, logs, and tracing. Lead or contribute to architectural discussions around system scalability, security, and best practices. Essential Skills & Qualifications Proven expertise with … including IaC, environment parity, and automation strategies. Proficiency in scripting languages such as Python for automation and tooling. Demonstrable experience with system monitoring and observability platforms. Solid understanding of security automation practices and experience working with modern infrastructure patterns. Active SC clearance (Security Check) is mandatory for this role. Excellent ...

AWS API Engineer / Data Architect

Hiring Organisation
Capgemini
Location
City and Borough of Birmingham, United Kingdom
Employment Type
Full Time
will design RESTful APIs and service boundaries, define integration standards, and collaborate with product, delivery and cybersecurity stakeholders to ensure secure access patterns, strong observability, and smooth deployments and cutovers. Where required, you will also contribute to data architecture decisions (e.g., PostgreSQL and document stores) to enable robust … will be able to design secure, scalable and well-governed RESTful APIs and microservices on AWS, applying consistent standards for identity, access control, observability, auditability and operational resilience. You will be comfortable collaborating with stakeholders to translate outcomes into pragmatic architectures and delivery plans, and you will bring strong engineering ...

Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
/CD pipelines in BitBucket Pipelines to test and deploy infrastructure changes automatically and safely. Apply SRE principles by codifying reliability, self‐healing, and observability directly into automated solutions. Set coding and automation standards for the team and mentor colleagues while remaining hands‐on; identify manual, repetitive tasks and replace … automate platform migration and modernisation work across Azure, AWS, VMware, Citrix, storage services, and Office 365. Instrument systems and build automated monitoring and observability using tools such as Splunk, Grafana, and Opsgenie. Participate in on‐call rotations and incident response, automating detection and remediation to minimise downtime. Embed security ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

AI Consulting –Sr. AI Architect & Client Partner

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
cost‐efficient infrastructure that empowers engineering teams to deliver quickly without sacrificing reliability or compliance. You will be the driving force behind our DevOps, observability, and compliance readiness , ensuring our systems are audit‐ready, highly available, and optimized for both performance and cost. Key Responsibilities Architect, implement, and maintain cloud … incredible journey and learning a lot along the way. Requirements Technical stack : Azure (also AWS is a plus), Terraform, AKS (Kubernetes), Docker, GitHub Actions. Observability : Experience implementing logging, metrics, and tracing frameworks. Security : Familiarity with best practices, secrets management, and security scanning tools. Networking : Solid understanding of VPCs, private networking ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
17918
Location
United Kingdom
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...