1,001 to 1,025 of 1,971 Observability Jobs

Engineering Lead

Hiring Organisation
Hays
Location
Cheshire, North West, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
Up to £500.0 per day + Inside IR35
guide implementation of scalable integration solutions across APIs, middleware, event-driven platforms, and external systems. Drive best practices across software engineering, DevOps, resiliency, observability, operational excellence, audit readiness, and governance. Conduct architecture reviews, code reviews, technical design assessments, and controls compliance reviews. Provide mentoring, coaching, and hands-on technical guidance … Services, or large-scale transformation programmes. Experience supporting regulated applications, risk technology platforms, or compliance-driven initiatives. Experience with Docker and Kubernetes. Knowledge of observability, monitoring, logging, and control monitoring frameworks. Experience working within Agile delivery environments. Experience supporting geographically distributed teams. Exposure to AI-enabled engineering tools, automation frameworks ...

Engineering Lead

Hiring Organisation
17918
Location
London, United Kingdom
guide implementation of scalable integration solutions across APIs, middleware, event-driven platforms, and external systems. Drive best practices across software engineering, DevOps, resiliency, observability, operational excellence, audit readiness, and governance. Conduct architecture reviews, code reviews, technical design assessments, and controls compliance reviews. Provide mentoring, coaching, and hands-on technical guidance … Services, or large-scale transformation programmes. Experience supporting regulated applications, risk technology platforms, or compliance-driven initiatives. Experience with Docker and Kubernetes. Knowledge of observability, monitoring, logging, and control monitoring frameworks. Experience working within Agile delivery environments. Experience supporting geographically distributed teams. Exposure to AI-enabled engineering tools, automation frameworks ...

Senior Platform Engineer

Hiring Organisation
Capgemini
Location
City and Borough of Birmingham, United Kingdom
Employment Type
Full Time
GitOps-driven platforms and internal developer platforms that give engineers true self-service. You’ll modernise legacy estates into cloud-native architectures, unify observability across complex environments using OpenTelemetry and modern APM platforms, and drive FinOps-focused optimisation. You’ll contribute to and help shape engineering standards, SRE-aligned operability … guardrails. • Lead technical workshops and design reviews with client stakeholders and engineering teams. • Set and evolve platform standards (security, reliability, operability, CI/CD, observability) and help teams adopt them in practice. • Define SLOs, improve reliability, and lead incident reviews. • Coach engineers through mentoring, pairing and pragmatic hands-on technical ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
operational efficiency. Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service … programming language, preferably Python, for automation and tool development. Tooling & Concepts CI/CD: Experience setting up and maintaining modern CI/CD pipelines. Observability: Practical experience implementing and managing monitoring and logging tools. Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud‐native networking within Kubernetes. ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
17918
Location
United Kingdom
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Senior Site Reliability Engineer - Python

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
operating and improving cloud-based platforms and services used across DBT and wider government. Working across Python development, cloud infrastructure, CI/CD pipelines, observability and automation, you'll help improve reliability, developer experience and service performance while supporting critical business services used by thousands of users. As a Senior … code approaches. Support teams to adopt Site Reliability Engineering practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. Contribute to observability across services, helping teams better understand performance, reliability and user impact. Develop and improve CI/CD pipelines to enable safe, frequent and low-risk ...

Machine Learning Systems & Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Ship workloads with Docker and Kubernetes; maintain IaC (Terraform) for the surfaces you own and CI/CD pipelines, including self‐hosted GPU runners. Observability and reliability: Monitoring, logging, and alerting for job performance, data‐pipeline health, and cost (e.g., Prometheus/Grafana, OpenTelemetry); define SLOs and incident response … stores; and object storage with caching layers. Familiarity with ML workflow orchestration and experiment tracking (e.g., Kubeflow Pipelines, MLflow). Experience with monitoring and observability tooling (e.g., Prometheus/Grafana, OpenTelemetry) and CI/CD for infra and ML workflows (e.g., GitHub Actions). At SpAItial, we are committed ...

Cloud Platform Engineer (Senior / Lead)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices—SLOs/SLIs, observability, capacity planning, resilience and blameless incident management—to keep the platform reliable and cost‐efficient. Partner with data engineering to design and optimise data pipelines, data … workloads and data—including secrets management, least‐privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost‐optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. ...

Senior Software Engineer - Agentic AI Platform

Hiring Organisation
Pinnacle Technical Resources
Location
Fort Worth, Texas, United States
Employment Type
Permanent
Salary
USD 60 Annual
maintain agent orchestration services, tool registries, and execution runtimes. Build APIs and microservices for LLM integration, prompt management, and agent lifecycle management. Implement observability, logging, and monitoring for agentic workflows. Write comprehensive tests (unit, integration, end-to-end) to ensure platform reliability. Collaborate with architects on design decisions … like LangChain, LangGraph, or similar orchestration tools. Nice to Have Skills: Experience with event-driven architectures (Kafka, EventBridge). Knowledge of vector databases and observability tools (Datadog, Splunk, OpenTelemetry). Familiarity with agent evaluation frameworks, FastAPI/Flask, async Python, and Redis/caching patterns. Experience in the airline ...

.NET Solution Architect

Hiring Organisation
HTC Global Services Inc
Location
Miami, Florida, United States
Employment Type
Permanent
Salary
USD Annual
tools. Understanding of Responsible AI and AI security principles. Preferred Skills TOGAF or cloud certifications preferred. Experience with AI-assisted development tools. Experience with observability and monitoring tools such as: Datadog Dynatrace Knowledge of AIOps and automated operational frameworks. Experience in ZeroOps/self-healing platform implementations. Education Bachelor … Exposure to GenAI-assisted SDLC workflows. Experience in platform engineering and reusable accelerator frameworks. Knowledge of enterprise integration ecosystems. Experience with AI-driven monitoring, observability, and self-healing platforms. Soft Skills Excellent communication and stakeholder management skills. Strong problem-solving and analytical abilities. Ability to lead distributed teams. Strong presentation ...

Senior Reliability Engineer

Hiring Organisation
Fitch Group
Location
Greater London, United Kingdom
Employment Type
Full Time
someone who is curious about the evolving role of AI in infrastructure engineering, someone who actively explores how AI-assisted tooling, automation, and intelligent observability can raise the bar for reliability and developer experience. You will collaborate closely with global development and engineering teams to deliver reliable, resilient, and high … efficiency Identify, contain, and mitigate risk across all cloud environments, maintaining a robust security posture for infrastructure and applications Implement proactive monitoring and observability practices to detect and prevent issues before they impact users Develop and maintain automation and tooling solutions, including AI-assisted approaches to reduce toil and accelerate ...

Forward Deployed Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
plus.* Hands-on experience with cloud platforms (AWS), Docker and Kubernetes; CI/CD in GitLab (or equivalent); infrastructure-as-code (e.g. Terraform); observability and monitoring stacks.* Solid understanding of database systems, SQL, data modelling, ETL pipelines, REST/gRPC APIs and microservices architecture; identity and access management with OIDC …/SAML and Azure Entra ID.* Experience with LLMs, RAG systems, prompt engineering and AI evaluation frameworks; familiarity with MLOps, model deployment and AI observability/guardrails; working use of AI-assisted development tools (e.g. Claude Code).* Familiarity with one or more of: trading, ERP or treasury platforms; workflow ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
operational efficiency.* Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment.* Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification.* Reliability & Performance: Define, measure, and enforce Service … programming language, preferably Python, for automation and tool development.**Tooling & Concepts*** CI/CD: Experience setting up and maintaining modern CI/CD pipelines.* Observability: Practical experience implementing and managing monitoring and logging tools.* Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud-native networking within Kubernetes. ...

Java Developer - Security & Intelligence

Hiring Organisation
Jobleads-UK
Location
Gloucester, England, United Kingdom
solutions meet stringent performance, reliability and security requirements. Contribute to architectural design, code reviews, automated testing and continuous improvement activities. Implement logging, monitoring and observability to support the operation of production systems. Collaborate closely with multi‐disciplinary teams including engineers, data specialists and operational stakeholders. Skills Required Strong experience developing … backend systems for data‐intensive or mission‐critical applications. Experience working in DevSecOps environments, including Docker, Kubernetes, CI/CD pipelines, automated testing and observability tooling. Ability to work directly with stakeholders to understand requirements and translate them into robust, secure and scalable software solutions. Strong collaboration and communication skills ...

Java Developer - Security & Intelligence

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
solutions meet stringent performance, reliability and security requirements. Contribute to architectural design, code reviews, automated testing and continuous improvement activities. Implement logging, monitoring and observability to support the operation of production systems. Collaborate closely with multi‐disciplinary teams including engineers, data specialists and operational stakeholders. Skills Required Strong experience developing … backend systems for data‐intensive or mission‐critical applications. Experience working in DevSecOps environments, including Docker, Kubernetes, CI/CD pipelines, automated testing and observability tooling. Ability to work directly with stakeholders to understand requirements and translate them into robust, secure and scalable software solutions. Strong collaboration and communication skills ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
productivity and experience across the SDLC Cost optimise infra through rightsizing and creating episodic, on-demand environments Proactively monitor for security and reliability using Observability tooling including SIEM, APM, tracing, infrastructure metrics, logs and dashboards Durably engineer away toil You will be a great fit here if you: Are passionate … infrastructure continuously using CI/CD tools such as GitlabCI, CircleCI, Github Actions, and GitOps using ArgoCD, FluxCD Troubleshooting and debugging applications using Observability tooling across microservices and serverless applications such as Splunk, DataDog Managing ephemer secrets and credentials using Hashicorp Vault Managing least privileged access to cloud resources using ...

Sr. Distinguished Machine Learning Engineer (Remote-Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
Sr. Distinguished Machine Learning Engineer (Remote-Eligible) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine ...

Founding Lead SRE - Remote, Equity

Hiring Organisation
Jobleads-UK
Location
United Kingdom
company. You’ll be a technical and operational leader for production services, mentoring others and setting on-call standards. You’ll own disaster recovery, observability, and incident response while working on our Cloud Application Platform, Kubernetes on AWS, and interconnected engineering teams. #J-18808-Ljbffr ...

Senior Cloud Platform Architect - IaC & FinOps Leader

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deliver speed, security, cost efficiency, and reliability. You will mentor engineers, drive policy‐based governance, and collaborate with SRE and Security to ensure observability and compliance from day one. #J-18808-Ljbffr ...