601 to 625 of 1,390 Permanent Observability Jobs

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
workflows for Kubernetes‐based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application‐level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real‐world, not greenfield ...

Compute Platform Engineer II

Hiring Organisation
GSK
Location
Greater London, United Kingdom
Employment Type
Full Time
/CD-driven platform represents and enables the entire application and analysis lifecycle including interactive development and explorations (notebooks), large-scale batch processing, observability and production application deployments. A Compute Platform Engineer II is a technical contributor who can consistently take a poorly defined business or technical problem, work … Knowledge and use of at least one common programming language: e.g., Python, Go, C++, Scala, Java, including toolchains for documentation, testing, and operations/observability Expertise in modern software development tools/ways of working (e.g. git/GitHub, devops tools, metrics/monitoring, ...) Cloud expertise (e.g., AWS, Google ...

Senior Platform Engineer (Remote UK Only)

Hiring Organisation
Jobleads-UK
Location
Newcastle upon Tyne, England, United Kingdom
workflows for Kubernetes-based deployments (using Helm, Kustomize, ArgoCD, or similar) with automated guardrails to ensure fast, repeatable, and safe code delivery. Implement application observability: Set up application-level metrics, logging, and alerting within the namespaces, ensuring engineering teams have the visibility they need to monitor workload health. Create developer … Docker Compose in production environments Understanding of networking fundamentals Strong scripting ability: Bash and Python Experience with GitOps tooling: ArgoCD or Flux Experience with observability tooling (Prometheus, Grafana, Loki, Alertmanager or equivalent) Ability to think creatively within constraints and plan pragmatically around them: our stack is real-world, not greenfield ...

Global Banking & Markets - Software Engineer - Vice President - London

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
standard for years to come. What You Will Do Design, build, and operate high‐availability, multi‐region, cloud‐native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer. Develop event‐driven architectures, multi‐stage processing pipelines, and optimized data paths for high‐throughput … patterns (retry, dead‐letter queues, error isolation). Cloud & Infrastructure: Cloud platforms (GCP, AWS), container orchestration (Kubernetes, Docker), and JVM tuning for containerized workloads. Observability & Operations: Application instrumentation (metrics, distributed tracing, structured logging) and production support in high‐availability environments. Data & Performance: Data modeling, SQL/NoSQL databases, caching strategies ...

AIML Software Engineer, AI for Science

Hiring Organisation
GSK
Location
Greater London, United Kingdom
Employment Type
Full Time
Salary
136125 to 226875 USD Annually
infrastructure as code. Strong problem-solving and debugging skills, and experience working in cluster settings or cloud-based environments. Experience operating production services - monitoring, observability and alerting, and diagnosing and resolving issues in live systems. Experience designing and administering SQL databases - schema design, query performance, and day-to-day operational … including defining and working to service-level objectives (SLOs/SLIs). Experience with incident response and post-incident review, and with building the observability that supports it. Infrastructure-as-code (e.g. Terraform) for provisioning and maintaining cloud environments. Experience developing and administering workloads on Kubernetes (e.g. GKE). Familiarity ...

Grafana Observability Engineer (6-month contract)

Hiring Organisation
17918
Location
London, United Kingdom
Grafana Observability Engineer (6-month contract) Location: Fully remote Our client They are delivering a company-wide Digital Transformation (DX) Programme that will modernise their technology landscape and transform how they deliver services. As part of this journey, their IT and Portfolio Delivery teams are implementing a new enterprise technology … making this an exciting opportunity to join the business and help shape their future. Role Overview Reporting to the Lead Platform Engineer, the Senior Observability Engineer will own and develop their observability capability, leading the design, implementation and continuous improvement of their monitoring and alerting platform. Working closely with infrastructure ...

Senior Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Leeds, England, United Kingdom
Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated … implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like ...

Senior Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated … implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like ...

Senior Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated … implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like ...

Senior Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Implement the data pipelines, workflow orchestrations, and specialised compute footprints needed to support enterprise AI applications, using technologies like AWS Bedrock and AWS AgentCore. Observability & Reliability: Build robust monitoring and observability pipelines to ensure the health, performance, and security of distributed cloud applications and AI models. FinOps Standards: Embed automated … implementing components of data pipelines, including real-time streaming tools (AWS Kinesis, Kafka), data orchestration (dbt, Airflow), and managing vector databases for RAG architectures. Observability & Cost Management: SRE/Platform experience with the practical application of real-time monitoring and cloud cost optimisation using native CSP tools or utilities like ...

Senior Azure SRE: Cloud Reliability & Automation

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
enterprise-scale Azure platforms. The role focuses on reliability, automation, and governance to reduce operational overhead. The successful candidate will lead reliability engineering, implement observability, and drive resilient architecture with Terraform, Azure DevOps, and strong SRE practices. UK-based, with flexible same-team collaboration across regions. #J-18808-Ljbffr ...

Multi-Cloud SaaS Release Lead (CI/CD & DevOps)

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
coordinating multiple teams to minimize risk and maximize cadence. The role requires strong DevOps practices, containerisation knowledge, and the ability to drive automation and observability across a SaaS platform. #J-18808-Ljbffr ...

Senior Platform Engineer: Kubernetes & CI/CD Lead

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
modernise deployment and scale workloads. You will drive Kubernetes standardisation, containerisation of existing workloads, and the development of self-service runbooks, while owning observability #J-18808-Ljbffr ...

Senior Platform Engineer — SRE & Cloud

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
response, and mentor engineers while partnering with product and leadership to deliver scalable services. You will own platform roadmap, advance Kubernetes, Terraform, GitOps, and observability, and champion cost-aware engineering across teams. Remote options not specified. #J-18808-Ljbffr ...

Senior SRE — Front-Office Trading Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
J.P. Morgan is seeking a Lead Site Reliability Engineer within the Trading Technology group to shape SRE patterns and observability across globally distributed systems. You will work directly with traders and software engineers to improve reliability, performance, and resilience in front‐office environments. You will design and implement automated remediation ...

Senior Backend Engineer – Platform Modernisation (Java/.NET)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
APIs. You will design, build, and optimize backend services, collaborate with product, operations and vendors, and champion automated testing, CI/CD and observability to ensure reliable, scalable platform delivery. This critical role requires strong Java/.NET skills and experience in financial services environments. #J-18808-Ljbffr ...

Senior MLOps & Data Platform Engineer

Hiring Organisation
Jobleads-UK
Location
England, United Kingdom
develop cloud-native data pipelines across AWS and GCP. The role focuses on turning research into reliable production systems with strong emphasis on observability, governance and collaboration with cross-functional teams. The position offers a hybrid working pattern with a competitive salary, bonus, equity and benefits, and opportunities to drive ...

Cloud Platform Engineer — AWS, Kubernetes & CI/CD

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
pipelines, while building internal tooling to improve reliability and developer productivity. You will work closely with the engineering team to evolve the platform, curating observability and automation to reduce firefighting and improve deployment outcomes. This role emphasizes pragmatic delivery and high standards. #J-18808-Ljbffr ...

Lead Compute Platform Engineer – Cloud & HPC

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
medicine research. The Sr. Compute Platform Engineer will design, build, and operate tools and workflows across on‐prem and cloud, with emphasis on automation, observability, and scalable batch processing. You will mentor teammates and set software engineering best practices. You will collaborate with science users to optimize performance on large ...

Azure Platform Engineer - Remote, Terraform & Security Focus

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
emphasis on automation and reliability. You will design, build, and operate Azure-based infrastructure, create reusable Terraform modules, enforce security controls, and contribute to observability and SRE practices within a collaborative #J-18808-Ljbffr ...

Solutions Architect

Hiring Organisation
Anson McCade
Location
England, United Kingdom
implement platform governance, security and RBAC. Deliver platform maturity assessments and future operating models. Develop Infrastructure as Code using Terraform. Drive automation, monitoring, observability and FinOps best practices. Provide technical leadership across architecture, engineering and platform operations. Essential Experience Proven experience designing and delivering enterprise-scale Snowflake platforms. Strong knowledge ...

Senior Consultant VMware

Hiring Organisation
COMPUTACENTER (UK) LIMITED
Location
South East London, London, United Kingdom
Employment Type
Permanent
design/build/operate experience Strong NSX (T0/T1, DFW, VRFs, EVPN), vSAN (ESA/OSA) Automation with Terraform, Ansible, PowerCLI, APIs Observability with Aria Ops/Ops for Networks/Logs Migration experience with HCX Strong communication, documentation, and stakeholder engagement Preferred Skills Kubernetes on vSphere ...

Staff Data Engineer | London | Con...

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Apache Airflow Cloud‐native data engineering architectures Large‐scale data pipeline modernisation Building reusable engineering patterns within enterprise environments Data testing, validation and observability AI‐assisted software development using modern coding assistants Experience working within regulated or highly governed organisations Contract £750–£900 per day (inclusive) Initial 6‐month contract ...

Senior Azure AI Platform Engineer | Azure AI Foundry | C# | APIs

Hiring Organisation
Intelligent Talent Solutions Ltd
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
Servers Model Context Protocol Semantic Kernel Prompt Flow Agentic AI AI orchestration frameworks Copilot Studio Power Platform integrations RAG architectures AI governance and observability Technology Stack Azure AI Foundry | Azure OpenAI | C# | .NET | Azure Functions | Azure API Management | Azure Service Bus | Azure Logic Apps | MCP Servers | Semantic Kernel | Copilot Studio ...

Senior AI Engineer - Agentic & Generative AI

Hiring Organisation
The Portfolio Group
Location
Manchester, United Kingdom
Employment Type
Permanent
Salary
£80000 - £100000/annum
AutoGen or similar. Experience with Microsoft Azure, including Azure AI Foundry, Azure OpenAI Service and Azure AI Search. Understanding of LLM evaluation, AI safety, observability and production deployment. Why Join? Build cutting-edge AI products for a global SaaS customer base. Shape the future direction of Agentic AI within ...