1,426 to 1,450 of 1,810 Permanent Observability Jobs

Azure Platform Engineering Consultant

Hiring Organisation
Morgan McKinley
Location
Newbury, Berkshire, England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
platform templates and landing zone patterns. CI/CD & Automation: Build and refine automated deployment pipelines, environment management, and release practices. Platform Quality: Embed observability (monitoring, logging, alerting), resilience, security, and FinOps principles directly into platform assets. Co-Delivery & Knowledge Transfer: Work closely alongside client engineering teams to pair, document … Core compute, networking, storage, identity, security, and platform services. Infrastructure as Code: Strong proficiency with Terraform AND Terragrunt using modular, reusable implementation patterns. DevOps & Observability: Strong experience with CI/CD tools (Azure DevOps/GitHub Actions) and monitoring stacks (Prometheus, Grafana, Azure Monitor, etc.). FinOps: Practical knowledge ...

AI DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
transitions from local/team execution to a centralised, hosted enterprise service. This person builds and manages the infrastructure, CI/CD, monitoring, and observability that keep the platform available, secure, and scalable across multiple teams and products. They automate deployments, maintain pipelines, and provide operational support, playing a critical … enterprise SLAs, security, and governance requirements. Key responsibilities Own AI-SDLC deployment, operations, and platform reliability (uptime, performance, incident response). Manage infrastructure, monitoring, observability, and operational/on-call support. Build and maintain CI/CD pipelines and automation for platform releases. Drive the shift to a centralised, hosted ...

Member of Technical Staff (AI Infrastructure Engineer)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical … workloads Experience developing APIs and managing distributed systems for both batch and real-time workloads Solid debugging and monitoring skills with expertise in observability tools for containerized environments Preferred Skills Experience with Kubernetes operators and custom controllers for ML workloads Advanced Slurm administration including multi-cluster federation and advanced scheduling ...

Permanent position: Senior Python Backend Engineer (AI Systems)

Hiring Organisation
Nicoll Curtin Technology
Location
United Kingdom
Employment Type
Permanent
Salary
GBP Annual
responsible for building and operating the infrastructure and workflows that allow AI features to execute end-to-end, with a strong focus on reliability, observability, latency, and user experience. Areas of responsibility include: Building production Back End systems that serve AI-powered product features Designing retrieval, inference, and orchestration workflows … workflows Tool calling and workflow orchestration LLM-based applications used by real users Vector databases and semantic search AI evaluation and monitoring frameworks Production observability and reliability engineering Distributed systems and high-throughput services End-to-end ownership of Back End systems in production Technical Environment Python Node.js SQL/ ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent
Salary
£75,000
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Consultant - AI Ops

Hiring Organisation
Consulting Point
Location
England, United Kingdom
closely with clients and multidisciplinary teams to assess, design, and improve cloud operating models, combining advisory capability with hands-on knowledge of IT operations, observability, service management, DevOps, and emerging AI technologies. Key Responsibilities Advise clients on cloud transformation strategy and the impact of cloud operating models on business performance … ideally in a process transformation or re-engineering environment Experience operating as an LSS Black Belt, preferably with certification Strong knowledge of IT operations, observability, and monitoring practices and tools Awareness of predictive and prescriptive analytics in an IT operations context Understanding of AI frameworks and AIOps or MLOps lifecycle ...

Senior Software Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£95,000
cloud Collaborating with engineers, designers, product managers within and outside of the team Implementing maintainable and well-tested solutions iteratively, keeping business impact and observability as a primary focus Understanding the wider context of the business and designing system architectures that meet short and long term business goals Your skillset … Algorithms, data structures Observability Web services, REST, HTTP Containers, cloud Testing, reliability, monitoring Strong knowledge in building and owning application end-to-end, from inception to maintenance, to retirement Designing systems that scale Expertise in some of our main programming languages - TypeScript, Java, Golang, Rust, Python Desirable Experience with Kubernetes ...

Java Developer with AI

Hiring Organisation
BC Forward
Location
Chicago, Illinois, United States
Employment Type
Permanent
Salary
USD 4,827 Annual
Prototype, validate, and productionize AI-enabled features with attention to latency, accuracy, cost, monitoring, fallback design, and reliability. Design for high availability, scalability, resiliency, observability, performance, security, privacy, and compliance. Integrate with enterprise systems, eventing platforms, data providers, communication providers, and vendor-managed services. Troubleshoot and resolve complex production issues … automation. Understanding of responsible and secure AI practices including data privacy, access control, guardrails, evaluation, monitoring, hallucination risk, human review, and auditability. Experience with observability and analytics tools such as Dynatrace, ELK Stack, or CloudWatch. Experience in agile environments with CI/CD, automated testing, code quality, deployment readiness ...

Senior Vice President, Senior Software Engineering, Developer Experience

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
build automation. Optimize developer workflows by streamlining build orchestration, dependency management, and CI/CD pipelines to accelerate software delivery while reducing friction. Drive observability and resilience by integrating telemetry, analytics, and automated remediation into build systems to proactively identify and resolve bottlenecks and failures. Champion secure software supply chain … Experience (DevEx), Platform Engineering, or Build/Release Engineering. Experience with artifact repositories (e.g., Artifactory, Nexus) and secure software supply chain practices. Knowledge of observability systems, telemetry frameworks, and performance optimization for build pipelines. Exposure to AI/ML-driven build acceleration and failure prediction tooling. Prior hands-on software ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
scalability. Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions. Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively. Lead cost optimisation: monitor spend, right‐size workloads, tune autoscaling, and improve infrastructure efficiency. Improve … operational experience (on‐prem and AWS EKS). Hands‐on experience defining and operating SLOs/SLIs, alerting, and incident workflows. Deep understanding of observability and telemetry (monitoring, logging, tracing). Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD. Scripting proficiency in Python, Bash ...

AI Ops Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
scalable across merchant payment use cases. You will also be accountable for the end‐to‐end engineering of GenAI and ML platforms, embedding governance, observability and operational resilience by design, hile enabling teams to deploy and run AI solutions with clarity, assurance and accountability at scale. To be successful … Bedrock for foundation models, agent orchestration patterns, Lambda and Step Functions, alongside demonstrated Python engineering capability and secure microservices and API design. AI governance, observability and cost optimisation, embedding governance by design through policy as code, alignment to model risk framework expectations, lifecycle traceability and audit‐ready evidence, supported ...

Network Automation Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self-service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with Network, Platform and Security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … protocols Experience with infrastructure-as-code and automation tools such as Ansible, Terraform and Jinja2 Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Desirable: Experience ...

Java Lead Software Engineer — Digital Markets Execution Technology, Execute

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
code written by others Leads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomes Establishes reliability goals and implements observability, resilience patterns, and operational readiness practices Leads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation … teams Preferred qualifications, capabilities, and skills Exposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace) Experience with observability stacks and resilience engineering for low‐latency/latency‐sensitive platforms Familiarity with Python Experience operating services in regulated environments with strong auditability and controls ...

Forward Deployed AI Engineer

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
enabled systems. You’ll bring deep expertise across modern full-stack technologies (.NET, Azure, SQL, React/Angular), along with experience in distributed systems, observability, and AI tooling such as LLMs, retrieval pipelines, and agentic workflows. Acting as a bridge between business and technology, you’ll work across product, data … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Principal Software Engineer (London)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
other teams shift faster and more reliably Taken AI and agentic systems from prototype to production, with a strong understanding of infrastructure, authentication, and observability Delivered on high‐stakes consulting engagements across multiple language paradigms, stacks, ecosystems, and client industries Built high‐quality, maintainable software collaboratively, incrementally, and through … legacy systems with short and long‐term business needs Led and delivered solutions to architecture‐level problems including scalability, security, reliability, performance, maintainability, and observability Facilitated alignment across technical and non‐technical stakeholders to move initiatives forward through ambiguity and complexity Provided mentorship and team support at scale, while sharing ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Head of Platform Engineering, you'll define and deliver robust platform services, establish the patterns for self-service infrastructure, CI/CD, and observability, and drive the integration of AI/ML capabilities across the organisation. You'll translate platform strategy into a coherent technical roadmap, setting standards that teams … these are sustained as the platform evolves. Defining and maintaining platform standards, patterns, and reusable components that drive consistency across teams. Leading improvements to observability, monitoring, and alerting capabilities across systems and services. Mentoring and developing engineers across the organisation, raising the collective standard of platform engineering practice. Identifying ...

Senior Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
foundational infrastructure layer behind ACRA. This is a deep infrastructure role. You will work across Kubernetes, Linux, networking, storage, service‐to‐service communication, observability, security boundaries, and production operations. You will be responsible for the systems that everything else depends on: clusters, networks, storage layers, ingress and egress paths, runtime … security of the whole platform. You should be comfortable operating close to the metal: debugging Kubernetes, understanding networking behaviour, reasoning about distributed storage, improving observability, and helping define the infrastructure patterns that ACRA will rely on as it scales. This is not a generic DevOps support role or internal ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real‐time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Context Plane Python Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end-to-end — from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands-on experience using enterprise-authorized AI-assisted software development ...

Senior Data Management Professional - Data Engineer - Commodities Data

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
improve performance, reliability, and maintainability. Design automated pipeline controls for validation, monitoring, schema change, exception handling, and data integrity. Develop workflow orchestration, alerting, observability, and remediation processes. Translate business and client needs into engineering‐ready requirements and scalable technical solutions. Partner with Engineering on platform evolution, architecture, tooling, system design … hands‐on experience with Python or similar programming/scripting languages. Experience with querying structured, semi‐structured, and unstructured datasets. Experience with workflow orchestration, observability, monitoring, alerting, and scalable architecture design. Ability to analyze, refactor, and modernize legacy systems. Strong understanding of data lifecycle management, data integration, data modelling, data ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
trading decisions. You will be supporting customer-facing rollouts with Product and Commercial teams, ensuring production reliability. Additionally, you will help improve engineering processes, observability, and system resilience as we scale. Behind every line of code and every dataset, there’s a team of curious, driven people who bring ideas … strong experience building, deploying and operating reliable, scalable, end-to-end production APIs and data pipelines for external clients strong understanding of cloud infrastructure, observability, monitoring, incident response, and operational best practices relevant technologies (we use): AWS incl. S3, ECS, API Gateway, Batch, Redshift, Airflow, GitHub Actions, Docker, Pulumi, Python ...

Lead Infrastructure Engineer - AWS Cloud Support Engineering

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
adherence to resiliency and security expectations Familiarity with working in a large distributed system across a range of technologies including compute, databases, messaging, observability, and telemetry Knowledge of incident, change, and problem management processes and the controls that govern them Understanding of data-driven decision making and a drive … working in a follow-the-sun or globally distributed on-call support model Familiarity with large-scale cloud migration or modernization initiatives Exposure to observability and telemetry tooling in complex distributed environments ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products ...

Senior Data Management Professional - Data Engineer - Commodities Data London, GBR Posted today

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
improve performance, reliability, and maintainability. Design automated pipeline controls for validation, monitoring, schema change, exception handling, and data integrity. Develop workflow orchestration, alerting, observability, and remediation processes. Translate business and client needs into engineering‐ready requirements and scalable technical solutions. Partner with Engineering on platform evolution, architecture, tooling, system design … hands‐on experience with Python or similar programming/scripting languages. Experience with querying structured, semi‐structured, and unstructured datasets. Experience with workflow orchestration, observability, monitoring, alerting, and scalable architecture design. Ability to analyze, refactor, and modernize legacy systems. Strong understanding of data lifecycle management, data integration, data modelling, data ...

Staff Machine Learning Ops Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
optimization. Set the technical direction for CI/CD for ML, embedding testing, validation, security, performance checks, and release confidence into deployment pipelines. Establish observability standards for ML systems, including model metrics, service health, alerts, drift detection, data quality, lineage, and business-impact monitoring. Lead the evolution of Preply … monitoring, performance benchmarking, and lifecycle management. Strong hands‐on experience with cloud platforms such as GCP or AWS, Kubernetes, distributed compute, CI/CD, observability, and infrastructure-as-code practices. Experience building enabling tools and platform capabilities for Applied Scientists, Data Scientists, and engineering teams. Strong technical judgment ...