1,526 to 1,550 of 1,915 Observability Jobs

Principal Software Engineer (London)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
other teams shift faster and more reliably Taken AI and agentic systems from prototype to production, with a strong understanding of infrastructure, authentication, and observability Delivered on high‐stakes consulting engagements across multiple language paradigms, stacks, ecosystems, and client industries Built high‐quality, maintainable software collaboratively, incrementally, and through … legacy systems with short and long‐term business needs Led and delivered solutions to architecture‐level problems including scalability, security, reliability, performance, maintainability, and observability Facilitated alignment across technical and non‐technical stakeholders to move initiatives forward through ambiguity and complexity Provided mentorship and team support at scale, while sharing ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Head of Platform Engineering, you'll define and deliver robust platform services, establish the patterns for self-service infrastructure, CI/CD, and observability, and drive the integration of AI/ML capabilities across the organisation. You'll translate platform strategy into a coherent technical roadmap, setting standards that teams … these are sustained as the platform evolves. Defining and maintaining platform standards, patterns, and reusable components that drive consistency across teams. Leading improvements to observability, monitoring, and alerting capabilities across systems and services. Mentoring and developing engineers across the organisation, raising the collective standard of platform engineering practice. Identifying ...

Senior Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
foundational infrastructure layer behind ACRA. This is a deep infrastructure role. You will work across Kubernetes, Linux, networking, storage, service‐to‐service communication, observability, security boundaries, and production operations. You will be responsible for the systems that everything else depends on: clusters, networks, storage layers, ingress and egress paths, runtime … security of the whole platform. You should be comfortable operating close to the metal: debugging Kubernetes, understanding networking behaviour, reasoning about distributed storage, improving observability, and helping define the infrastructure patterns that ACRA will rely on as it scales. This is not a generic DevOps support role or internal ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real‐time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Context Plane Python Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end-to-end — from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands-on experience using enterprise-authorized AI-assisted software development ...

Senior Data Management Professional - Data Engineer - Commodities Data

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
improve performance, reliability, and maintainability. Design automated pipeline controls for validation, monitoring, schema change, exception handling, and data integrity. Develop workflow orchestration, alerting, observability, and remediation processes. Translate business and client needs into engineering‐ready requirements and scalable technical solutions. Partner with Engineering on platform evolution, architecture, tooling, system design … hands‐on experience with Python or similar programming/scripting languages. Experience with querying structured, semi‐structured, and unstructured datasets. Experience with workflow orchestration, observability, monitoring, alerting, and scalable architecture design. Ability to analyze, refactor, and modernize legacy systems. Strong understanding of data lifecycle management, data integration, data modelling, data ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
trading decisions. You will be supporting customer-facing rollouts with Product and Commercial teams, ensuring production reliability. Additionally, you will help improve engineering processes, observability, and system resilience as we scale. Behind every line of code and every dataset, there’s a team of curious, driven people who bring ideas … strong experience building, deploying and operating reliable, scalable, end-to-end production APIs and data pipelines for external clients strong understanding of cloud infrastructure, observability, monitoring, incident response, and operational best practices relevant technologies (we use): AWS incl. S3, ECS, API Gateway, Batch, Redshift, Airflow, GitHub Actions, Docker, Pulumi, Python ...

Lead Infrastructure Engineer - AWS Cloud Support Engineering

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
adherence to resiliency and security expectations Familiarity with working in a large distributed system across a range of technologies including compute, databases, messaging, observability, and telemetry Knowledge of incident, change, and problem management processes and the controls that govern them Understanding of data-driven decision making and a drive … working in a follow-the-sun or globally distributed on-call support model Familiarity with large-scale cloud migration or modernization initiatives Exposure to observability and telemetry tooling in complex distributed environments ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products ...

Senior Data Management Professional - Data Engineer - Commodities Data London, GBR Posted today

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
improve performance, reliability, and maintainability. Design automated pipeline controls for validation, monitoring, schema change, exception handling, and data integrity. Develop workflow orchestration, alerting, observability, and remediation processes. Translate business and client needs into engineering‐ready requirements and scalable technical solutions. Partner with Engineering on platform evolution, architecture, tooling, system design … hands‐on experience with Python or similar programming/scripting languages. Experience with querying structured, semi‐structured, and unstructured datasets. Experience with workflow orchestration, observability, monitoring, alerting, and scalable architecture design. Ability to analyze, refactor, and modernize legacy systems. Strong understanding of data lifecycle management, data integration, data modelling, data ...

Staff Machine Learning Ops Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
optimization. Set the technical direction for CI/CD for ML, embedding testing, validation, security, performance checks, and release confidence into deployment pipelines. Establish observability standards for ML systems, including model metrics, service health, alerts, drift detection, data quality, lineage, and business-impact monitoring. Lead the evolution of Preply … monitoring, performance benchmarking, and lifecycle management. Strong hands‐on experience with cloud platforms such as GCP or AWS, Kubernetes, distributed compute, CI/CD, observability, and infrastructure-as-code practices. Experience building enabling tools and platform capabilities for Applied Scientists, Data Scientists, and engineering teams. Strong technical judgment ...

Advanced Solutions Architect

Hiring Organisation
Jobleads-UK
Location
Chiswick, England, United Kingdom
concept development when needed to de-risk architectural decisions## ## **Engineering Quality & Standards*** Define and champion non-functional standards: performance baselines, resilience patterns, observability requirements, and security controls* Lead or contribute to performance and load testing design, interpreting results in the context of SLAs and platform growth targets* Establish … delivery leads, product managers, and executive stakeholders* Drive alignment between iGaming and wider L&W engineering teams on shared concerns: API strategy, data platforms, observability, and shared services* Participate in hiring and technical assessment processes for engineering and architecture candidates**Qualifications**Degree in Computer Science, Software Engineering, or a related ...

QA Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
accessibility, regression and performance testing activities. Contribute to test management using Testmo and help drive continuous improvements across our quality processes. Use monitoring, observability and production insights to improve platform reliability and customer experience. Champion accessibility, resilience and operational excellence as core engineering principles. Evaluate emerging technologies, including AI‐assisted … Desirable Experience using BrowserStack or similar real‐device testing platforms. Experience with visual regression tools such as Percy. Experience with LogRocket, Grafana or similar observability platforms.Exposure to performance testing tools such as k6 or Gatling. Experience working with Google Cloud Platform. Understanding of SQL databases and backend services. Experience working ...

Senior AI Engineer| London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Vector databases & retrieval: Pinecone, Weaviate, Chroma, pgvector, FAISS; embeddings, semantic and hybrid search, reranking MLOps/LLMOps & deployment: Docker, Kubernetes, FastAPI, CI/CD; observability, tracing, evaluation tooling (LangSmith, LangFuse); guardrails, prompt/version management Responsible AI & safety: bias & fairness, hallucination mitigation, evaluation, privacy, security, governance of AI and agentic … optimization. LLMOps, Evaluation & Optimization : Experience operationalizing LLM and agentic applications—building evaluation harnesses and offline/online metrics for quality, groundedness, and safety; implementing observability, tracing, and monitoring; continuously optimizing accuracy, cost, and latency. Familiarity with guardrails, red‐team, and responsible deployment of AI systems in production. Communication Skills : Excellent ...

Staff .Net Backend Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
performance. Championing a high‐quality engineering culture – test coverage, peer review, CI/CD discipline via GitHub Actions, Infrastructure as Code, secure coding, observability and performance – aligned to the Reapit Global Technology Strategy, Reapit Connect and agentic tooling. Mentoring and up‐levelling engineers around you through pairing, PR review, architectural …/CD (ideally GitHub Actions), Infrastructure as Code (AWS CDK or Terraform), comprehensive testing (unit, integration and contract), and a genuine commitment to observability, performance and secure‐by‐default coding. Technical leadership without the title – a track record of lifting teams through pairing, mentoring, PR review and example, rather than ...

QA Test Architecture Engineer - DV Cleared

Hiring Organisation
Fuel Recruitment
Location
Taunton, Somerset, UK
Employment Type
Full-time
Job Description \n We are looking for a QA Test Infrastructure Engineer to work with one of our clients in the defence sector. You'll work alongside engineers, architects, and delivery specialists to develop technology ...

AI Platform Architect

Hiring Organisation
83zero Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £100000/annum
AI Platform Architect Location: London (Hybrid - 1-2 days per week) Salary: £80,000 - £100,000 + Bonus & Excellent Benefits We're partnering with a global technology consultancy delivering one of the UK's largest ...

SRE Managing Consultant - Cloud Operating Model

Hiring Organisation
Capgemini
Location
Manchester, United Kingdom
Employment Type
Full Time
Budgets : Establish service measures and targets (SLIs/SLOs) and introduce Error Budgets to enable data-driven trade-offs between reliability and delivery velocity. Observability & Operational Insight: Shape observability approaches (metrics/logs/traces) and operational monitoring models that make reliability risks visible and actionable, improving operational decision-making. … large‐scale delivery contexts; associate‐level certifications are desirable but not mandatory. Design, establish, and evolve SRE‐led centres of excellence (e.g. Reliability, Observability, or Operational Excellence), setting enterprise‐level standards for SLIs/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms. Exposure to modern observability ...

SRE Managing Consultant - Cloud Operating Model

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Budgets**: Establish service measures and targets (SLIs/SLOs) and introduce Error Budgets to enable data-driven trade-offs between reliability and delivery velocity.* **Observability & Operational Insight:** Shape observability approaches (metrics/logs/traces) and operational monitoring models that make reliability risks visible and actionable, improving operational decision-making. … large‐scale delivery contexts; associate‐level certifications are desirable but not mandatory.* Design, establish, and evolve SRE‐led centres of excellence (e.g. Reliability, Observability, or Operational Excellence), setting enterprise‐level standards for SLIs/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms.* Exposure to modern observability ...

GCP DevOps

Hiring Organisation
Pracyva ltd
Location
Bristol, City of Bristol, United Kingdom
Employment Type
Contract
Contract Rate
£400 - £425/day
Actions, Harness, Jenkins). Networking & Security: Experience with GCP Cloud Armor, GCP Networking, and embedding secure-by-design controls from design to runtime. Automation & Observability: Implementing actionable observability, performance tuning, and automation to reduce toil. Defining and operating against SLOs/SLIs. Scripting & Tooling: Scripting in Bash, PowerShell, or Python. ...

Site Reliability Engineer (SRE)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
approach. Key Responsibilities Integrating tightly with our Product Engineering teams Following SRE practices and maintaining high standards of compliance Implementing a new standard of observability utilising SLI/SLO/Error Budgets Continually evolving our observability platforms for greater coverage Using a code-first approach to build and changes … ongoing communication with stakeholders Skills Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer Observability product experience (eg Datadog) Managing services using SLI/SLO & Error Budgets Experience with AWS or other cloud providers Experience in HA environments Automation skills ...

Applied Scientist, Observability, Prime Video

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
innovating on behalf of our customers is at the heart of everything we do. If this sounds exciting to you, please read on. PV observability team's mission is to deliver efficient, zero-touch observability solutions that combine log management, tracing, and AI-powered analytics, enabling teams to detect, diagnose … will work alongside other scientists and engineering teams to deliver your research into production systems. About the team Our team owns Prime Video observability features for development teams. We consume PBs of data daily which feed into multiple observability features focussed on reducing the customer impact time. Basic Qualifications ...

Senior Data Practice Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
pick the right approach for the problem Data testing: know what to catch at build time (schema contracts, assertions, transformation logic) versus defer to observability, and can make that call for a team Data observability: treat data reliability like site reliability, with measurable indicators, alerting, incident response, and root cause ...

DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions Ltd
Location
Leeds, West Yorkshire, England, United Kingdom
Employment Type
Contractor
Contract Rate
£400 - £450 per day
InsideIR35 | Hybrid 1 day onsite in Leeds | 6 month Initial contract DevOps Engineer to work on the development of an large scale observability platform, experienced in cloud techologies, infrastracture as code, building and operating distributed systems at scale, proficient in devops engineering, reviewing, writing and testing code, working with source … control mechanisms and and deploying infrastructure. Key experience we are looking for: Previous experience building and supporting large-scale AWS observability and monitoring platforms. Strong Python development background with experience creating automation and engineering tooling. Hands-on Kubernetes (K8s) experience deploying, managing and troubleshooting containerised workloads. Experience using Grafana ...

Cloud Advisory Senior Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security and cost outcomes: Embed zero trust IAM, security‐by‐design, HA/DR, and operational controls. Define SLOs/SLIs and bake in observability (metrics, logs, traces). Apply AIOps to reduce noise, accelerate incident triage, and improve reliability, and embed FinOps to manage performance and run‐cost value …/DR, containers/orchestration, API management, and iPaaS, plus modern engineering patterns such as microservices, event‐driven architecture, and DDD. Operational excellence, observability and AIOps: Translate NFRs into pragmatic architecture decisions, define SLOs/SLIs, and design modern observability (metrics, logs, traces). Apply AIOps for alert reduction, anomaly ...

DevOps Engineer

Hiring Organisation
Oscar Associates (UK) Limited
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£70,000
scalable, reliable and cost-efficient as it moves into full production. Working closely with engineering teams, you'll drive automation, improve deployment pipelines, strengthen observability and ensure the platform performs under high-volume, real-time workloads. This is a hands-on position with genuine ownership and plenty of opportunity … enhancing CI/CD pipelines with blue/green deployments and automated rollback Driving platform reliability, resilience and scalability Developing monitoring, alerting and observability across the environment Managing cloud costs and implementing best FinOps practices Participating in a small production on-call rota Technology AWS ECS Fargate Terraform Aurora ...