1,451 to 1,475 of 1,810 Permanent Observability Jobs

Advanced Solutions Architect

Hiring Organisation
Jobleads-UK
Location
Chiswick, England, United Kingdom
concept development when needed to de-risk architectural decisions## ## **Engineering Quality & Standards*** Define and champion non-functional standards: performance baselines, resilience patterns, observability requirements, and security controls* Lead or contribute to performance and load testing design, interpreting results in the context of SLAs and platform growth targets* Establish … delivery leads, product managers, and executive stakeholders* Drive alignment between iGaming and wider L&W engineering teams on shared concerns: API strategy, data platforms, observability, and shared services* Participate in hiring and technical assessment processes for engineering and architecture candidates**Qualifications**Degree in Computer Science, Software Engineering, or a related ...

QA Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
accessibility, regression and performance testing activities. Contribute to test management using Testmo and help drive continuous improvements across our quality processes. Use monitoring, observability and production insights to improve platform reliability and customer experience. Champion accessibility, resilience and operational excellence as core engineering principles. Evaluate emerging technologies, including AI‐assisted … Desirable Experience using BrowserStack or similar real‐device testing platforms. Experience with visual regression tools such as Percy. Experience with LogRocket, Grafana or similar observability platforms.Exposure to performance testing tools such as k6 or Gatling. Experience working with Google Cloud Platform. Understanding of SQL databases and backend services. Experience working ...

Senior AI Engineer| London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Vector databases & retrieval: Pinecone, Weaviate, Chroma, pgvector, FAISS; embeddings, semantic and hybrid search, reranking MLOps/LLMOps & deployment: Docker, Kubernetes, FastAPI, CI/CD; observability, tracing, evaluation tooling (LangSmith, LangFuse); guardrails, prompt/version management Responsible AI & safety: bias & fairness, hallucination mitigation, evaluation, privacy, security, governance of AI and agentic … optimization. LLMOps, Evaluation & Optimization : Experience operationalizing LLM and agentic applications—building evaluation harnesses and offline/online metrics for quality, groundedness, and safety; implementing observability, tracing, and monitoring; continuously optimizing accuracy, cost, and latency. Familiarity with guardrails, red‐team, and responsible deployment of AI systems in production. Communication Skills : Excellent ...

Staff .Net Backend Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
performance. Championing a high‐quality engineering culture – test coverage, peer review, CI/CD discipline via GitHub Actions, Infrastructure as Code, secure coding, observability and performance – aligned to the Reapit Global Technology Strategy, Reapit Connect and agentic tooling. Mentoring and up‐levelling engineers around you through pairing, PR review, architectural …/CD (ideally GitHub Actions), Infrastructure as Code (AWS CDK or Terraform), comprehensive testing (unit, integration and contract), and a genuine commitment to observability, performance and secure‐by‐default coding. Technical leadership without the title – a track record of lifting teams through pairing, mentoring, PR review and example, rather than ...

QA Test Architecture Engineer - DV Cleared

Hiring Organisation
Fuel Recruitment
Location
Taunton, Somerset, UK
Employment Type
Full-time
Job Description \n We are looking for a QA Test Infrastructure Engineer to work with one of our clients in the defence sector. You'll work alongside engineers, architects, and delivery specialists to develop technology ...

AI Platform Architect

Hiring Organisation
83zero Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £100000/annum
AI Platform Architect Location: London (Hybrid - 1-2 days per week) Salary: £80,000 - £100,000 + Bonus & Excellent Benefits We're partnering with a global technology consultancy delivering one of the UK's largest ...

SRE Managing Consultant - Cloud Operating Model

Hiring Organisation
Capgemini
Location
Manchester, United Kingdom
Employment Type
Full Time
Budgets : Establish service measures and targets (SLIs/SLOs) and introduce Error Budgets to enable data-driven trade-offs between reliability and delivery velocity. Observability & Operational Insight: Shape observability approaches (metrics/logs/traces) and operational monitoring models that make reliability risks visible and actionable, improving operational decision-making. … large‐scale delivery contexts; associate‐level certifications are desirable but not mandatory. Design, establish, and evolve SRE‐led centres of excellence (e.g. Reliability, Observability, or Operational Excellence), setting enterprise‐level standards for SLIs/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms. Exposure to modern observability ...

SRE Managing Consultant - Cloud Operating Model

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Budgets**: Establish service measures and targets (SLIs/SLOs) and introduce Error Budgets to enable data-driven trade-offs between reliability and delivery velocity.* **Observability & Operational Insight:** Shape observability approaches (metrics/logs/traces) and operational monitoring models that make reliability risks visible and actionable, improving operational decision-making. … large‐scale delivery contexts; associate‐level certifications are desirable but not mandatory.* Design, establish, and evolve SRE‐led centres of excellence (e.g. Reliability, Observability, or Operational Excellence), setting enterprise‐level standards for SLIs/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms.* Exposure to modern observability ...

Site Reliability Engineer (SRE)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
approach. Key Responsibilities Integrating tightly with our Product Engineering teams Following SRE practices and maintaining high standards of compliance Implementing a new standard of observability utilising SLI/SLO/Error Budgets Continually evolving our observability platforms for greater coverage Using a code-first approach to build and changes … ongoing communication with stakeholders Skills Good experience in DevOps or SRE, with a keen interest to learn and grow as a Site Reliability Engineer Observability product experience (eg Datadog) Managing services using SLI/SLO & Error Budgets Experience with AWS or other cloud providers Experience in HA environments Automation skills ...

Applied Scientist, Observability, Prime Video

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
innovating on behalf of our customers is at the heart of everything we do. If this sounds exciting to you, please read on. PV observability team's mission is to deliver efficient, zero-touch observability solutions that combine log management, tracing, and AI-powered analytics, enabling teams to detect, diagnose … will work alongside other scientists and engineering teams to deliver your research into production systems. About the team Our team owns Prime Video observability features for development teams. We consume PBs of data daily which feed into multiple observability features focussed on reducing the customer impact time. Basic Qualifications ...

Senior Data Practice Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
pick the right approach for the problem Data testing: know what to catch at build time (schema contracts, assertions, transformation logic) versus defer to observability, and can make that call for a team Data observability: treat data reliability like site reliability, with measurable indicators, alerting, incident response, and root cause ...

Cloud Advisory Senior Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security and cost outcomes: Embed zero trust IAM, security‐by‐design, HA/DR, and operational controls. Define SLOs/SLIs and bake in observability (metrics, logs, traces). Apply AIOps to reduce noise, accelerate incident triage, and improve reliability, and embed FinOps to manage performance and run‐cost value …/DR, containers/orchestration, API management, and iPaaS, plus modern engineering patterns such as microservices, event‐driven architecture, and DDD. Operational excellence, observability and AIOps: Translate NFRs into pragmatic architecture decisions, define SLOs/SLIs, and design modern observability (metrics, logs, traces). Apply AIOps for alert reduction, anomaly ...

DevOps Engineer

Hiring Organisation
Oscar Associates (UK) Limited
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£70,000
scalable, reliable and cost-efficient as it moves into full production. Working closely with engineering teams, you'll drive automation, improve deployment pipelines, strengthen observability and ensure the platform performs under high-volume, real-time workloads. This is a hands-on position with genuine ownership and plenty of opportunity … enhancing CI/CD pipelines with blue/green deployments and automated rollback Driving platform reliability, resilience and scalability Developing monitoring, alerting and observability across the environment Managing cloud costs and implementing best FinOps practices Participating in a small production on-call rota Technology AWS ECS Fargate Terraform Aurora ...

Senior Backend Engineer - Databases Pyroscope | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
United Kingdom (Remote) Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations … cost that makes sense. Turn Pyroscope into a platform capability inside Grafana: bi-directional trace-to-profile correlation, integration with Kubernetes Monitoring and App Observability, and profiles surfaced where engineers already start their investigations. Prepare Pyroscope for an agent-driven world: APIs, CLI, and docs designed so AI agents ...

Senior AI Enablement Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
. You will help build the paved roads that allow engineers to use AI tools safely and effectively: standards, reusable workflows, automation, evaluation approaches, observability, documentation, and internal platform capabilities. The goal is to make AI‐assisted engineering reliable, measurable, secure, and aligned with enterprise software delivery standards. You will … testing, code review, onboarding, and knowledge retrieval. Contribute to AI governance implementation by helping translate policy and security expectations into usable engineering workflows. Support observability and measurement for AI adoption, including usage insights, effectiveness, quality signals, and operational risks. Partner with platform, architecture, AppSec, and infrastructure teams to ensure ...

Software Engineer - Synthetic Monitoring | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … Monitoring as code and with confidence Break down complex, ambiguous problems into incremental deliverables and iterate quickly based on feedback Ensure quality through testing, observability of your own systems, and documentation : our checks are something customers alert on, so reliability is a feature Be a part of the team ...

Agentic AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
AIOps. The role will focus on designing, building, deploying and operating enterprise AI agents and multi-agent workflows on AWS, with strong emphasis on observability, reliability, cost control, security and continuous optimization in production environments. Required Skills Agentic AI, Engineering, AWS Platform, AIOps/LLMOps, Dev Key Responsibilities AI Agent … DynamoDB and SQS/SNS. Create reusable libraries, patterns and accelerators to standardize AI agent development across teams. AIOps, Production Monitoring & Operations Establish monitoring, observability and operational governance for production AI workloads. Track agent performance, model latency, cost, prompt effectiveness, error rates and quality signals. Define alerting, incident response ...

Senior Backend Software Engineer (AI Infrastructure / Artifact Management) – Developer Services

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automated remediation, advancing AI from assisted analysis toward closed-loop resolution; Architect and optimize large-scale distributed systems to continuously improve throughput, reliability, observability, and developer experience. Work closely with engineering, security, infrastructure, and AI business teams to understand real-world requirements and drive complex projects from design to adoption … solid engineering skills; Familiarity with distributed systems and hands‐on experience in one or more areas such as storage, computing, task scheduling, caching, messaging, observability, or reliability engineering; Strong problem‐solving, system design, and execution skills, with the ability to diagnose issues across complex system paths and drive long‐term ...

SRE Technical Lead

Hiring Organisation
Adecco
Location
Reading, Berkshire, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 90,000 Annual
remediation Act as the technical escalation point for major incidents and high-risk releases Lead blameless post-incident reviews and ensure continuous improvement Establish observability and capacity management practices using modern tooling Identify and eliminate systemic reliability risks and operational inefficiencies Collaborate with engineering, platform, security, and operations teams across … Experience working in multi-cloud or hybrid cloud environments Strong understanding of SRE principles (SLOs, SLAs, error budgets, reliability engineering) Hands-on experience with observability tooling (eg, Prometheus, Grafana, OpenTelemetry, Loki, Tempo) Strong knowledge of Infrastructure as Code and GitOps (eg, Helm, Kustomize, ArgoCD, Tekton) Experience with CI/ ...

Senior Observability Solution Architect – Pre-Sales

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
leading observability platform in Greater London is seeking an experienced Solution Engineer to join their team. This role involves collaborating with account executives on technical sales cycles, delivering impactful presentations, and overseeing technical aspects of the process. The ideal candidate will have a minimum of 5 years in a customer ...

Staff Analytics Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Senior Cloud Engineer - Contract

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Engineering builds and runs the new Azure cloud that everything else at Flagstone depends on. It's the foundation for our security tooling, our observability, and the hosting for our AI. The team works in infrastructure-as-code (Terraform and Bicep), owns the landing zones and hub-and-spoke networking … roll out golden paths that cut delivery cost and speed up engineering squads. Stand up our AI platform foundations: an AI gateway, an observability stack, and hosting for AI tools including Flagstone Concierge, with model access through AWS Bedrock. Build and maintain infrastructure pipelines with security scanning, plan validation ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
platform engineering and operational backbone of our solutions. Responsibilities Define and maintain Auriga’s DevOps and platform engineering roadmap across infrastructure, CI/CD, observability, reliability and security. Design scalable platform architectures across cloud, on-premises, hybrid and customer-site environments. Establish reusable standards, deployment patterns and reference architectures across … deployments across development, test, staging, production and customer environments. Design and support customer-isolated, multi-site and multi-tenant deployment models where required. Implement observability, monitoring, alerting, SLOs and operational health practices for production and customer-facing systems. Lead incident response, post-incident reviews and continuous improvement of platform reliability ...

Global DevOps Lead

Hiring Organisation
Stott & May Professional Search Limited
Location
United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
with engineering, cloud, and operations teams to deliver a modern, automated, and scalable platform. You'll drive DevOps strategy across infrastructure, CI/CD, observability, SRE, and cloud optimisation while influencing senior stakeholders across the business. Key Responsibilities - Define and implement a global DevOps operating model, including governance, standards … initiatives. - Partner with engineering and cloud teams to establish clear ownership across DevOps and infrastructure. - Lead the implementation and optimisation of enterprise monitoring and observability using Datadog. - Build scalable deployment pipelines that improve release quality and speed. - Establish and monitor DORA metrics, driving improvements in deployment frequency, lead time, change ...