3,326 to 3,350 of 3,962 Observability Jobs

Senior Product Manager - Platform

Location
Belfast City District, Northern Ireland, United Kingdom
THIS ROLE We're hiring a Senior Platform Product Manager to own two of our most critical platform areas: our API gateway and our observability platform. The API gateway is the public entry point for all Apex APIs, handling hundreds of millions of requests per month (and growing). Observability … with a track record of owning cloud platform infrastructure Good understanding of API concepts - latency, rate limiting, response codes Experience with Datadog or comparable observability tooling Demonstrated ability to set product vision and strategy, and communicate it credibly to technical and non-technical stakeholders Excellent prioritization judgment in ambiguous, highly ...

Forward Deployed Strategist

Location
Greater London, England, United Kingdom
Dynamo’s products within enterprise environments Translate customer requirements into deployment strategies and implementation plans Help customers design scalable workflows around AI governance, evaluation, observability, and real-time guardrails Navigate ambiguity and unblock cross-functional execution across internal and customer teams Balance speed, technical feasibility, governance, and business impact during … from ambiguity to execution Nice To Have Experience with Generative AI, LLM deployments, AI governance, AI evaluation, or guardrail systems Familiarity with enterprise AI observability, monitoring, or security workflows Experience working with regulated industries such as financial services, healthcare, or government Background in management consulting, enterprise SaaS, or technical implementation ...

AI Engineer

Location
Rochdale, England, United Kingdom
shipping LLM features to real users. Fluent in Python or TypeScript; comfortable across both. Understand embeddings, retrieval, prompt design and evaluation. Rigorous about testing, observability and cost. Right to work in the UK. #J-18808-Ljbffr ...

AI AGENTS ENGINEER

Location
Manchester, England, United Kingdom
status, support history, and customer‐specific metadata. Design agent workflows that sit inside real task surfaces rather than generic chatbot experiences. Build evaluation and observability for agent behaviour, including tool‐call history, task success metrics, regression tests, failure modes, traceability, and guardrails. Collaborate with ML, product, hardware, client, and deployment … models or multimodal AI over images, video, inspection evidence, diagrams, screenshots, or technical records. Experience with RAG over structured and unstructured data. Familiarity with observability, log analysis, incident response, support tooling, runbooks, or developer tools. Experience building workflow UIs where AI assists a specific operational task. Familiarity with manufacturing, quality ...

Application Support Production Operations

Hiring Organisation
vaaridatech
Location
Massachusetts, United States
Employment Type
Permanent
Salary
USD Annual
experience supporting applications hosted on Azure with Azure platform administration. Experience managing application servers, service availability, and operational health. Experience with monitoring and observability tools such as Azure Monitor, Application Insights, Log Analytics, Splunk, Dynatrace, AppDynamics, or Grafana. Experience in incident management, problem management, root cause analysis, and operational support ...

Clickhouse Solutions Architect

Location
Slough, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...

Clickhouse Solutions Architect

Location
City Of London, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...

Senior GTM System Builder

Location
Wedmore, England, United Kingdom
work under staff technical direction but make your own architecture and build-vs-buy decisions within your domain. You will manage vendor relationships, define observability standards, and be the person who picks up the phone when something breaks in production. What You Would Do Own the full lifecycle of automation … Design the prompting architecture and agent logic that commercial reps, BDRs, and AEs interact with daily — reliable outputs for non‐technical end users, with observability and error handling built in from the start Own operational health of your domain’s live systems: first responder when things break, with adoption metrics ...

Technical Services Manager

Hiring Organisation
Hackajob Ltd
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
features from design -> build -> test -> release -> operational readiness, contributing directly through coding and code reviews Establish and uphold standards for maintainability, security, reliability, and observability across the mainframe DevEx/tooling estate Act as senior escalation for complex build/deploy/tooling issues across UK/US/Mexico … toolchain for source, build, artefact, and deployment across regions; design and deliver production-grade APIs wrapping SDLC/platform tooling (clear contracts, versioning, resilience, observability) Introduce and scale AI-assisted engineering practices with safe working and review expectations; run structured pilots, quantify outcomes, and scale what works across UK/ ...

Cloud Infra Engineer

Hiring Organisation
DNS Info Ltd
Location
Glasgow, Glasgow City, City of Glasgow, United Kingdom
Employment Type
Contract
Contract Rate
£350 - £400/day Inside IR 35
experience: ForgeRock, Okta, Ping, Auth0 Kafka/event-streaming: producing/consuming events, setting up topics/consumers, reacting to events in a service Observability: Prometheus/Grafana/OTel ...

Senior Product Manager, Performance

Location
Greater London, England, United Kingdom
Partner with experimentation and analytics teams to evaluate how product, content, and feature changes affect performance and traveler outcomes. Collaborate with teams responsible for observability, monitoring, and developer tooling to improve visibility into performance bottlenecks and degradations. Influence roadmaps across multiple product and platform teams without direct authority, aligning leaders … ability to partner closely with engineering teams on complex technical challenges. Deep understanding of web and/or mobile application performance, including performance measurement, observability, monitoring, diagnostics, release protections, and performance regression prevention. Experience using data, analytics, experimentation, and customer insights to evaluate impact and make informed product decisions ...

Sr Director, Platform Engineering - Data Platform & Agentic Platform

Hiring Organisation
RELX Group
Location
England, UK
Employment Type
Full-time
mechanismsImplement and operate agent workflow platform capabilities aligned to product-defined standards and interfaces, including traceability, state handling, and convergence patternsImplement production-grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflowsRequirements12+ years leading platform engineering teams delivering shared services or large-scale SaaS systems with production operations … including ingestion and integration patterns, data quality enforcement, metadata and lineage, and secure access controlsExperience enabling AI-powered systems in production, including evaluation, monitoring, observability, auditability, and guardrailsStrong ability to partner with product leaders on contract-driven platforms and deliver outcomes without creating bottlenecksProven capability to hire and develop high ...

Staff Software Engineer, Ingestion Systems

Location
Greater London, England, United Kingdom
training and research Building interfaces and systems that reduce bottlenecks and improve reliability and scalability Setting best practices for pipeline reliability, including observability, alerting and monitoring Leading work to reduce pipeline latency, improve failure recovery and meet SLAs Building a culture of transparency, collaboration and shared ownership across teams … pipelines Familiarity with third-party dataset ingestion and transformation Understanding of compliance and data governance frameworks (e.g. GDPR, TISAX, ASPICE) Hands-on experience integrating observability and monitoring solutions More about Wayve: Wayve is building the leading AI platform for autonomous driving. We are pioneering an end to end AI approach ...

Applied AI Engineer, Digital Natives

Location
Greater London, England, United Kingdom
code to build prototypes, evaluation harnesses, reference implementations, integrations, and production accelerators. Make sound technical decisions across models, agents, retrieval, tools, data, reliability, observability, latency, cost, safety, security, and governance. Diagnose complex implementation challenges, reproduce failures, test hypotheses, and drive blockers toward resolution. Help customers progress from promising prototypes … evaluate AI systems systematically using representative data, graders, production signals, and human judgment. Have navigated enterprise production requirements such as integrations, reliability, observability, security, privacy, data governance, performance, and cost. Can connect technical decisions to customer workflows, adoption, and measurable business outcomes. Communicate with clarity and credibility across hands ...

Lead Engineer, AI

Location
United Kingdom
systems, rather than ML research or model training. We’re looking for a Lead Engineer to help us build out the evaluation frameworks, observability tooling, and diagnostic infrastructure that tell us whether our AI Agents are working well, where to improve them, and how to make them faster and more … built on and proficiency in Python is a strong plus, especially for the data and evaluation side of the work. Experience building evaluation or observability infrastructure for ML/AI systems. You've built eval pipelines, scorers, dashboards, or CI/CD for evals before and have experience with evaluation ...

Senior Product Manager, Performance

Location
Greater London, England, United Kingdom
remediation.Partner with experimentation and analytics teams to evaluate how product, content, and feature changes affect performance and traveler outcomes.Collaborate with teams responsible for observability, monitoring, and developer tooling to improve visibility into performance bottlenecks and degradations.Influence roadmaps across multiple product and platform teams without direct authority, aligning leaders on shared … ability to partner closely with engineering teams on complex technical challenges.Deep understanding of web and/or mobile application performance, including performance measurement, observability, monitoring, diagnostics, release protections, and performance regression prevention.Experience using data, analytics, experimentation, and customer insights to evaluate impact and make informed product decisions and trade-offs.Proven ...

Lead Architect (Mission) - Airline Operations

Location
Greater London, England, United Kingdom
communicate complex technical concepts to both technical and non-technical audiences. · Balance strategic thinking with practical delivery considerations and trade-offs. · Familiarity with observability platforms, telemetry, event streaming, and modern data architectures. · Knowledge of automation frameworks, digital operations platforms, and emerging technology trends. · Experience with cloud-native architectures and platform … Establishing architecture standards, governance frameworks, and assurance processes. · Strong understanding of distributed systems, cloud platforms, APIs, integration patterns, and event-driven architectures. · Familiarity with observability platforms, telemetry, event streaming, and modern data architectures. · Knowledge of automation frameworks, digital operations platforms, and emerging technology trends. · Understanding of data architecture, operational data ...

Lead Engineer, AI Platform

Location
Greater London, England, United Kingdom
systems, rather than ML research or model training. We're looking for a Lead Engineer to help us build out the evaluation frameworks, observability tooling, and diagnostic infrastructure that tell us whether our AI Agents are working well, where to improve them, and how to make them faster and more … built on and proficiency in Python is a strong plus, especially for the data and evaluation side of the work. Experience building evaluation or observability infrastructure for ML/AI systems. You've built eval pipelines, scorers, dashboards, or CI/CD for evals before and have experience with evaluation ...

Staff Android Engineer

Location
Greater London, England, United Kingdom
modern Android platform practices. Improve reliability and delivery confidence: Shape how Android work is designed, reviewed, tested, shipped, and operated, including the automation, observability, rollout practices, and AI‐enabled validation needed to move quickly with confidence. Create clear mobile‐facing system contracts: Partner with iOS, backend, product, and design … deep Android base, with enough iOS fluency to shape cross‐functional architecture decisions where backend, web, and product choices affect identity, reliability, release quality, observability, and native product outcomes. Experience shaping mobile‐facing platform contracts, API behavior, or backend service boundaries that improve product quality across native client experiences. Ability ...

Senior Product Manager - Digital Experience

Location
Greater London, England, United Kingdom
Web.* Use data, incidents, service insight and stakeholder feedback to identify improvements and inform roadmap decisions.* Balance feature delivery with security, resilience, observability, scalability, technical debt and ongoing platform health.* Provide clear roadmap, progress, risk and dependency updates, influencing senior stakeholders and resolving competing priorities across teams.**Knowledge, Skills … volume consumer App and Web products in a complex, multi-brand or international environment.* Understanding of non-functional requirements including security, privacy, performance, reliability, observability and scalability.* Experience with product analytics, service monitoring and incident insight to measure outcomes and improve platform health.* Familiarity with agile product delivery tools ...

FP&A Lead

Location
Greater London, England, United Kingdom
ITRS, we make society's critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries - trusted by 90% of Tier 1 capital markets ...

Lead Architect (Mission) - Airline Operations

Hiring Organisation
Easyjet
Location
London, UK
Employment Type
Full-time
communicate complex technical concepts to both technical and non-technical audiences.· Balance strategic thinking with practical delivery considerations and trade-offs.· Familiarity with observability platforms, telemetry, event streaming, and modern data architectures.· Knowledge of automation frameworks, digital operations platforms, and emerging technology trends.· Experience with cloud-native architectures and platform … Establishing architecture standards, governance frameworks, and assurance processes.· Strong understanding of distributed systems, cloud platforms, APIs, integration patterns, and event-driven architectures.· Familiarity with observability platforms, telemetry, event streaming, and modern data architectures.· Knowledge of automation frameworks, digital operations platforms, and emerging technology trends.· Understanding of data architecture, operational data ...

Applied AI Engineer

Location
Greater London, England, United Kingdom
code to build prototypes, evaluation harnesses, reference implementations, integrations, and production accelerators. Make sound technical decisions across models, agents, retrieval, tools, data, reliability, observability, latency, cost, safety, security, and governance. Diagnose complex implementation challenges, reproduce failures, test hypotheses, and drive blockers toward resolution. Help customers progress from promising prototypes … evaluate AI systems systematically using representative data, graders, production signals, and human judgment. Have navigated enterprise production requirements such as integrations, reliability, observability, security, privacy, data governance, performance, and cost. Can connect technical decisions to customer workflows, adoption, and measurable business outcomes. Communicate with clarity and credibility across hands ...

Senior Software Engineer

Hiring Organisation
Civica
Location
United Kingdom, UK
Employment Type
Full-time
your product manager and your principal engineer both understand Define what "good" looks like for your area of the codebase - testing strategy, code quality, observability, and documentation Collaborate with product, design, and other engineers to understand the real problem before building anything RequirementsA track record of building and shipping complex … before 'how' and can translate a vague business need into a structured technical approach Production-quality discipline - you care about testing, observability, error handling, and operational readiness, not just making things work locally Communication that reaches beyond engineers - you can explain a technical trade-off to a product manager, write ...

Technical Lead- IAM

Location
Sheffield Green, England, United Kingdom
activities across internal teamsand third-party suppliers. Support planning, estimation and technical delivery. Define and assure non-functional requirements including: Availability - Resilience - Performance - Scalability -Observability - Security - Ensure solutions are production-ready andoperationally supportable. Supportarchitecture governance, design authorities and technical steering groups. Championengineering best practices, DevSecOps and continuous improvement. #J ...