1,701 to 1,725 of 1,808 Permanent Observability Jobs

Technical Product Manager — Data Manufacturing Infrastructure London, GBR

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
role will partner closely with DMO, partner Engineering Infrastructure, AI, and domain teams to define a product roadmap for infrastructure capabilities that support automation, observability, process analysis, semantic data readiness, and scalable production workflows. This is not a traditional project management role. You will apply product discipline to infrastructure: translating … goals and strategy. Prioritize needs across multiple stakeholders to construct a coherent backlog that reduces complexity and achieves focus. Balance competing infrastructure needs, including observability, pipeline analysis, and technical migrations. Possess a robust knowledge of data manufacturing approaches across Data, and develop strategies that improve adoption while respecting Engineering architecture ...

Cribl Engineer Expert

Hiring Organisation
EMC
Location
Reston, Virginia, United States
Employment Type
Permanent
Salary
USD 180,000 Annual
counterintelligence (CI) polygraph. Job Description: Job Summary: We are seeking a highly experienced Cribl Engineer to serve as the principal technical authority for observability pipelines built on Cribl Stream and Cribl Edge. This role is designed for a senior technologist with deep expertise in log/telemetry routing, large-scale … data engineering, and enterprise-grade observability architectures. You will shape pipeline strategy, design complex routing and transformation logic, drive platform reliability, mentor senior engineers, and serve as the top technical escalation point for Cribl-related challenges. Key Responsibilities: Lead architecture and design for Cribl Stream/Edge across multiple enclaves ...

Managing Engineer – Cyber Platform Engineering

Hiring Organisation
Jobleads-UK
Location
Belfast, Northern Ireland, United Kingdom
secure, and observable services, tools, platforms, or data pipelines aligned to shared engineering standards. Own service and platform lifecycle expectations, including reliability, operational readiness, observability, and continuous improvement. Partner across product, platform, and engineering stakeholders to shape architecture, integration design, and delivery sequencing. Drive platform maturity through automation, reuse, standardization … experience leading engineering teams or technical delivery in product-based environments. Experience operating production systems, services, tools, or pipelines with strong expectations around reliability, observability, security, and maintainability. Strong understanding of integration patterns, automation, cloud-native engineering practices, and scalable platform design. Experience building reusable services, workflows, platforms, or data ...

Lead Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
continuous improvement* Define and evolve engineering standards, frameworks and best practices across the entire engineering organisation* Drive improvements in software quality, testing strategy, observability and release confidence* Partner with engineering, platform, product and security to deliver large-scale, cross-functional improvements* Shape and deliver internal tooling and AI-assisted engineering … engineering-first mindset and influencing people with a broader organisational impact* Previous experience on building enterprise level, highly scalable projects, focused on performance, observability and security best practices.* Experience in a technology-driven organisation with strong engineering standards**You’ll also bring:*** Strong experience improving engineering quality, reliability and operational ...

Inference Engineering - Platform Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
production system our customers route real traffic through. When it degrades, their product degrades. You will own its operational health end-to-end: SLOs, observability, incident response, deployment discipline, and capacity planning across heterogeneous compute backends. As the platform scales to millions of concurrent requests across heterogeneous compute, the work … with the hardware and orchestration teams to expose heterogeneous backends reliably through the platform. Who You'll Build Service-level objectives, monitoring, alerting, and observability On-call and incident response: runbooks, escalation, blameless postmortems, and follow-through Capacity planning and the operational side of running across heterogeneous compute backends What ...

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
operate agent workflow platform capabilities aligned to product‐defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production‐grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large‐scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands‐on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Senior Machine Learning Engineer, Developer Advocacy | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Senior ML Engineer Recommender Systems, Developer Advocacy | UK | Remote Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud … Experience with directed graphs, sequence models, or prerequisite‐aware recommendations Experience with contextual bandits or other exploration strategies Familiarity with Grafana or the broader observability ecosystem Experience with open source software or transparent development practices Experience working with privacy, fairness, explainability, or responsible personalization constraints Compensation & Rewards ...

Cloud FinOps Analyst

Hiring Organisation
Manufacturing Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£50,000
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Cloud FinOps Analyst

Hiring Organisation
Manufacturing Recruitment Limited
Location
Leicester, Leicestershire, East Midlands, United Kingdom
Employment Type
Permanent
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Principal Software Engineer – Customer Platforms

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
customers to a desired outcome, without prescribing it. Authoritative skills at cloud computing (network, security, serverless, Kubernetes etc) and automation. Experience with implementation of Observability and Reliability using market technologies (e.g.: New Relic). Good experience with Performance Engineering (load testing, derivations, tuning, core web vitals, page speed etc.). … reliability testing. Able to influence people at senior levels and from the highly technical to non‐technical. Hard Skills System Design Quality Assurance Observability Implementation Reliability Testing Prototyping Methods Soft Skills Strategic Thinking Mentoring Influencing Team Leadership Innovation Advocacy #J-18808-Ljbffr ...

Staff Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering standards, best practices and reusable patterns while partnering with Enterprise Architecture and influencing technical direction Drive engineering excellence by improving code quality, testing, observability, reliability and operational practices Support end‐to‐end delivery by guiding teams through complex technical challenges, improving decision‐making, and contributing to planning and risk … data lakes/lakehouse architectures, Iceberg or similar table formats, as well as batch and streaming processing Knowledge of data quality, governance, cataloguing and observability tools (e.g. Datadog), with DBT or AI‐assisted engineering practices as a plus Your benefits 29 days holiday allowance + bank holidays Private medical ...

Software Development Engineer - Amazon Ads, Creative-X (Advertising)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
right audiences with meaningful, high-quality experiences. Within this space, the Creative-X organization is at the forefront of redefining ad quality, delivery, and observability through agentic AI systems, large language models, and intelligent automation across the creative lifecycle. We're looking for an experienced Software Development Engineer to join … major third‐party ad servers Optimize service infrastructure for cost, throughput, and latency to keep up with Streaming TV creative volume growth, and improve observability so advertiser‐driven failures are clearly distinguished from service‐side regressions Deploy features safely through CI/CD, participate in on‐call rotations for advertiser ...

Software Engineer, ChatGPT Infrastructure

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
diagnosing performance, scalability, or reliability issues in production environments. Understanding of distributed systems, data storage, concurrency, asynchronous processing, or networking. Familiarity with modern deployment, observability, and cloud infrastructure practices. Ability to lead complex technical work and collaborate effectively across product and infrastructure teams. Experience with cloud infrastructure, containerized environments … observability tools is useful but not required. We welcome candidates from backend engineering, distributed systems, platform engineering, and reliability backgrounds. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general‐purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities ...

Senior Engineering Manager- Payment Experience

Hiring Organisation
Jobleads-UK
Location
United Kingdom
pragmatic RFC/ADR process for architectural changes, ensuring decisions are documented, reviewed, and reversible where possible. Own production excellence across the domain, including observability with Prometheus, Grafana, and Sentry; end-to-end payment-journey traceability; incident management and post-incident reviews; and reporting pipelines to Snowflake/… modes, and reconciliation. A track record of improving conversion, approval rates, or performance through measurable, data-driven work. Experience improving operational maturity, including SLOs, observability, on-call health, incident response, and continuous reliability work. A history of partnering effectively with Product, Design, and commercial stakeholders while balancing experience, cost ...

AI Engineer (Contract)

Hiring Organisation
Addition
Location
United Kingdom
calling and context management. Build secure, scalable AI back-end services using AWS Lambda, DynamoDB, S3 and other serverless technologies. Implement monitoring, tracing and observability to ensure AI applications are reliable, performant and production-ready. Work closely with engineering, product and client teams to shape AI architectures and contribute … Solid understanding of AWS services, particularly Lambda, DynamoDB and S3. Experience integrating LLMs using prompt engineering, context management and function calling techniques. Knowledge of observability, monitoring and tracing tools such as CloudWatch. Strong software engineering principles with a focus on scalable, maintainable production systems. AWS certifications, particularly AI/ ...

Staff Engineer

Hiring Organisation
17918
Location
London, United Kingdom
engineering standards, best practices and reusable patterns while partnering with Enterprise Architecture and influencing technical direction Drive engineering excellence by improving code quality, testing, observability, reliability and operational practices Support end-to-end delivery by guiding teams through complex technical challenges, improving decision-making, and contributing to planning and risk … data lakes/lakehouse architectures, Iceberg or similar table formats, as well as batch and streaming processing Knowledge of data quality, governance, cataloguing and observability tools (e.g. Datadog), with DBT or AI-assisted engineering practices as a plus Additional Information Your benefits Werea community here that cares as much about ...

Staff Engineer

Hiring Organisation
Stepstone UK
Location
South East London, London, United Kingdom
Employment Type
Permanent
engineering standards, best practices and reusable patterns while partnering with Enterprise Architecture and influencing technical direction Drive engineering excellence by improving code quality, testing, observability, reliability and operational practices Support end-to-end delivery by guiding teams through complex technical challenges, improving decision-making, and contributing to planning and risk … data lakes/lakehouse architectures, Iceberg or similar table formats, as well as batch and streaming processing Knowledge of data quality, governance, cataloguing and observability tools (e.g. Datadog), with DBT or AI-assisted engineering practices as a plus Additional Information Your benefits Werea community here that cares as much about ...

Software Engineer - AI Accelerated Development

Hiring Organisation
Jobleads-UK
Location
United Kingdom
workflows. Implement RAG pipelines and integrate vector stores (e.g., Qdrant, pgvector, Pinecone) to ground agent outputs in trusted data. Instrument agents for evaluation and observability so their outputs can be verified, measured, and audited. Apply AI‐specific risk controls appropriate to a regulated clinical environment, addressing hallucination, determinism, and auditability … contracts, not one‐off experiments. Python proficiency as the preferred language for AI/agent tooling; comfort with async patterns a plus. Evaluation and observability literacy—able to reason about whether an agent’s output is correct and instrument it to prove so. Claude Agent SDK experience—building custom agents ...

Staff Software Engineer - AI Tools

Hiring Organisation
Jobleads-UK
Location
United Kingdom
test frameworks and headless runs, and wire results into CI and dashboards; Implement production LLM integrations (prompting, RAG/tool calls, guardrails/safety, observability, timeouts/retries, and cost controls); Collaborate with studio engineering, tech art, design, and QA to land integrations across varied codebases and pipelines; Document patterns … well‐tested, maintainable code. Nice to Have You have built networking/service integrations for game backends (gRPC/HTTP) and implemented telemetry/observability (logs, metrics, traces); You have packaged SDKs/plugins for multiple studios and managed semantic versioning and migration tooling; You understand console platform requirements ...

Backend Engineer, Web3 & High-Performance APIs

Hiring Organisation
Jobleads-UK
Location
United Kingdom
reliability, maintainability, and developer experience across backend systems Design and implement integrations with smart-contract platforms, blockchain data sources, and external services Contribute to observability, monitoring, incident response, and post-mortem processes Optimise backend services for performance, resilience, and operational efficiency Collaborate closely with product, blockchain, and infrastructure teams … contributing to architectural decision-making Nice to haves Experience working with multiple blockchain networks or non-EVM ecosystems Familiarity with on-chain data indexing, observability tooling, or protocol debugging systems Experience building performance-sensitive systems such as routing engines, risk controls, or high-throughput pipelines Experience working within fintech, crypto ...

AI Engineer (Amazon Bedrock)

Hiring Organisation
Jobleads-UK
Location
United Kingdom
other sources. Apply the A2A (Agent‐to‐Agent) protocol to enable interoperability between agents across systems and workflows. Instrument AI systems with observability and tracing tooling – CloudWatch, spans, and traces – to support debugging, performance monitoring, and compliance requirements. Integrate LLMs into client applications through prompt engineering, context management, and function … strong, applied proficiency in an AWS and AI context. AWS infrastructure – working knowledge of Lambda, DynamoDB, and S3 as components of AI system backends. Observability – experience instrumenting AI systems with tracing, logging, and monitoring tooling (CloudWatch preferred). LLM integration – prompt engineering, tool/function calling, context window management ...

Product Delivery Manager - Public Cloud SRE

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
bottlenecks. Ensure operational readiness and change safety requirements are executed in delivery workflows, improving release outcomes and reducing change-related incidents. Drive delivery of observability improvements (metrics/logs/traces coverage, alert quality, dashboards) that improve signal quality and operational decision-making. Deliver measurable automation and toil reduction, increasing … design, and data analytics Deep experience in multi-cloud platforms, infrastructure services, automation, and operational tooling. SRE domain expertise: SLIs/SLOs, error budgets, observability, incident/problem management, operational readiness, and reliability analytics. Proven transformation leadership across matrixed, global organizations; strong executive communication and stakeholder influence. ABOUT US J.P. ...

Senior Director, Technology Operations

Hiring Organisation
Jobleads-UK
Location
Guildford, England, United Kingdom
telephony lead and specialist team holding the hands‐on work. DevOps and developer experience (run side): CI/CD reliability, environments, deployment, and observability; currently delivered by the outsourced partner, to be shaped and, over time, selectively insourced. Run‐side Service Delivery, including telephony provisioning and change. Run‐side management … hours/on‐call model for the live service, acting as the senior escalation point. Replace firefighting with proactive reliability: root‐cause discipline, observability, and measurable reductions in downtime and change‐failure rate. Platform, telephony, and estate Own the evolution of the on‐prem and cloud estate against a quarterly ...

Staff Software Engineer, AI Reliability Engineering

Hiring Organisation
Jobleads-UK
Location
England, United Kingdom
Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems across the token path. Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers Lead incident response … more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and systematic resilience testing. Have contributed to open-source infrastructure or ML tooling. The annual compensation range for this ...

IT Infrastructure Solutions Architect

Hiring Organisation
Jobleads-UK
Location
Cambridge, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...