1,501 to 1,525 of 4,175 Observability Jobs

Lead AI Engineer

Location
Manchester, England, United Kingdom
traffic, within real latency budgets and reliability realities. Hold a high engineering bar on AWS and TypeScript — clean CI/CD, infrastructure as code, observability, testing, and LLMOps for running model‐backed systems in production. Shape delivery against the roadmap with the VP, product managers, data analysts and Staff Engineers … powered applications. Extensive experience building and operating cloud‐native applications on AWS, with strong knowledge of CI/CD, infrastructure as code, observability and modern engineering practices. Experience delivering high‐performance, real‐time systems that operate reliably at scale with demanding latency requirements. Demonstrated experience leading the successful transition ...

Infrastructure Python Developer

Hiring Organisation
Experis
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£350 - £402/day
deploy application/services using Docker and Kubernetes Administration for Kubernetes Resources (Pods, Ingress, Services, Secrets, CRDs, etc) Nice to have Exposure on enhancing observability with knowledge of tools such as Prometheus, Grafana, and OpenTelemetry. Advantageous to have enterprise tools knowledge (i.e., Control M, True sight, Guardium, Tenable Nessus, Delinea ...

Principal Product Engineer

Hiring Organisation
Zapp
Location
London, UK
Employment Type
Full-time
home in NestJS (or a similar Node.js framework) and React. Cloud and DevOps minded. Comfortable across GCP or a similar cloud, CI/CD, observability, and modern infrastructure tooling. Use AI daily. AI assistants are part of how you build. You bring back patterns that help others get more … code. You measure your work by what changed for the customer, not the lines of code shipped. Production minded. Solid grasp of API design, observability, scaling, reliability, and security. Care about craft. Clean APIs, attention to detail, and how the system feels to work in. Comfortable in uncertainty. You move ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
data retrieval layers (pgVector, Pinecone) and ensure efficient embedding pipelines for AI contexts. Partner with AI teams to monitor data latency, cost efficiency, and observability metrics. Collaboration & Governance Partner with Platform Operations and Security to enforce privacy, compliance, and access control frameworks (GDPR, SOC2). Work cross‐functionally with analysts … platform data (Google Ads, Meta, TikTok, DV360, Amazon Ads). Demonstrated expertise in code versioning (GitHub), CI/CD integration, and data observability practices. Ability to write clean, modular, testable code and review peers’ contributions for maintainability and performance. Additional Information Publicis Groupe has fantastic benefits on offer ...

Cloud Advisory Architect

Location
Greater London, England, United Kingdom
where GenAI and Agentic play a role. Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps. Manage & Mentor: Lead teams of architects and engineers, providing technical coaching, career counselling, performance management, and coaching ...

Infrastructure/DevOps Engineer

Location
Greater London, England, United Kingdom
enjoys building strong, cost-efficient infrastructure. Your work will keep systems audit-ready and high-performing. In this senior role, you’ll lead DevOps, observability, and compliance. Tasks include architecting cloud infrastructure for growth, prepping for SOC 2/ISO 27001 audits, and setting up CI/CD pipelines ...

Agentic Commerce Architect

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
intelligence — including embedding pipelines, vector retrieval, semantic search, and structured product data — ensuring agent outputs are accurate, grounded, and commercially reliable Embedding AI governance, observability, and responsible AI principles into architecture design — including audit logging, human-in-the-loop escalation points, and performance monitoring — as first-class concerns rather than … deployment, including familiarity with cloud-native services and CI/CD-based delivery practices Understanding of AI governance and responsible AI principles — including observability, tracing, auditability, and how to build guardrails into production AI systems Preferred Experience Experience contributing to technical proposals, reference architectures, or delivery accelerators in a consulting ...

Senior Software Engineer (Java)

Location
City Of London, England, United Kingdom
market connectivity workflows Knowledge of Linux engineering, troubleshooting, and performance optimisation Experience with Spring Boot or Google Guice dependency injection frameworks Experience with observability stacks (Open Telemetry, Grafana) Experience with distributed caching solutions such as Hazelcast Experience with BDD and automation frameworks (Cucumber) #J-18808-Ljbffr ...

Consultant - Service Management Transformation

Location
Greater London, England, United Kingdom
service management, including intelligent automation, virtual agents, knowledge management, workflow optimisation and service analytics.AIOps & Intelligent Operations – Develop an understanding of modern operational practices including observability, monitoring, intelligent alerting, event management and AIOps capabilities, supporting clients in improving operational performance through automation and data-driven insights.Cloud & Digital Operations – Support the design … automation can be applied within service management to improve operational efficiency, employee experience and service outcomes.Awareness of modern AIOps and observability capabilities, including monitoring, event correlation, anomaly detection, intelligent alerting, operational analytics and automation.Familiarity with enterprise service management platforms such as ServiceNow, Jira Service Management, Freshworks or similar technologies.Experience contributing ...

Data Platform Engineer- BPL-CIO

Location
Greater London, England, United Kingdom
automation, GitOps operating models AWS Platform Engineering Hands‐on AWS engineering with a focus on: IAM and security patterns, Networking integration, Storage and encryption, Observability, Resilience and operational readiness, Automation and supportability Databricks on AWS Experience with: Workspace deployment, Databricks Terraform provider, Identity integration, Storage integration, Network connectivity patterns Platform … Engineering/Internal Developer Platforms Experience building: Self‐service capabilities, Golden paths, Paved roads, Reusable engineering patterns, Internal platform products Observability, Reliability & Operational Engineering Experience with: Monitoring and alerting, Operational telemetry, Incident management, Resilience testing, Recovery automation, Production support You may be assessed on the key critical skills relevant ...

Software Engineer (Backend)

Location
Greater London, England, United Kingdom
engineering, performance, and reliability problems that sit outside of their scope. Contribute to the shared foundations other engineers depend on: deployment pipelines, service templates, observability, and the environments their work runs in. You should apply if You have built and worked on production backend systems. You will spend your time … whether that means learning a new skill, reaching out to stakeholders to clarify requirements, or suggesting an alternate approach. You treat quality, security, and observability as engineering fundamentals rather than optional extras, and you flag technical debt rather than letting it accumulate silently. You actively seek feedback ...

Principal AI Engineer - Microsoft Azure AI Foundry (Contract)

Hiring Organisation
Hackajob Ltd
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
tooling. Define reusable architecture patterns for AI workloads, including development, testing and production environments. Establish platform standards covering resource structure, environments, deployment patterns, observability and operational management. Work closely with AI Engineers to design and deliver AI Agents and agentic workflows that meet business and technical requirements. Governance, Risk & Compliance … provisioning and configuration using Bicep, Terraform or ARM templates. Build automated processes for prompt-flow evaluation, model deployment, testing and release management. Establish platform observability covering availability, performance, usage, cost and AI workload health. Experience 5+ years' experience in Azure cloud architecture, engineering or platform engineering. 12+ years' experience specifically ...

Lead Platform Engineer

Location
Manchester, England, United Kingdom
using Terraform Supporting container platforms using Kubernetes and Docker Developing CI/CD pipelines to automate delivery and deployment Improving platform monitoring, reliability and observability Providing technical leadership and mentoring engineers across delivery teams LEAD PLATFORM ENGINEER ESSENTIAL SKILLS Infrastructure as Code experience using Terraform Container technologies such as Kubernetes ...

Agentic Commerce Architect

Location
City Of London, England, United Kingdom
intelligence — including embedding pipelines, vector retrieval, semantic search, and structured product data — ensuring agent outputs are accurate, grounded, and commercially reliable Embedding AI governance, observability, and responsible AI principles into architecture design — including audit logging, human-in-the-loop escalation points, and performance monitoring — as first-class concerns rather than … deployment, including familiarity with cloud-native services and CI/CD-based delivery practices Understanding of AI governance and responsible AI principles — including observability, tracing, auditability, and how to build guardrails into production AI systems Preferred Experience Experience contributing to technical proposals, reference architectures, or delivery accelerators in a consulting ...

Director, AI Engineer (Remote Eligible)

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
development, testing, deployment, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Make high judgment build-vs-buy decisions across a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs … execution plans across multiple product areas, balancing innovation with delivery discipline Scale AI engineering practices across teams through shared infrastructure, reusable components, and unified observability and governance frameworks Establish enterprise standard for Responsible AI, including fairness metrics, model evaluation protocols, documentation requirements, and audit readiness Partner with research, compliance ...

Director, AI Engineer (Remote Eligible)

Hiring Organisation
Capital One
Location
Richmond, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
development, testing, deployment, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Make high judgment build-vs-buy decisions across a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs … execution plans across multiple product areas, balancing innovation with delivery discipline Scale AI engineering practices across teams through shared infrastructure, reusable components, and unified observability and governance frameworks Establish enterprise standard for Responsible AI, including fairness metrics, model evaluation protocols, documentation requirements, and audit readiness Partner with research, compliance ...

Director, AI Engineer (Remote Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
development, testing, deployment, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Make high judgment build-vs-buy decisions across a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs … execution plans across multiple product areas, balancing innovation with delivery discipline Scale AI engineering practices across teams through shared infrastructure, reusable components, and unified observability and governance frameworks Establish enterprise standard for Responsible AI, including fairness metrics, model evaluation protocols, documentation requirements, and audit readiness Partner with research, compliance ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Carmarthenshire, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Lancashire, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Herefordshire, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Suffolk, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Anglesey, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Cornwall, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Nottinghamshire, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...

Remote Engineering Manager, Cloud Platform (UK)

Hiring Organisation
Flosum
Location
Wokingham, United Kingdom
8+ years of software engineering experience, with 3+ years leading engineering teams Deep expertise architecting distributed systems on AWS (compute, networking, data, IAM, observability) Strong Node.js background and a clear point of view on building maintainable services at scale Experience designing platform-agnostic systems that can run on any major ...