1,426 to 1,450 of 5,223 Permanent Observability Jobs

Senior Cloud Infrastructure Engineer - Scale & Observability

Location
Greater London, England, United Kingdom
scale. You will own tooling in Python/Go, extend Kubernetes and Docker usage, and help advance multi-cloud readiness. You will contribute to observability through Prometheus/OpenTelemetry/Grafana, strengthen DR/BCP, and collaborate with a talented team to deliver robust, scalable infrastructure across customers. #J ...

Senior Cloud Platform Engineer – Multi-Cloud & Observability

Location
Greater London, England, United Kingdom
Thought Machine is hiring Senior Software Engineers for Infrastructure to deploy and maintain cloud-native platform infrastructure. You will work on multi-cloud orchestration, observability, and data layers to reduce cognitive load for developers and clients at scale. The role focuses on building resilient, scalable tooling with Python or Golang ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions Ltd
Location
London, United Kingdom
Employment Type
Full-Time
Salary
£780.00 - £830.00 per day
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone … shape how thousands of infrastructure and application events are detected, correlated and actioned across EMEA, and you'll have genuine scope to modernise observability capability rather than simply keep the lights on. This is a hands-on technical leadership role with no direct reports, so your influence comes from your ...

Senior Product Manager for AI Observability

Location
Greater London, England, United Kingdom
Role Profile As part of the LSEG AI, we are hiring a Senior Product Manager–AI Observability to own the strategy, design and rollout of telemetry systems that monitor, measure and analyse how AI models, MCPs, and AI-enabled features behave across all LSEG products. This role is responsible … model performance tracking cost efficiency user experience optimisation operational reliability auditability and the long‐term evolution of our AI platform. Key Responsibilities & Accountabilities Telemetry & Observability Strategy Define the end‐to‐end telemetry vision and roadmap for LLMs, MCPs, vector stores, embeddings, inference layers and AI‐powered user experiences. Establish ...

Senior Product Manager for AI Observability

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
Role Profile As part of the LSEG AI, we are hiring a Senior Product Manager – AI Observability to own the strategy, design and rollout of telemetry systems that monitor, measure and analyse how AI models, MCPs, and AI-enabled features behave across all LSEG products. This role is responsible … model performance tracking cost efficiency user experience optimisation operational reliability auditability and the long-term evolution of our AI platform. Key Responsibilities & Accountabilities Telemetry & Observability Strategy Define the end-to-end telemetry vision and roadmap for LLMs, MCPs, vector stores, embeddings, inference layers and AI-powered user experiences. Establish ...

Senior Developer Advocate - Data Observability

Location
Greater London, England, United Kingdom
They own projects from beginning to end, facilitating collaboration and enabling Datadog's community to solve real-world problems. With a focus on Data Observability, this role will enable our community of engineers around Datadog to be part of a movement of building better software. This is a unique opportunity … engineering expertise and advocacy skills to shape the ever‐evolving technological landscape. What You’ll Do: Act as a subject matter expert for data observability on behalf of the Datadog advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader ...

Site Reliability Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
software engineering and infrastructure. You'll develop internal platforms, tooling, and automation across Linux, distributed systems, and cloud-native technologies, helping improve reliability, observability, and operational efficiency across a global production environment. Responsibilities: Design and develop internal infrastructure tooling and automation. Build and maintain monitoring, observability and configuration management platforms. …/Must Have: Strong experience programming with Python, Go and/or C++ Strong Linux knowledge and understanding of distributed systems. Experience with monitoring, observability or SRE practices. Experience with CI/CD pipelines, Git and infrastructure automation. Familiarity with Kubernetes and containerised workloads. Strong analytical and troubleshooting skills. Benefits ...

Lead DevOps Engineer - Real Time Platform

Location
Greater London, England, United Kingdom
global scale. This is an individual contributor leadership role , where you’ll define infrastructure architecture, raise operational standards, and ensure resilience, security, and observability across a mission‐critical platform. AI-First Engineering This team operates with an AI-first approach. We expect hands‐on experience with AI development tooling: terminal … augmented IDEs, and automated workflows. You are ultimately accountable for production quality, security, and correctness. This means owning infrastructure review, security validation, system observability, and operational guardrails. WHAT YOU'LL DO Design and operate cloud infrastructure on AWS to support low‐latency, always‐on real‐time workloads. Own infrastructure ...

Senior / Principal Applied AI Engineer (UK / Europe, Remote)

Location
United Kingdom
production-grade AI agents, agentic workflows, and LLM integrations for internal automation and in-product features for a web-native trading platform. Own implementation, observability, and security while partnering with Product and Platform teams to deliver measurable AI systems in production. Job Description Role Senior/Principal Applied AI Engineer … integrate LLMs into internal systems and the trading product. This is a hands-on engineering role: you will write production code, put evaluation and observability on everything shipped, and operate autonomously in a lean team. Key Responsibilities Build agent loops, tool-calling, structured outputs, planning/state management, retries, guardrails ...

Observability SRE AVP: Cloud Observability & Migrations

Location
Greater London, England, United Kingdom
Citi is seeking an experienced Site Reliability Engineer to lead end-to-end observability and resiliency initiatives in a large-scale environment. You will migrate monitoring tooling to Google Cloud Observability and Grafana, implement OpenTelemetry instrumentation, and author reusable deployment solutions for OpenShift/Kubernetes and VM environments. The role ...

SRE / Platform Engineer - Remote

Hiring Organisation
Genesis10
Location
New York, United States
Employment Type
Permanent
Salary
USD 105 Hourly
infrastructure and operational problems. Engineers on the team write code every day and work across application and infrastructure layers to improve reliability, performance, scalability, observability, and system integration. A major initiative for the team is establishing a centralized observability capability across an environment where monitoring and operational data have historically … been siloed. The organization is bringing telemetry together using Datadog and enterprise data lake capabilities, creating a common observability foundation that can ultimately support AIOps, agentic AI, automated remediation, and self-healing systems. This is not a traditional operations or Solutions Architecture position. The successful candidate will be expected ...

Core Platform Developer

Location
City Of London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment. Key Responsibilities Design and build shared backend services, frameworks … developer tooling that support internal application and service development. Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities. Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams. Help ...

DevOps & Infrastructure Engineer

Location
Gloucester, England, United Kingdom
Security customers, spanning both on-premise environments and cloud-based solutions. You’ll lead hands-on DevOps and infrastructure engineering across CI/CD, observability, infrastructure-as-code and platform automation, helping teams build secure, reliable and scalable services in demanding environments. What you’ll be doing: You’ll lead … cloud-based solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver ...

MLOps Engineer

Hiring Organisation
DGH Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
platform reliability. Key Responsibilities - Design, deploy, and manage AI platforms and agent infrastructure - Build and maintain CI/CD pipelines and DevOps workflows - Implement observability, monitoring, and logging solutions - Optimise performance, scalability, and cost efficiency - Support AI teams with infrastructure, deployment, and integration - Ensure platform security, compliance, and high availability …/CD, automation, and DevOps best practices - Experience with Kubernetes/containerisation technologies - Strong programming skills (e.g. Python, Go, Node.js) - Experience with observability tools (e.g. OpenTelemetry, Datadog) - Understanding of security, performance optimisation, and scalability Desirable Skills - Experience working on AI/ML platforms or deployments - Exposure to large-scale distributed ...

Senior Software Engineer, GoLang

Location
Greater London, England, United Kingdom
continuous integration and delivery (CI/CD) Make data-guided decisions affecting core business metrics and processes Apply platform and reliability engineering practices, including observability, performance optimisation, analytics, and security best practices Facilitate collaboration between teams and promote continuous improvement Mentor junior engineers on engineering practices, coding standards, and troubleshooting … DevOps Continuous Integration and Delivery (CI/CD) Infrastructure-as-Code Hard Skills Application Development Performance Optimisation Automation Development Containerisation Development Methodologies Design Patterns Observability Analytics Security Best Practices Troubleshooting Soft Skills Mentoring Collaboration Continuous Improvement Industry Keywords Public-Facing Systems Internal Insurance Systems Tools & Technologies AWS Lambda DynamoDB Azure ...

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering … solving, prioritization, and decision‐making skills. What We’re Looking For: Basic Required Qualifications: 10+ years of experience in product management, IT operations, SRE, observability, platform engineering, or related enterprise technology roles. Strong understanding of AIOps concepts, including event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics ...

AI Native SW Engineering

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … Vertex AI, and open-source models with fallback routing, token, cost, and latency management ImplementLLMOpsin production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions workshops, proofs ...

AI Native SW Eng

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
Design and build production-grade agentic systems end-to-end: multi-agent orchestration, RAG pipelines, policy-based routing, tool invocation, memory management, and lifecycle observability Build and own RAG pipelines: embeddings, chunking strategy, vector search, context window engineering and tuning against real quality targets Integrate and abstract across multiple … Vertex AI, and open-source models with fallback routing, token, cost, and latency management ImplementLLMOpsin production: eval harnesses with real quality metrics, prompt versioning, observability tooling (LangSmith, Braintrust, or equivalent), cost and safety monitoring Embed directly with client engineering teams to design, prototype, and deploy agentic solutions workshops, proofs ...

Software Engineer - Cloud Compute Platform

Location
Hursley, England, United Kingdom
Contribute To Include Workload orchestration across Kubernetes clusters Platform APIs and Kubernetes operators Cloud platform integrations Multi-tenant workload isolation and security Observability, health checks, and operational tooling Workload scheduling, disruption management, and rolling updates What You Will Do Design and implement services, APIs, and Kubernetes controllers using Go. Contribute … engineers to understand requirements and deliver reliable solutions. Participate in technical design discussions and help evaluate implementation trade-offs. Improve the reliability, scalability, observability, and maintainability of existing systems. Write automated tests, documentation, and operational runbooks. Participate in code reviews and provide constructive feedback to teammates. Help investigate and resolve ...

AI Native SW Engineering

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
govern production-grade agentic systems at enterprise scale: multi-agent orchestration across complex environments, RAG pipelines, policy-based routing, memory management, andprogramme-level lifecycle observability Define RAG pipeline standards across engagements:establishchunking and embedding strategies, set quality benchmarks, and ensure metric-backed tradeoff decisions are documented and transferable Set multi … default, fallbackroutingand cost governance as standard design practice across providers including OpenAI, Anthropic, Vertex AI, and open-source models OwnLLMOpsatprogrammescale: eval strategy, prompt governance, observability tooling standards, safetymonitoringand cost controls across multiple concurrent systems Lead client engineering engagements at senior level facilitatearchitecture design sessions, lead proof-of-concept delivery ...

Platform Engineer

Location
Milton Keynes, England, United Kingdom
Terraform, CloudFormation, or CDK), CI/CD pipelines, API Gateway, Lambda, and Aurora PostgreSQL — and confident applying fundamentals such as high availability, fault tolerance, observability, and cost control in a live environment. Payments or fintech exposure is an advantage but not required; what matters more is a practical, first-principles … rollback processes Support deployment of AI generated applications and tooling Support change control and release management alongside Engineering and the Information Security Officer Reliability, Observability & Data Infrastructure Design and maintain systems for high availability, fault tolerance, and resilience by default Implement logging, monitoring, and alerting (for example CloudWatch) across services ...

AI Platform & Site Reliability Engineering Managing Consultant

Location
United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services. You will work with technology, engineering, operations … operational requirements. AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. Reliability Engineering & SRE: Establish SRE practices including ...

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Workday Integrations Product Engineer

Location
Hook, England, United Kingdom
ecosystem supporting our global workforce. This is a hands‐on engineering role with a strong focus on Workday integration development, architecture, automation, reliability, security, observability, and operational excellence. Your Responsibilities: Design, develop, test, deploy, and support robust integrations between Workday and enterprise applications using Workday Studio, EIB, Core Connectors, Cloud … technology ecosystem. Apply strong software engineering principles including version control, peer reviews, automated testing, CI/CD, and structured release management. Build integrations with observability, operational resilience, proactive monitoring, alerting, logging, automated recovery, and self‐healing error handling by design. Troubleshoot complex production issues, conduct root‐cause analysis, and continuously ...

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services.You will work with technology, engineering, operations … operational requirements.• AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices including ...