1,526 to 1,550 of 1,764 Permanent Observability Jobs

Mainframe Hardware Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Mainframe Hardware Infrastructure Engineer You’ll engineer infrastructure technology for public and private cloud environments, complying with security, resilience, sustainability, and operational requirements with observability and guardrails built in You’ll also use automation to provide testing and a route to live for the product, working with customers to help … replication, virtual tape technologies and disaster recovery automation Working knowledge of the development of CI or CD pipelines using modern tooling Experience of using observability tools and techniques and the ability to use data, information and user sentiment to continuously improve solutions Public cloud vendor knowledge covering ...

Senior Vice President, Full-Stack Engineer

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise-scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. Qualifications Bachelor’s or Master’s degree … platform design and optimisation. Proven track record of delivering production-grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands-on technical leader who can actively contribute to solution design and critical builds, while defining ...

Senior AI-Driven Sales Engineer: Observability & Cloud

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technical resource for the Sales team. You will drive technical closing of opportunities by demonstrating the value of Observe to prospective customers, combining deep observability expertise with strong presentation skills. The role requires experience in modern observability tools and cloud-native architectures, with travel to client sites and events ...

Platform Operations Director

Hiring Organisation
ClearCourse
Location
City of London, London, United Kingdom
business continuity across the group. Internal IT & Systems Manages internal IT and business systems administration (M365, NetSuite, SuccessFactors, SharePoint) -infrastructure, integrations, and IAM. Ensures observability and SRE capability is fit for purpose across cloud, hosted, and end-user environments. Vendor & Cost Management Drives cloud and vendor cost discipline - manages …/CD infrastructure requirements. • Head of Infrastructure & Cloud - Direct report. Hosting strategy, cloud platform, and FinOps execution. Head of SRE - Direct report. Observability, on-call, and DR/BCP processes. • Head of Internal Services - Direct report. Internal IT, business systems, and end-user support. Finance - Direct report. Cloud cost visibility ...

Senior Presales Consultant (Enterprise Payments, Technical)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
site) You will lead and contribute to RFI and RFP responses, providing detailed technical input across areas such as integration, security, resiliency, observability, and compliance You will identify risks, gaps, and constraints, and clearly articulate assumptions and solution trade‐offs You will plan and support Proofs of Concept (POCs), enabling … REST APIs and integration patterns Security and identity (OpenID Connect/OAuth) Cloud platforms (AWS, Azure, GCP, IBM Cloud, Oracle Cloud) Monitoring and observability tools (e.g. Splunk, SIEM) Pre‐Sales & Customer‐Facing Experience Proven experience in a customer‐facing technical, solution architecture, or pre‐sales role Confident presenting to senior ...

Senior AI Engineer

Hiring Organisation
MarkIT Placements
Location
West London, London, United Kingdom
Employment Type
Permanent, Work From Home
execution Deploy AI systems into cloud, on-premises, and air-gapped environments Build production-ready pipelines from data ingestion through to inference Experience with observability for AI systems, including agent behaviour, model performance, and failure modes Collaborate with engineers, product leads, and customers to translate requirements into working systems Contribute … with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for monitoring and managing deployed applications MLOps experience - model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter-agent communication protocols (e.g. A2A) Familiarity with MCPs ...

ML Platform Engineer

Hiring Organisation
Jobleads-UK
Location
United Kingdom
operation across research and product environments. Improve workload scheduling, monitoring, debugging, and resource management for GPU-based and cloud infrastructure environments. Drive improvements across observability, automation, reliability, developer experience, and platform usability. Create abstractions and developer tools that enable engineering teams to work more effectively with complex ML systems. Collaborate … infrastructure, model serving systems, or data-intensive workloads. Experience working with GPU-based systems, performance-sensitive environments, or large-scale computing resources. Familiarity with observability, monitoring, and debugging practices for distributed systems. Knowledge of infrastructure and development tools such as Terraform, Datadog, GitHub Actions, or similar technologies. Ability to work ...

HPC Engineer, Metal Net

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
NVSwitch platforms across large data centre environments. This role is a strong fit for engineers who enjoy production troubleshooting, hardware‐adjacent systems work, automation, observability, and learning specialized infrastructure deeply. You will be responsible for troubleshooting Linux, networking, hardware, firmware, performance, and stability issues in production, while building automation … Preferred: Experience with Ansible or other infrastructure‐as‐code and configuration automation tooling. Kubernetes application development or live platform operations experience. Familiarity with modern observability systems, including Grafana, Prometheus, PromQL, or similar stack components. Experience managing large fleet operations across Linux systems, network devices, GPUs, or infrastructure components. Deep understanding ...

AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Develop evaluation frameworks that measure model performance across accuracy, latency, cost, reliability, and safety Ship Production-Grade Systems: Ensure systems meet high standards for observability, reliability, security, and maintainability Raise the Engineering Bar: Improve development workflows, evaluation practices, and deployment strategies as our AI platform continues to evolve … Guardrails, MCP, or agent frameworks Experience implementing RAG pipelines, tool use, or multi‐step AI workflows Strong understanding of AI system evaluation, debugging, and observability Experience building reliable production systems with modern DevOps practices Experience deploying AI systems in enterprise environments is a plus Experience working across cloud platforms ...

Senior Software Engineer, Event Streaming Systems

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
reliable, event-driven products at scale. Design, build and operate large-scale event streaming infrastructure that supports critical production workloads. Improve the reliability, scalability, observability and automation of Kafka, EventBus and related streaming platforms. Debug complex distributed systems issues, including performance bottlenecks, capacity constraints and production reliability challenges. Build infrastructure … with a focus on reliable, maintainable production systems. Experience operating scalable production infrastructure, including debugging, troubleshooting and improving reliability. Strong understanding of infrastructure automation, observability, capacity management and operational excellence. Good fundamentals in Linux, networking and JVM-based systems, or curiosity to deepen expertise in these areas. Strong collaboration ...

AI Technical Project Manager / AI Program Manager

Hiring Organisation
BC Forward
Location
Chicago, Illinois, United States
Employment Type
Permanent
Salary
USD 6,769 Annual
risk, bias and fairness, explainability, and human-in-the-loop controls. Drive operational readiness for production support including runbooks, incident playbooks, SLAs/SLOs, observability, and ServiceNow Knowledge Management alignment. Partner on release planning, change management, and operational handoffs and lead post-launch retrospectives and continuous improvement. Maintain project artifacts … integration patterns, vector and relational database basics, and LLM evaluation metrics including quality, safety, latency, and cost. Exposure to LLMOps or MLOps practices including observability, rollback, feature flags, and CI/CD for models. Financial services or regulated industry experience and production support background. Certifications such ...

Senior Monitoring/Observability Architect

Hiring Organisation
BC Forward
Location
Pennington, New Jersey, United States
Employment Type
Permanent
Salary
USD 7,023 Annual
Title: Senior Monitoring/Observability Architect or Production Support Architect Location: Pennington, NJ/Phoenix, AZ Duration: Contract - 11 months Pay Range: $70.23/hr (W2) Job ID: 407285 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune ...

Security Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Auth provider, expanded multi‐tenant APIs, connector executions, agent tool‐calling paths etc. Build detection and response capability around credential and authentication flows, with observability that closes incidents fast. Partner with engineering to raise the bar day‐to‐day: architecture reviews, written standards, and security embedded in code review. … Terraform Security tooling: Aikido (SAST, DAST, container scanning, pen testing), 1Password, GitHub (org‐level enforcement, Advanced Security) Compliance & ops: Drata, Iru, EasyLlama Observability & IR: Datadog, Sentry, Logfire, Incident.io Languages: TypeScript (Node.js), Python Benefits Meaningful share options (EMI) - share in the company’s success as we grow 25 days holiday + ...

Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
manage dependencies and deliver high‐quality solutions. Operational Focus – Familiarity with monitoring, incident response and maintaining service reliability at scale, including working with SLOs, observability platforms and optimising for cost efficiency. Process & Best Practices – Knowledge of modern software development workflows, including CI/CD, testing strategies and iterative improvement … high‐scale SaaS or cloud‐based environment. Hands‐on experience with modern DevOps practices, including CI/CD and cloud infrastructure. Familiarity with observability tools, monitoring and error tracking. Passion for building a strong engineering culture focused on ownership, collaboration and continuous improvement. At DISCO, our employees have told ...

Global Network Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
access controls* Ensure compliance with security best practices and enterprise policies.* Collaborate closely with data security teams to maintain and protect network integrity**Automation & Observability*** Drive network automation and infrastructure-as-code using:* Python, Ansible, APIs* Build and enhance monitoring, telemetry, and analytics platforms to track:* Network performance* Capacity … Experience integrating on-prem networks with cloud environments**Automation & Tooling*** Proficiency in Python and/or automation frameworks (Ansible)* Experience with network monitoring and observability tools* Familiarity with APIs and infrastructure-as-code practices**Certifications (Preferred)*** CCNP/CCIE (or equivalent experience)* Relevant cloud or security certifications (AWS, Azure, Zscaler ...

Windows Server Engineering Lead

Hiring Organisation
Jobleads-UK
Location
Knutsford, England, United Kingdom
role in advancing our Core Infrastructure Platform strategy. You will define and deliver the roadmap for the Windows Server platform, driving standardisation, automation, security, observability, and continuous service improvement across a large-scale enterprise estate. As part of the product leadership team, you will lead the adoption of new operating … Security and compliance, including Windows Server hardening, vulnerability and patch management, and privileged access management technologies such as BeyondTrust Endpoint Privileged Management (EPM). Observability and telemetry platforms, including Observe and/or Elastic. Cloud and hybrid infrastructure experience across VMware, AWS and Azure. You may be assessed ...

Quality Test Engineer III

Hiring Organisation
Elsevier
Location
Greater London, United Kingdom
Employment Type
Full Time
REST and evolving GraphQL APIs. Integrate and maintain automated tests within CI/CD pipelines, supporting testing at multiple levels of the stack. Use observability tooling such as New Relic and logging to support diagnosis, monitoring, and quality analysis. Technically analyse problems and proposed solutions, recognising the right level … Karate, Cucumber with RestAssured. Good exposure to CI/CD pipelines, supporting the integration of testing at various levels of the stack. Exposure to observability tooling such as New Relic or logging. The ability to technically analyse a problem and a solution, recognising the right level of detail and abstraction ...

Applied AI Product Builder (FDE based in Singapore)

Hiring Organisation
Riverchelles International
Location
London Area, United Kingdom
Partner with architecture, cloud, platform and DevOps teams to deliver solutions across cloud, hybrid and enterprise environments. Establish automated testing, deployment, evaluation, monitoring and observability practices. Monitor application quality, model performance, latency, reliability, availability, usage and cost. Apply sound engineering practices, including version control, CI/CD, documentation and code … APIs or orchestration frameworks. Understanding of cloud platforms, APIs, databases, data pipelines, containerisation, enterprise integration and production software practices. Experience with model evaluation, AI observability, guardrails, responsible AI or production monitoring. Ability to translate customer needs into practical product and technical solutions. Strong problem-solving, communication and execution capabilities, with ...

Software Engineer - Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
regressions are caught before production. Take end‐to‐end ownership from discovery and design through build, rollout, and operational excellence. Instrument systems with the observability, cost tracking, and audit trails needed to know when they degrade. Apply FinOps and cost‐optimization practices to AI workloads, tracking and managing token, inference … example LangChain, LlamaIndex, or the Model Context Protocol). Experience with MLOps/LLMOps tooling such as experiment tracking, model versioning, monitoring, evaluation/observability platforms, or CI/CD for ML and LLM systems. Experience implementing responsible‐AI or AI‐governance controls such as guardrails, human‐in‐the‐loop ...

Senior Software Engineer - Global KYC and Onboarding - Java

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
# Senior Software Engineer - Global KYC and Onboarding - Javaat Wise • Worship Square, 65 Clifton Street, London, United KingdomBack to jobs1. Home2. Jobs3. United Kingdom4. Senior Software Engineer - Global KYC and Onboarding - Java12h agoW## Senior Software ...

Research Software Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and ...

AI Deployment Consultant

Hiring Organisation
CGI
Location
United Kingdom
Employment Type
Full Time
This job is with CGI, an inclusive employer and a member of myGwork – the largest global platform for the LGBTQ+ business community. Please do not contact the recruiter directly. AI Deployment Consultant Position Description Help ...

Engineering Manager – Payments, BACS, CHAPS

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Responsibilities Deliver and continuously improve payment services across UK payment rails (BACS, CHAPS, SWIFT). Ensure systems meet high standards for availability, resilience, security and regulatory compliance. Own delivery outcomes in partnership with Product, Technology ...

Staff Engineer (Tech Lead)

Hiring Organisation
Talent Locker
Location
Warminster, Wiltshire, South West, United Kingdom
Employment Type
Permanent
Staff Engineer (Tech Lead) - Warminster, Hybrid - Security Cleared - Up to £120,000 We're looking for an experienced Staff Engineer/Technical Lead who thrives on solving complex engineering challenges and wants to play a ...

Senior Reliability Engineer – Platform & Observability

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Identity’s OneLogin team is seeking Senior Software Engineers who own production systems and drive reliability, observability, and operability across the stack. You’ll tackle complex issues, design for resilience, and apply AI-assisted development approaches to accelerate debugging and root-cause analysis. The role emphasizes end-to-end ownership ...