2,476 to 2,500 of 5,534 Observability Jobs

Software Engineer - AI Agents (Satori)

Location
Belfast City District, Northern Ireland, United Kingdom
means at Proofpoint: the eval suites, safety checks, telemetry, and AuthX that gate every rollout. The primitives you help shape — from prompt management to observability — influence how mission-critical agentic software ships across the company. You will partner with product tech leads to turn one-off integrations into reusable platform … LangChain, or similar) and keep up with the field. Experience contributing to agentic or LLM-based systems in production, including familiarity with the eval, observability, and rollback story. Demonstrated learning velocity and curiosity; driven to understand systems beyond surface-level usage at enterprise scale. Development experience with Python; TypeScript/ ...

Data Scientist - BAU Analytics

Hiring Organisation
Executive Facilities
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£500.00 per day
sources that help build the AI/ML user experience Build and maintain scalable data pipelines for BigQuery using Cloud Composer and Airflow. Provide observability and monitoring using Monte Carlo and Looker as well as operational tools such as NewRelic and Splunk, driving reliability, data quality, data quality and robustness. ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
captures and governs new categories of data (e.g. smart badge telemetry, location/proximity signals) in a privacy-compliant, lawful-basis-aware way Build observability into data systems — structured logging, freshness and quality monitoring, SLOs, and pipeline health dashboards Champion high-quality technical communication: proposals, specifications, and documentation that other … product engineering teams, even if your focus is data and AI infrastructure DevOps fluency: AWS, Kubernetes/EKS, Terraform, CI/CD pipelines Excellent observability practices — structured logging, metrics, distributed tracing, SLOs (Datadog, Sentry) Feature flags, canary deployments, and gradual rollout patterns Track record of driving data quality, governance ...

Senior DevOps Engineer CGEMJP00357990

Hiring Organisation
Experis
Location
Nationwide, United Kingdom
Employment Type
Contract
Build and manage AWS infrastructure using Terraform. Automate environment provisioning and routine operational processes. Maintain Docker-based deployment and runtime patterns. Improve platform resilience, observability and security. Lead incident investigation, remediation and platform upgrades. All profiles will be reviewed against the required skills and experience. Due to the high number ...

Data Scientist - BAU Analytics

Location
City Of London, England, United Kingdom
sources that help build the AI/ML user experience Build and maintain scalable data pipelines for BigQuery using Cloud Composer and Airflow. Provide observability and monitoring using Monte Carlo and Looker as well as operational tools such as NewRelic and Splunk, driving reliability, data quality, data quality and robustness. ...

Software Engineer - Affiliate Operations

Location
Greater London, England, United Kingdom
remain highly reliable while evolving to support new markets and acquisition channels. Reliability is fundamental to everything we build. We invest heavily in automation, observability, and operational excellence to reduce manual effort. We are also exploring how AI can transform our engineering productivity and marketing platforms. We operate … improve platform effectiveness, launch new capabilities, and reduce manual operational effort across the business. Improve You'll continuously improve our systems through better observability, automation, and thoughtful refactoring. You'll help evolve our architecture and engineering practices to ensure our platforms remain resilient as they scale. Own You'll take ...

Cloud DevOps Platform Engineer

Location
Greater London, England, United Kingdom
their failure modes, and a good understanding of network concepts and fundamentals Expertise in managing and maintaining Kubernetes clusters in production Excellent knowledge on observability tooling and best practices Experience working within or alongside engineering teams to deliver Preferred qualifications, capabilities and skills: Systems and database experience, with an understanding ...

Technical Lead, Lending & Savings

Hiring Organisation
Blockchain
Location
London, UK
Employment Type
Full-time
performance, security, and maintainability. Remain hands-on, contributing production-quality code and reviewing critical changes. Drive engineering best practices across system design, testing, deployment, observability, and operational excellence. Mentor engineers through code reviews, technical coaching, and day-to-day leadership. Engineering DeliveryOwn the delivery of technical initiatives from design through … design. Experience working with Redis or other NoSQL technologies. Deep understanding of microservices architecture, APIs, distributed systems, and cloud-native applications. Experience with monitoring, observability, incident response, and production operations. Strong debugging and performance optimisation skills. Demonstrated experience shipping reliable production systems that process financial transactions. LeadershipExperience leading engineering teams ...

Site Reliability Engineer (Edv) - National Security

Location
Cheltenham, England, United Kingdom
keep broadening their technical remit. What you'll be working with AWS Kubernetes Terraform Linux CI/CD Python/Bash Monitoring and observability Automation Reliability and performance engineering You’ll be working across secure, live National Security environments, helping teams improve deployment, resilience, monitoring and operational performance. ...

DevOps and Infrastructure Engineer

Hiring Organisation
Sanderson Government and Defence
Location
Gloucestershire, South West, United Kingdom
Employment Type
Permanent
cloud-based solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver ...

Senior Full Stack Engineer (Realtime & Voice) Customer Experience Platform

Location
Greater London, England, United Kingdom
Build the safety and compliance plumbing enterprise partners audit, including guardrails, content filtering, and PII redaction integration points Keep revenue-critical deployments healthy through observability, alerting, incident response, and SLA performance Build the platform capabilities forward-deployed engineers configure for partner telephony integrations and go-lives Raise the engineering … standard part of their workflow, with the judgment to review, correct, and own everything that ships An operable-systems mindset, covering SLAs, observability, on-call rotations, and rollback plans The ability to break down complex problems, make pragmatic tradeoffs, and ship iteratively, backed by strong communication across product, design ...

MongoDB Site Reliability Engineer

Location
Knutsford, England, United Kingdom
paced environment, your role will be essential to ensuring our infrastructure remains resilient, secure, and scalable. You’ll work on automating operations, enhancing system observability, and driving continuous improvements that reduce downtime and improve efficiency. If you’re motivated by solving, multi-layered problems and building systems that perform reliably ...

Senior Applied AI Engineer (Defence Contractor)

Location
United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Senior Data Architect

Hiring Organisation
Radley James
Location
City of London, Greater London, UK
stakeholder management skills. Desirable Experience with crypto or digital asset data. Familiarity with regulatory and audit reporting. Exposure to analytics layers, BI tools, or observability platforms. ...

Senior Applied AI Engineer (Defence Contractor)

Location
Greater London, England, United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Senior Back-End Developer - .NET / C# -High Growth SaaS

Hiring Organisation
Robert Half
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
integrations and data models Develop secure, scalable payment and FX services Work with PostgreSQL, Redis and Kubernetes/GKE Improve CI/CD, testing, observability and reliability Lead code reviews and mentor other developers Work closely with DevOps, Product and Operations Requirements Strong commercial C#/.NET experience Microservices ...

Senior Full Stack Engineer (Realtime & Voice) Customer Experience Platform

Location
United Kingdom
Build the safety and compliance plumbing enterprise partners audit, including guardrails, content filtering, and PII redaction integration points-Keep revenue-critical deployments healthy through observability, alerting, incident response, and SLA performance-Build the platform capabilities forward-deployed engineers configure for partner telephony integrations and go-lives-Raise the engineering … standard part of their workflow, with the judgment to review, correct, and own everything that ships-An operable-systems mindset, covering SLAs, observability, on-call rotations, and rollback plans-The ability to break down complex problems, make pragmatic tradeoffs, and ship iteratively, backed by strong communication across product, design ...

Quality Engineer

Hiring Organisation
E-Solutions IT Services UK Ltd
Location
Burgess Hill, West Sussex, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
practices, particularly GitHub Actions. · Test-data creation and reusable test-data approaches. · Automation reliability, including identifying and resolving flaky or inefficient tests. · Logs, observability and other engineering tools used for defect investigation and root-cause analysis. AI & Modern Quality Engineering Experience or a strong interest in using AI-assisted ...

python developer for an investment platform

Location
Greater London, England, United Kingdom
system integration Knowledge of automated testing and quality engineering Knowledge of CI/CD and version control Knowledge of application security, monitoring, and observability Experience with Snowflake or comparable modern data platforms Experience working with data‐intensive applications or analytics solutions Strong stakeholder communication and requirements analysis skills Experience mentoring ...

Platform Technical Lead (OpenShift)

Location
Welwyn Garden City, England, United Kingdom
platform across its full stack: OpenShift (HCP/Virtualization), GitOps delivery (Argo CD), multi‐cluster management and policy (ACM, Gatekeeper/OPA), observability, networking and storage integration, and infrastructure‐as‐code. Set and enforce the paved‐road patterns the team builds to. Evaluate new technologies against product outcomes, not novelty … first platform design: authoring and versioning APIs consumed by other product teams, managing breaking changes and deprecation, Kubernetes API extension patterns (CRDs, operators). Observability engineering: Prometheus, Alertmanager, Grafana or equivalent; defining SLIs/SLOs and designing alerting that reflects service health rather than component noise. Designing and operating fault ...

DevOps Platform Engineer

Hiring Organisation
Context Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £100,000 per annum, Inc benefits
sharded MongoDB clusters. Building and managing Kubernetes environments, including GKE and self-managed clusters. Designing and supporting highly available, resilient infrastructure. Monitoring and observability tools such as Grafana, Prometheus, ELK Stack and rsyslog. Infrastructure automation using Terraform, Ansible, Packer and Bash. Strong Linux administration experience, particularly Ubuntu environments. Exposure ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
Using BGP, EVPN-VXLAN, JunOS (Juniper QFX/MX), OSPF, Spine-Leaf/IP Fabric, VLANs, VRFs, Linux networking, Ansible, Terraform, Prometheus/Grafana, Observability tooling The adventures that await you after becoming Senior Network Engineer at Hack The Box: Design and implement spine-leaf network architectures across multiple data … series) running JunOS Develop and maintain network automation using Ansible, Terraform or similar infrastructure‐as‐code tooling Establish and improve network monitoring, alerting and observability (Prometheus, Grafana, SNMP, streaming telemetry) Plan and execute network capacity upgrades, site bring‐ups and hardware refresh cycles Collaborate with the platform/systems engineering ...

Endpoint Engineer

Location
Greater London, England, United Kingdom
platform engineering across Windows, Microsoft Intune, and related technologies. This role delivers secure, reliable, and frictionless user experiences through modern device management practices, automation, observability, and Zero Trust principles. How You’ll Make An Impact Endpoint Platform Engineering (Windows CSP/Intune) Own the migration of legacy Group Policy configurations … e.g., Patch My PC). Strong PowerShell automation expertise and familiarity with Git‐based workflows. Experience with vulnerability management tools and endpoint telemetry/observability platforms. Experience supporting local AI/LLM developer tooling (e.g., Claude Code, Ollama, LM Studio) and GPU‐accelerated workstations. Microsoft certifications related to Endpoint, Azure ...

Software engineering specialist

Hiring Organisation
Randstad Digital
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£650 - £700 per day
container deployments, and tool setups directly into developer workflows. Tech Stack Delivery: Automate the provisioning and configuration of our operational ecosystem, including tools across observability ( Prometheus, Elastic, Checkmk ), security/compliance ( Tenable, Red Hat Satellite ), service registry/IPAM ( NetBox ), container orchestration ( ArgoCD ), and artifact management ( Artifactory, GitLab ). Engineering ...

Fullstack Engineer

Location
Greater London, England, United Kingdom
product strategy, engineering direction and future hiring. Develop cloud-native systems using AWS Serverless, Lambda, PostgreSQL, Supabase and modern API architectures. Improve platform reliability, observability, security, performance and scalability. Support CI/CD, Infrastructure as Code and software engineering best practices. Tech Stack TypeScript, React, Next.js, Node.js, Go, Python ...