726 to 750 of 911 Remote/Hybrid Observability Jobs

Staff Reliability Engineer: Backend & Mobile (Remote)

Location
Greater London, England, United Kingdom
will lead incident response, drive cross‐team improvements, and strengthen operational practices in a remote-first environment. You will own reliability outcomes, build observability, and mentor engineers while collaborating with product and design to balance delivery and risk. #J-18808-Ljbffr ...

Remote AI Platform Support Architect

Location
United Kingdom
success across enterprise AI workloads. You will collaborate with NVIDIA partners, OEMs, and PS teams, delivering reusable assets and clear escalation paths while shaping observability dashboards and best practices #J-18808-Ljbffr ...

Hybrid Data Architect: Multi-Cloud & AI Platforms

Location
West of England, England, United Kingdom
data and AI transformations for clients, defining data blueprints and end-to-end architectures. You will work with multi-cloud platforms, governance, security, and observability while translating business needs into scalable solutions. The role emphasizes strong client-facing skills, collaboration across engineering teams, and a commitment to continuous learning ...

Hybrid Enterprise Integration Product Manager

Location
Sheffield, England, United Kingdom
Sheffield is seeking an Enterprise Integration Product Manager for a hybrid, contractor role. You will own a multi-quarter roadmap across middleware, messaging, and observability, driving platform standards and governance, while aligning with regulatory and regional constraints. The role requires deep IBM MQ knowledge, experience with ACE/ ...

Principal Engineer I, Prepurchase Platform (Remote)

Location
City Of London, England, United Kingdom
demand on-sales. You will write production code daily, influence technical direction, and collaborate across multiple teams within the Prepurchase domain. You will drive observability, resilience patterns, and AI-assisted enhancements while embedding across services or working horizontally. #J-18808-Ljbffr ...

Privacy Lawyer

Hiring Organisation
Elevate Legal Talent
Location
United Kingdom
Start ASAP We are seeking an experienced Data Privacy Lawyer to join a leading global technology business focused on Search AI, cloud, security and observability , on an initial 6-month interim contract . This is an excellent opportunity for a strong Data Privacy/Data Protection Lawyer to join ...

API Engineer

Location
Greater London, England, United Kingdom
languages, we will give you plenty of room to learn and grow with your team. Maintain the reliability of our systems by building for observability and participating in incident response as needed. You can find more about how we work on our engineering page. Salary range for Senior Engineer ...

Observability Solutions Architect - AIOps

Hiring Organisation
Lorien
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
Observability Solutions Architect - AIOps We are currently recruiting for a Observability Architect with strong IT Operations background/AIOps experience to join one of our Insurance Clients on a 6 month contract. Please note this role is Inside IR35 and mostly remote with travel to London on adhoc basis This … data-science role. The right candidate designs systems that carry enterprise-scale telemetry. Experience Required: Strong enterprise IT operations experience Demonstrable experience architecting enterprise observability or monitoring estates - designing telemetry pipelines and platform deployments at scale, not only operating them. Deep, hands-on expertise with at least one major platform ...

Observability & AIOps Engineer

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
services. Together, we support users globally with data, information, and systems to achieve their business objectives. Aufgaben The impact you'll make As an Observability & AIOps Engineer, you will build the intelligence layer that enables enterprise systems and AI agents to understand operational health in real time. You will transform … autonomous operations by enabling AI agents to perceive, reason about, and respond to complex system behavior. What you'll do Design and implement enterprise observability strategies across cloud, infrastructure, applications, and platforms. Build telemetry pipelines that ingest logs, metrics, traces, events, and topology information. Deploy and tune observability platforms ...

Senior / Principal Applied AI Engineer (UK / Europe, Remote)

Location
United Kingdom
production-grade AI agents, agentic workflows, and LLM integrations for internal automation and in-product features for a web-native trading platform. Own implementation, observability, and security while partnering with Product and Platform teams to deliver measurable AI systems in production. Job Description Role Senior/Principal Applied AI Engineer … integrate LLMs into internal systems and the trading product. This is a hands-on engineering role: you will write production code, put evaluation and observability on everything shipped, and operate autonomously in a lean team. Key Responsibilities Build agent loops, tool-calling, structured outputs, planning/state management, retries, guardrails ...

Senior AI Product Engineer

Hiring Organisation
Elliptic
Location
London, UK
Employment Type
Full-time
more junior engineers through pair programming, code review, and design feedback. Raise the engineering bar across the team by promoting good practices in testing, observability, and AI system reliability. Influence cross-team decisions on how AI capabilities integrate with the rest of the Elliptic platform. What you will achieve … technical direction of an AI workstream, including architecture, evaluation, and rollout. Established or improved at least one team practice for building AI systems (evals, observability patterns, prompt management, rollout safety).Mentored junior engineers on AI engineering practices and contributed to their growth. Built strong working relationships across product, web engineering ...

Platform Engineer

Location
United Kingdom
Terraform, CloudFormation, or CDK), CI/CD pipelines, API Gateway, Lambda, and Aurora PostgreSQL — and confident applying fundamentals such as high availability, fault tolerance, observability, and cost control in a live environment. Payments or fintech exposure is an advantage but not required; what matters more is a practical, first‐principles … rollback processes Support deployment of AI generated applications and tooling Support change control and release management alongside Engineering and the Information Security Officer Reliability, Observability & Data Infrastructure Design and maintain systems for high availability, fault tolerance, and resilience by default Implement logging, monitoring, and alerting (for example CloudWatch) across services ...

Senior DevOps Engineer

Hiring Organisation
Experis
Location
Derby, Derbyshire, United Kingdom
Employment Type
Contract
/CD pipelines and deployment orchestration. Support Kubernetes and OpenShift platform troubleshooting and optimisation. Deliver secure and compliant infrastructure solutions. Implement monitoring, logging, and observability tooling across environments. Collaborate with engineering, architecture, and delivery teams to improve deployment efficiency and platform reliability. Champion automation-first approaches to infrastructure and application … designing and maintaining enterprise-scale CI/CD pipelines . Strong understanding of cloud security and secure delivery practices. Experience implementing monitoring, logging, and observability solutions. Ability to define technical standards, governance, and reusable deployment frameworks. Experience working within large-scale enterprise transformation programmes. Desirable Skills Experience within highly regulated ...

SRE | Permanent | London, Hybrid, AWS

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
scalability. Key responsibilities Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions. Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively. Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling, and improve infrastructure efficiency. Improve … Kubernetes operational experience (on-prem and AWS EKS).Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows. Deep understanding of observability and telemetry (monitoring, logging, tracing).Infrastructure as Code with Terraform; experience with GitOps workflows and CI/CD.Scripting proficiency in Python, Bash, or Go. Proven ...

DevOps Engineer

Location
Greater London, England, United Kingdom
build and operate the infrastructure that every engineering team at Invisible depends on — Kubernetes, CI/CD, identity, secrets management, networking, and observability — along with the internal tooling that allows engineers and AI coding agents to ship safely and quickly. The Platform team is small relative to the surface area … entire engineering organisation. We are looking for engineers who don’t just apply best practices in IaC, Kubernetes, CI/CD and observability, but understand the problems those practices were designed to solve. Most were built around the pace and failure modes of human engineers. Those assumptions are changing ...

Platform Engineer London, UK · Full time · Hybrid

Location
Greater London, England, United Kingdom
engineers. The role combines infrastructure and software development, with an approximate 65/35 split. You’ll work on cloud infrastructure, deployment automation, observability, CI/CD, and internal tooling. The goal is to make development and releases more reliable, efficient, and straightforward. What you’ll do Improve developer experience … Find the causes of slow builds, failed pipelines, and flaky tests. Develop internal tools, improve test infrastructure, and occasionally work on product features. Improve observability across our infrastructure and delivery workflows using OpenTelemetry, ELK. Leverage AI to improve engineering workflows by building cloud-based agent tools and helping teams adopt ...

Developer - Scala

Hiring Organisation
Hackajob Ltd
Location
Newcastle Upon Tyne, Tyne and Wear, North East, United Kingdom
Employment Type
Permanent, Work From Home
frontend engineers, QA, product owners, solution designers, and other backend developers to deliver high-quality product increments. Support production stability by investigating issues, improving observability, and continuously reducing technical debt. Required Skills and Experience: Professional backend development experience, ideally in enterprise, SaaS, or cloud-based product environments. Strong hands … problems. Nice to Have: Experience with Squeryl, Doobie, or similar Scala data access libraries. Familiarity with Grafana or Kibana dashboards, alerting, logging, and production observability practices. Experience with large distributed systems, horizontal scaling, resilient service design, or high-throughput Play/Pekko applications. Previous experience in Payroll ...

Senior Backend Engineer (London, Barcelona, Madrid)

Location
Greater London, England, United Kingdom
product and stakeholders to stay aligned on direction. Contribute to design reviews with a clear view on the trade-offs. Keep your services healthy - observability, feedback loops, incident response. Raise the bar around you through code review, pairing and knowledge sharing. ABOUT YOU: 4+ years as a software engineer, with … equivalent is fine. A clear communicator and a pragmatic problem-solver. NICE TO HAVE EXPERIENCE: Data-intensive applications at scale. Building and interpreting observability - metrics, logging, tracing. Working within or alongside DevOps, SRE or infrastructure teams. Complex distributed environments - high throughput, low latency, large datasets. Any exposure to commodities, energy ...

Agentic Platform Engineer · Manchester, UK ·

Location
Manchester, England, United Kingdom
Protocol (MCP). Build production-grade agent services using Python, cloud-native architectures, event-driven design, automation and Infrastructure as Code. Implement robust evaluation, observability and continuous improvement capabilities, including testing, tracing, telemetry and performance optimisation. Embed security, governance and responsible AI principles through least-privilege access, policy enforcement, auditability … workflow state, retrieval-augmented generation and human-in-the-loop patterns. Experience building secure integrations with enterprise APIs, repositories, cloud services, CI/CD, observability or ITSM platforms. Experience creating evaluation frameworks for agent quality, task completion, safety, reliability, latency and cost. Strong understanding of agent security, including workload identity ...

Software Engineering Manager

Location
Greater London, England, United Kingdom
appropriate technical documentation. Champion strong engineering practices, including collaborative programming, testing approaches such as TDD, CI/CD and production ownership. Ensure strong observability and operational health, with useful code‐quality and Production metrics, appropriate KPIs, well‐calibrated alarms and clear ownership of actions following incidents and PIRs. Build relationships … Data/Data Science, CRM, Security, Legal, Privacy and other engineering teams. Experience establishing strong operational ownership of production software, including CI/CD, observability, support practices and continuous improvement. Commitment to inclusive leadership, frequent feedback and creating an environment where engineers can grow, challenge ideas and take meaningful ownership. ...

Senior Engineer

Location
Greater London, England, United Kingdom
turn ideas into working product quickly Contribute to product decisions and pragmatic technical trade-offs Diagnose and fix bugs quickly Improve testing, monitoring, and observability Maintain data integrity, system stability, and security Work closely with technical leadership Collaborate with our senior technical advisor on architecture and technical direction Implement technical … processes Design and maintain data-compliant systems with privacy, security, and regulatory standards embedded by default (e.g. UK GDPR, NHS DSPT, MHRA guidance) Implement observability and metrics to understand product performance, user behaviour, and system health Ensure our systems and features are audit-ready and aligned with healthcare regulatory requirements ...

Senior Software Engineer

Location
City Of London, England, United Kingdom
Make the architectural calls on your domain, write them down, and defend them in front of the tribe. Build in security, availability, reliability and observability from the start, and instrument your services so the answer to what's happening is already in a dashboard. Use AI tooling well. … awkward parts that come with them: idempotency, ordering, retries, poison messages, exactly-once as a promise nobody can keep. Security, availability, reliability and observability built in from the start, not bolted on at the end. You think about who can reach what, you know how your service behaves when ...

Software Engineering Manager

Location
Bristol, England, United Kingdom
risks early and transparently. Establish engineering guardrails across scope, quality, and non-functional requirements, enabling teams to design optimal solutions within them. Champion observability and operational excellence, ensuring system health, SLOs, and alerting are visible and actively managed. Partner with Tech Leads and Architects on system design and evolution, bringing … architecture, including API design and integration, performance optimisation, security, and microservice or event‐driven patterns. Experience with CI/CD, modern development workflows, and observability practices. Proven track record of leading high‐performing teams in fast‐paced, complex or regulated environments. Passion for mentoring and developing engineers through coaching, feedback ...

Sr. Software Engineer, Cloud Platform (Hybrid, London)

Location
Greater London, England, United Kingdom
culture emphasizes innovation, knowledge sharing, and technical excellence while maintaining a strong focus on security best practices. Our Tech Stack: Cloud Platforms: AWS, GCP Observability: LogScale (Humio), Grafana Infrastructure: Kubernetes, Kafka Programming: Golang (APIs), Python Data: PostgreSQL, Cassandra, OpenSearch Integration: REST, GraphQL, OAuth, JSON What You’ll Do: Design … processes, improve efficiency and drive business outcomes. Bonus Points: Experience working directly with customers or partners Security certifications Container orchestration experience Experience with observability tools Knowledge of CrowdStrike Falcon platform #LI-GO1 #LI-HYBRID Benefits of Working at CrowdStrike: Market leader in compensation and equity awards Comprehensive physical and mental ...

Senior Security Engineer, Detection and Response

Location
Greater London, England, United Kingdom
writing code, building AI-powered tooling, and automating workflows end-to-end. This role operates across the full detection lifecycle—from identifying gaps in observability to shipping high-signal detections and leading incident response when it matters most. You’ll help scale what a small team can accomplish by embedding … Principles Problem Solving to identify root causes and improve long-term resilience Partner cross-functionally with engineering and platform teams to expand logging, improve observability, and embed detection capabilities into the development lifecycle Continuously improve detection quality by analyzing alert performance, tuning for signal, and building feedback loops between incidents ...