926 to 950 of 2,476 Remote Observability Jobs

Platform Engineer

Location
Greater London, England, United Kingdom
tooling and improvements that would have real impact, then take ownership of building them once we agree they're worth it. Support reliability and observability, using Datadog to make sure we find out about problems before our customers do. Take part in the operational work: migrations, upgrades, incident response ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Livingston, Scotland, United Kingdom
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Dunfermline, Scotland, United Kingdom
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Broughton, Wales, United Kingdom
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Product Manager - Data

Location
Greater London, England, United Kingdom
connected vehicles, and compliance with automotive industry standards such as ISO 26262 and ISO 21434 Application Management - Open source solutions in the enterprise including Observability, IAM, App Stores and technologies such Grafana, GitOps, and Juju Charms We will route you to the most suitable team. Location: These roles are home ...

Backend Software Engineer - Infrastructure

Hiring Organisation
Palantir Technologies
Location
London, UK
Employment Type
Full-time
efficiently scheduling hundreds of thousands of containers every hourDesigning architecture and opinionated APIs to keep application developers on the happy pathTracing and performance observability in high scale distributed microservice architecturesBuilding reliant, performant, and scalable systems for storage, auth, or asset serving to enable other product teams to build robust applications ...

Lead Platform Engineer

Location
City of Westminster, England, United Kingdom
testing, documentation and code organisation. Responsible for evolving those standards over time, protecting core principles while adapting to new tools and approaches. Ensure appropriate observability, logging and error-handling patterns are in place across applications. Responsible for ensuring documentation exists where it adds long-term value, and remains accurate. Problem ...

Staff Software Engineer (Provenance)

Hiring Organisation
Cloudsmith
Location
Belfast, UK
Employment Type
Full-time
wider engineering team to understand how enterprise security teams consume provenance data and translate that into features they love. Quality: Prioritise correctness, security, and observability — this is critical infrastructure customers trust with their software supply chain decisions. Mentor: Share your expertise across the team through code reviews, documentation, and open ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)—to Google Cloud Platform. Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs … Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development ...

Senior Site Reliability Engineer

Location
Knutsford, England, United Kingdom
drive reliability, scalability and performance across critical banking systems. This role combines hands‐on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities Build and maintain reliable, scalable and secure infrastructure platforms and solutions. Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. Develop automation using programming and scripting to reduce manual intervention and improve efficiency. Develop and improve observability, monitoring, instrumentation and performance capabilities. Use data and reliability metrics to drive continuous improvement and optimisation. Lead technical discussions, blameless retrospectives and problem‐solving activities. Work ...

Senior Site Reliability Engineer

Hiring Organisation
GCS
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£75000 - £95000/annum Bonus
drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem-solving activities. * Work ...

Remote SRE: Platform Reliability & Observability

Location
United Kingdom
Orexnova is seeking an experienced SRE/Platform Engineer to join a fully remote UK team. You’ll own incident response, blameless post-mortems and drive reliability improvements across services while partnering with the SRE ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Location
United Kingdom
managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform …/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud‐native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem‐solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating ...

Senior DevOps Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Permanent
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...

Data Technical Lead

Hiring Organisation
PA Consulting
Location
Pimlico, Hertfordshire, UK
Employment Type
Full-time
Data governance: Catalogue, metadata, lineage, quality, security, privacy and access controls Data products: Reusable, discoverable and well-governed data products and marketplaces DataOps and observability: Testing, monitoring, operational controls and platform reliability Platform engineering: Terraform, CloudFormation, Azure Bicep and infrastructure-as-code CI/CD: GitHub Actions, Azure DevOps, Jenkins … serving. Applying strong software engineering practices to data platforms, including testing, CI/CD and infrastructure-as-code. Establishing effective approaches to data quality, observability, governance, metadata and security. You can lead data platform modernisation and migration, including coexistence, cutover and decommissioning. Making pragmatic technology choices and understanding the trade ...

Observability SME/Architect/Consultant

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
national client is looking for various observability candidates for upcoming projects. Observability Architect/SME/Consultant6 months+ contracts Inside IR35 Hybrid working Role PurposeThe Observability Architect/SME/Consultant will play a key role within the client's Operational Resilience services, helping customers establish, enhance, and optimise observability … hands-on subject matter expertise, requiring the ability to design strategic solutions while supporting implementation and continuous improvement initiatives. Essential Skills & ExperienceObservability & Monitoring Enterprise Observability Architecture Monitoring Strategy & Design Telemetry, Metrics, Logs, and Distributed Tracing Service Health Monitoring Dashboard Design & Reporting Alerting & Event Correlation AIOps and Operational Analytics Performance & Availability ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic … technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit ...

Remote Platform Engineer: AI, Cloud & CI/CD Automation

Location
Greater London, England, United Kingdom
design and operate pipelines, IaC, and cloud services across AWS/Azure and Kubernetes, aligning with GitOps and security best practices. You will implement observability, self‐service tooling, and automation while collaborating with developers to ship reliable software and AI workloads at scale. #J-18808-Ljbffr ...

Senior Full Stack Engineer C# Vue.js

Hiring Organisation
Client Server
Location
Bracknell, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
architecture and technical decisions within your domain, collaborate closely with Product, Design and Compliance to take features from idea to production and improve reliability, observability, testing and production performance, modernising legacy services and contributing to wider architectural direction. Location/WFH: You can work from home most of the time ...

Backend Software Engineer Python LLM - Finance

Hiring Organisation
Client Server
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
development You have hands-on experience deploying LLMs and multi-modal models at scale in production You have a strong understanding of scalable MLOps, observability and cloud-native AI deployment You're collaborative and pragmatic with advanced stakeholder communication, problem-solving and project management skills in Agile environments Experience with ...

DevOps Engineer

Hiring Organisation
Signature Recruitment Limited
Location
Bristol, Avon, United Kingdom
Employment Type
Full-Time
Salary
£50,000 - £60,000 per annum
extend infrastructure using Infrastructure as Code, Docker and GitOps approaches Manage edge devices and on-site collection systems Build and maintain monitoring and observability tools Support integration between hardware, control systems and software platforms Create and maintain technical documentation and operational runbooks What We're Looking For Strong experience ...

Senior Data Engineer (Databricks)

Hiring Organisation
Tenth Revolution Group
Location
Surrey, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £90,000 per annum
implementing modern Lakehouse architectures using Delta Lake and medallion design principles. Optimising Databricks workloads for performance, reliability and scalability. Implementing data quality, monitoring, and observability frameworks. Supporting governance, security, and access management within the data platform. Collaborating with Data Scientists, Analysts and Product teams to deliver business-critical data products. ...

Technical Engineer, Full Stack Java

Hiring Organisation
Clarify Consultancy Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
Kafka. You will demonstrate the ability to build robust, scalable and highly available solutions, with hands-on experience across Docker, Kubernetes, MongoDB and modern observability tooling. A solid grounding in design patterns, Domain-Driven Design and front-end developmentideally Reactis important, along with the ability to produce production-quality code ...