426 to 450 of 1,405 Remote/Hybrid Observability Jobs

WAF Security Engineer

Location
United Kingdom
repeatable and scalable outcomes. Playing a key role in platform modernisation and service improvements, including refactoring, automation and the reduction of technical debt Driving observability and operational insight by improving monitoring, alerting and telemetry to strengthen platform reliability and incident response Support the diagnosis and resolution of major incidents, partnering ...

Senior Mobile Engineering Manager (iOS & Android SDK)

Location
Greater London, England, United Kingdom
Kotlin and Swift. Required: experience with Git-based development workflows and CI/CD release pipelines (We use Github Actions). Required: experience with observability and crash tooling used in mobile environments (for example Sentry, Firebase, Crashlytics). Preferred: Python for data analysis and validation workflows. Preferred: experience with feature ...

Senior QA Engineer Reading, United Kingdom

Location
Reading, England, United Kingdom
quality, ensuring testing and quality considerations are part of the conversation from day one. Partner with developers during design to build testability and observability into the product from the start, not bolt it on afterward. Build and maintain internal testing tools/frameworks that raise the bar for the whole ...

Backend Software Engineer - Infrastructure

Location
Greater London, England, United Kingdom
scheduling hundreds of thousands of containers every hour Designing architecture and opinionated APIs to keep application developers on the happy path Tracing and performance observability in high scale distributed microservice architectures Building reliant, performant, and scalable systems for storage, auth, or asset serving to enable other product teams to build ...

Platform Engineer

Location
Greater London, England, United Kingdom
tooling and improvements that would have real impact, then take ownership of building them once we agree they're worth it. Support reliability and observability, using Datadog to make sure we find out about problems before our customers do. Take part in the operational work: migrations, upgrades, incident response ...

Expert Data Scientist

Hiring Organisation
BP Energy
Location
Sunbury-On-Thames, London, United Kingdom
Employment Type
Work From Home
innovations from experimentation through to productised, maintainable solutions that deliver measurable value. Drive engineering excellence across ML systems, including CI/CD, testing, observability, reliability, and MLOps guidelines. Define technical standards, patterns, and protocols for ML engineering and applied ML science across teams. Lead complex, multi-team technical initiatives ...

Product Manager - Data

Location
Greater London, England, United Kingdom
connected vehicles, and compliance with automotive industry standards such as ISO 26262 and ISO 21434 Application Management - Open source solutions in the enterprise including Observability, IAM, App Stores and technologies such Grafana, GitOps, and Juju Charms We will route you to the most suitable team. Location: These roles are home ...

Lead Platform Engineer

Location
City of Westminster, England, United Kingdom
testing, documentation and code organisation. Responsible for evolving those standards over time, protecting core principles while adapting to new tools and approaches. Ensure appropriate observability, logging and error-handling patterns are in place across applications. Responsible for ensuring documentation exists where it adds long-term value, and remains accurate. Problem ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
APIs, and serverless services Infrastructure & tooling: AWS, Terraform, Docker, Kubernetes, Redis, CI/CD pipelines Practices: Automation-first, metrics-driven, incident write-ups and observability baked in Our 2026 Engineering Strategy We’re not just here to write code, we’re here to redefine how insurance works. Our 2026 engineering ...

Principal Data Engineer (we have office locations in Cambridge, Leeds and London)

Location
Greater London, England, United Kingdom
into platforms handling sensitive clinical and research data Strengthening engineering practices and platform resilience through DataOps, DevOps, CI/CD, infrastructure as code, automation, observability and security by design Evaluating emerging technologies and helping teams make thoughtful, appropriate and sustainable technology choices Mentoring experienced engineers, sharing knowledge and contributing ...

AI Data/Graph Engineer

Location
Greater London, England, United Kingdom
Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, OpenTelemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyperscaler or on-prem. A tool-for-tool match is not expected ...

Distinguished Engineer, Core DevOps

Hiring Organisation
GitLab
Location
United Kingdom
Salary
£ 70 K
defining evaluation frameworks, running benchmarks, and translating findings into scalable architecture decisions.Strong background in scalable, multi-tenant distributed systems, including service decomposition, fault tolerance, observability, and operational resilience.Experience designing and implementing human-in-the-loop controls, safety guardrails, and responsible AI practices for production systems.Experience mentoring senior engineers and influencing ...

Distinguished Engineer, Core DevOps

Hiring Organisation
GitLab
Location
Moffat, Dumfries & Galloway, UK
Employment Type
Full-time
evaluation frameworks, running benchmarks, and translating findings into scalable architecture decisions. Strong background in scalable, multi-tenant distributed systems, including service decomposition, fault tolerance, observability, and operational resilience. Experience designing and implementing human-in-the-loop controls, safety guardrails, and responsible AI practices for production systems. Experience mentoring senior engineers ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)—to Google Cloud Platform. Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs … Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development ...

Remote SRE: Platform Reliability & Observability

Location
United Kingdom
Orexnova is seeking an experienced SRE/Platform Engineer to join a fully remote UK team. You’ll own incident response, blameless post-mortems and drive reliability improvements across services while partnering with the SRE ...

Remote .NET Core Engineer - CI/CD, Azure & Observability

Location
Bishop's Stortford, England, United Kingdom
The Delta Group is seeking an experienced software developer to join our Microsoft-focused team. This remote role requires occasional travel to Head Office in Bishops Stortford or Melksham to collaborate on deployments and reviews. ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Location
United Kingdom
managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform …/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud‐native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem‐solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating ...

DevOps Engineer, Blockchain Infra (Fully Remote)

Hiring Organisation
Binance
Location
London, United Kingdom
Salary
£ 70 K
/CD pipelines for application and infrastructure deployment.Deploy and operate middleware platforms such as Kafka, Redis and NGINX.Automate operational tasks using Golang and Python.Improve observability through monitoring, logging, alerting, and distributed tracing.Ensure platform reliability, scalability, security, and disaster recovery.Troubleshoot production incidents and perform root cause analysis.Work closely with software engineers … Kafka/Redis/NGINXStrong scripting and programming skills in: Golang/PythonExperience with Git, GitOps workflows, and CI/CD platforms.Strong understanding of observability tools such as Prometheus, OpenTelemetry.Familiarity with container technologies including Docker and Kubernetes. Strong troubleshooting and problem-solving skills.Preferred QualificationsExperience operating blockchain infrastructure or Web3 platforms.Experience ...

DevOps Engineer, Blockchain Infra (Fully Remote)

Hiring Organisation
Binance
Location
London, UK
Employment Type
Full-time
pipelines for application and infrastructure deployment. Deploy and operate middleware platforms such as Kafka, Redis and NGINX.Automate operational tasks using Golang and Python. Improve observability through monitoring, logging, alerting, and distributed tracing. Ensure platform reliability, scalability, security, and disaster recovery. Troubleshoot production incidents and perform root cause analysis. Work closely …/Redis/NGINXStrong scripting and programming skills in: Golang/PythonExperience with Git, GitOps workflows, and CI/CD platforms. Strong understanding of observability tools such as Prometheus, OpenTelemetry. Familiarity with container technologies including Docker and Kubernetes. Strong troubleshooting and problem-solving skills. Preferred QualificationsExperience operating blockchain infrastructure ...

Senior DevOps Engineer

Location
United Kingdom
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...

Senior DevOps Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Permanent
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...

AI Platform Engineer- Senior Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
Implement MLOps and LLMOps pipelines (model deployment, monitoring, retraining, and fine-tuning where relevant) using Infrastructure-as-Code, GitOps, and CI/CD• Establish observability, security, and governance frameworks specific to AI systems, including cost attribution and lifecycle management• Work with clients and internal teams to develop new opportunities … equivalent), including familiarity with the Model Context Protocol (MCP)• Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)• Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust)MLOps & LLMOps• Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone ...

Data Technical Lead

Hiring Organisation
PA Consulting
Location
London, UK
Employment Type
Full-time
Data governance: Catalogue, metadata, lineage, quality, security, privacy and access controls Data products: Reusable, discoverable and well-governed data products and marketplaces DataOps and observability: Testing, monitoring, operational controls and platform reliability Platform engineering: Terraform, CloudFormation, Azure Bicep and infrastructure-as-code CI/CD: GitHub Actions, Azure DevOps, Jenkins … serving. Applying strong software engineering practices to data platforms, including testing, CI/CD and infrastructure-as-code. Establishing effective approaches to data quality, observability, governance, metadata and security. You can lead data platform modernisation and migration, including coexistence, cutover and decommissioning. Making pragmatic technology choices and understanding the trade ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic … technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit ...