301 to 325 of 761 Remote/Hybrid Observability Jobs

Product Engineer (all levels)

Location
Greater London, England, United Kingdom
help shape that), here’s a current snapshot: Full‐stack TypeScript, React, Postgres, and Temporal for long‐running orchestration. Infra: AWS, Terraform, strong observability via Sentry and Datadog (full‐stack, not just logs). Data: we lean heavily on Snowflake, Omni, dbt, Fivetran, Amplitude and Segment — and we’ve even ...

Site Reliability Engineer - Hybrid, Observability

Location
United Kingdom
bet365 Group is seeking a Site Reliability Engineer to strengthen the stability, performance and resilience of the systems behind every click and live change across our global product. In this full-time role, you will ...

Custody Support Applications Support - Assistant Vice President

Location
Belfast City District, Northern Ireland, United Kingdom
operational excellence of a suite of business-critical custody and settlement applications. The role focuses on distributed systems, cloud-native technologies, microservices, and modern observability platforms supporting securities processing and settlement functions. This position combines traditional application support responsibilities with Site Reliability Engineering (SRE) principles, automation, resiliency engineering, and operational … enterprise relational databases Database performance analysis and tuning Linux/Unix fundamentals Application troubleshooting and performance diagnostics Incident and Problem Management processes Monitoring and observability platforms Desirable: OpenShift/Kubernetes Cloud technologies (Google Cloud Platform, Azure, AWS) Microservices and distributed architectures Event-driven architectures and messaging platforms (Kafka, MQ) Elastic ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
Southampton, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£60000 - £70000/annum
Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using … with tools such as Jenkins, GitLab CI/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
opportunity to grow and make a difference in ways that matter to you. Role Summary In this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability and recoverability … Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application‐specific logging Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk ...

AI-Enabled Product Engineer Aladdin, Investments & Trading Engineering, Vice President

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
compliance and investment information. Build scalable, resilient, and operable systems that balance functional requirements with non-functional concerns such as performance, latency, reliability, security, observability, and maintainability. Influence architecture and technical direction across initiatives, articulating designtrade-offsand aligning stakeholders around pragmatic engineering choices. Champion engineering excellence through test-driven development … augmented generation, semantic search, embeddings, knowledge retrieval, or agentic workflow patterns. Strong focus on software quality, including automated testing, TDD, BDD, code review discipline, observability, resiliency, and production support. Ability to reasonaboutscalability, reliability, security, cost, maintainability, and delivery velocity when making technical decisions. Experience working in large, global engineeringorganisationsand collaborating ...

Lead Performance Test Engineer

Location
Greater London, England, United Kingdom
distributed systems.* Knowledge of cloud-native, containerised and modern application architectures.* Experience integrating performance testing within CI/CD pipelines.* Experience using monitoring, observability and telemetry tools to identify performance bottlenecks.* Strong root-cause analysis, diagnostics and performance tuning skills.* Experience defining performance requirements, workloads, test data, baselines and acceptance … Level Objectives (SLOs) and operational resilience.* Experience working within regulated, public sector or healthcare environments.* Relevant certifications in Performance Testing, Cloud, DevOps, SRE or Observability disciplines.## **Your security clearance**To be successfully appointed to this role, it is a requirement to obtain Security Check (SC) clearance. To obtain SC clearance ...

Senior Fullstack Engineer (Python + React.js)

Location
Greater London, England, United Kingdom
minimal downtime. Write unit and integration tests to maintain code reliability and ensure high- quality releases. Continuously monitor and optimize backend performance using observability tools such as Datadog, Cloud Watch or similar. Participate in design discussions and decision-making to enhance system robustness and scalability. Maintain technical documentation to ensure … handling asynchronous communication. Experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation for managing cloud infrastructure. Knowledge of observability and monitoring tools, such as Cloud Watch or Datadog, to track and troubleshoot system performance. Familiarity with serverless architectures (e.g., AWS Lambda) and event‐-driven programming paradigms. Exposure ...

Senior DevOps Engineer

Location
United Kingdom
pipelines for zero-touch deployments. Security hardening – lead security-posture reviews, implement GuardDuty, CloudWatch and IAM best practices. SRE & monitoring – uphold SLAs through observability stacks, proactive alerting and performance tuning of distributed systems. Collaboration & enablement – automate repetitive tasks, mentor developers and champion DevSecOps best practice across teams. Policy & audit ownership … Preferred/Bonus Experience with MLOps/LLMOps (Softwares such as Sagemaker, Kubeflow or ZenML) Deployment of on-premise Kubernetes Prometheus (or other stacks) observability Experience with AWS Karpenter & Compute Optimizer Compliance literacy - ISO 27001, NIST SSDF/OWASP SAMM, GDPR basics An active SC or DV clearance ...

Remote Senior AWS Infra & DevOps Engineer

Location
Birmingham, England, United Kingdom
infrastructure behind a real-time health-tech platform. You will design and operate high-availability systems, implement IaC, and own observability and incident response. The role emphasizes security, automation, and scalable delivery. You will collaborate with engineering to optimize performance, cost, and resilience, with a path toward expanding ...

Senior Java Developer – Remote Cloud‐Native Microservices

Location
Greater London, England, United Kingdom
building a real-time banking platform. You will design, develop and operate cloud-native services using Spring Boot and automated delivery, focusing on reliability, observability and scalable systems. You will work with distributed architectures, Kafka and RESTful APIs, contributing across the development lifecycle with a bias toward automation and continuous ...

Backend Platform Engineer – Java/AWS (Hybrid London)

Location
Greater London, England, United Kingdom
merchandise innovator, is hiring a Software Engineer for its Forge/Platform Team in Camden. You’ll join a small crew to enhance resilience, observability, and scalability of internal systems that power post-purchase fulfilment and vendor integrations. You will collaborate with senior engineers, own technical areas over time ...

Network & Infrastructure Tooling / Automation Specialist/ Architect - freelance - hybrid, London, UK

Location
Greater London, England, United Kingdom
designs. Build tooling to support large-scale migration activities with minimal operational risk. Integrate SD-WAN, MPLS, and cloud connectivity into unified operational and observability platforms. Enable traffic engineering, QoS, and policy enforcement through code-driven workflows. Observability, Operations & Resilience Develop automation for fault detection, root cause analysis, and remediation. … Integrate telemetry, logs, and metrics into observability platforms. Support SRE-style practices including error budgets, reliability metrics, and continuous feedback loops. Improve MTTR and operational consistency through standardised tooling and workflows. Governance, Compliance & Secure Design Embed security, segmentation, and compliance controls into automation pipelines. Ensure tooling supports audit trails, approval ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
TypeScript experience. Solid AWS and Serverless experience. Experience building distributed, resilient APIs. Experience with MongoDB or another document database. A strong understanding of testing, observability and performance. Experience working in financial services, fintech, startups or another regulated environment. Experience mentoring engineers and taking ownership of production services. ...

Data Engineer

Hiring Organisation
Sanderson Recruitment
Location
South West, United Kingdom
Employment Type
Contract
Contract Rate
£600 - £700 per day
ready data capabilities. Key Responsibilities Build and enhance data pipelines, models and data products Design scalable ingestion and transformation solutions Improve data quality, testing, observability and performance Support strategic platform and business transformation initiatives Skills & Experience Strong Data Engineering experience with advanced SQL and data modelling skills Hands-on experience ...

Data Engineering Lead

Hiring Organisation
Eutopia Solutions Limited
Location
Andover, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
warehouse solutions Drive best practice across SQL, Python, data modelling, and engineering standards Improve platform reliability, performance, and scalability Lead data quality, monitoring, and observability initiatives Work closely with senior stakeholders across the business to deliver data-driven outcomes Evaluate new tools, technologies, and future platform requirements Technology Microsoft Azure ...

Databricks Lead Engineer - Hybrid - Contract

Hiring Organisation
Anson Mccade
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £500 to £575 per day
Unity Catalog experience covering catalog design, access controls, lineage and secure data sharing Experience developing Declarative Pipelines with Expectations, including failure handling and observability Knowledge of Auto Loader, Lakeflow Connect and batch, streaming or incremental ingestion patterns Production-level PySpark and SQL development skills Strong knowledge of Delta Lake, including ...

Business Systems Engineer

Location
City of Edinburgh, Scotland, United Kingdom
Stack PHP/Laravel TypeScript, Node.js, NestJS React MySQL & PostgreSQL AWS (Lambda, ECS, RDS, S3, SQS, SNS) Docker & CI/CD pipelines Monitoring and observability tools What We're Looking For Strong experience in either PHP/Laravel or TypeScript/Node.js . Proven ownership of production software and business ...

DevOps Infrastructure Engineer

Hiring Organisation
Sterling Bridge Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
Doing Designing, managing and scaling production Kubernetes environments. Building fully automated infrastructure using Terraform, Ansible and modern DevOps tooling. Driving reliability, resilience and observability across the platform. Administering Linux infrastructure across physical and virtual environments. Working directly alongside talented engineers in a collaborative, fast-moving team. Owning projects from concept ...

Devops Infrastructure Engineer

Hiring Organisation
Sterling Bridge Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Ansible or similar Excellent Linux administration skills Experience supporting highly available production platforms Strong understanding of infrastructure automation, reliability and scalability Experience with monitoring, observability and infrastructure optimization Ability to take ownership of infrastructure projects from concept through to production UK based and able to commute to Southampton Desirable Skills ...

Lead Data Engineer

Hiring Organisation
Hexwired Recruitment Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
£80000 - £115000/annum
storage solutions for operational and AI workloads. Build and optimise vector search, retrieval systems and ML data pipelines. Ensure data reliability, governance, monitoring and observability across the platform. Work closely with AI, backend and product teams to support model training, inference and product development. Optimise large-scale datasets and database ...

Senior Engineer (Fincrime)

Location
Greater London, England, United Kingdom
with Python, but strong OOP fundamentals are what matter most. Adept at both constructing and managing services, with proficiency in establishing standard APIs, implementing observability (logging, metrics, alerting), and integrating external systems. Experience with event-driven architectures and Kafka. A quality-first mindset: your code should be testable and well ...

Salesforce Technical Architect - Service Cloud & Agentforce

Hiring Organisation
Focus on SAP
Location
West Yorkshire, England, United Kingdom
Govern technical designs, integrations, code quality, test automation, security and release readiness throughout the delivery lifecycle. Ensure architecture meets requirements across performance, resilience, security, observability, compliance, maintainability and Salesforce platform limits . Mentor technical leads and developers, resolving complex design challenges and supporting cutover, production deployment and post-release assurance ...

Machine Learning Engineer

Location
United Kingdom
ready code. Deploy models into batch and real-time environments, managing versioning, promotion, rollback, and scheduled workflows via MLflow and APIs. Implement monitoring and observability, including data and model drift detection, performance alerts, logging, and automated retraining. Collaborate with Data Engineering and Platform teams on CI/CD integration, pipeline ...

Senior Software Engineer

Location
Cambridge, England, United Kingdom
will be impactful and varied, including: Building new product features (full-stack, with frontend focus) Designing non-functional app features (e.g. offline support, improving observability) Contributing to technical roadmap planning Working with product/UX to iterate on the app design and requirements Participating in bug triage/diagnosis ...