15 of 15 Observability Jobs in East Lothian

Remote Staff Software Engineer - Data Platforms

Hiring Organisation
Our Future Health
Location
North Berwick, East Lothian, UK
review, TDD, CI/CD and pairing using tools like Git and GitHub. Experience in operationally managing software components/service once live, including: observability best practises, logging best practises, error reporting, debugging and live incident management. Experience using tools such as Grafana, Prometheus, New Relic etc. Experience of working ...

Remote Software Engineer (Cloud & Integration)

Hiring Organisation
Arden University
Location
Tranent, East Lothian, UK
code (IaC), automation, and CI/CD pipelines. Proficiency in writing and debugging code using TypeScript (Node.js) and Python. Experience with monitoring, logging, and observability tools such as Grafana, Prometheus, and AWS CloudWatch. Proficiency with Infrastructure as Code (IaC) tools, particularly AWS CDK and Terraform. Exposure to unit testing, integration ...

Remote Senior Software Engineer (Python)

Hiring Organisation
Aveni
Location
Tranent, East Lothian, UK
Background in agentic systems or autonomous workflows Experience in financial services or regulated environments Familiarity with Node.js/TypeScript, Terraform or CDK Knowledge of observability tools and ML fundamentals You’ll Thrive If You Are Pragmatic and able to balance speed with long-term quality Curious about AI, especially safety ...

Remote Technical Support Engineer - Remote EMEA

Hiring Organisation
The Factory
Location
North Berwick, East Lothian, UK
that complements Factory’s mission. Prior work at startups or in high‐growth environments where you built processes and tools from scratch. Familiarity with observability tools, automation frameworks or CI/CD pipelines. Basic coding or scripting abilities beyond your primary language (e.g., writing integrations or automations). ...

Remote Staff Data Engineer Subscriptions User Understanding

Hiring Organisation
Spotify
Location
Musselburgh, East Lothian, UK
teams to turn complex business challenges into durable, well-designed data solutions. Lead architectural decisions and establish engineering best practices for data quality, governance, observability, reliability, and operational excellence across multiple squads. Design data models and platform architecture that support sustainable growth toward one billion users while balancing performance, cost ...

Remote Senior Software Engineer

Hiring Organisation
Aveni
Location
Tranent, East Lothian, UK
with React Solid understanding of cloud-native engineering on AWS Experience with microservices, messaging patterns and distributed systems A commitment to clean code, testing, observability and operational excellence A proactive and motivated mindset — someone who wants to build, ship and iterate quickly Interest in AI-powered products and a drive ...

Remote (Senior or Staff) Backend Engineer, AI tooling

Hiring Organisation
grabjobs
Location
Tranent, East Lothian, UK
least 2-3 years working on senior/staff level Proven hands-on experience building and operating developer platforms, CI/CD systems, observability/metrics infrastructure, code quality automation, testing infrastructure and/or developer tooling. You have previously shipped infrastructure or tooling that measurably improved developer velocity ...

Remote Business Systems Software Engineer/Architect

Hiring Organisation
grabjobs
Location
Tranent, East Lothian, UK
Bachelor's degree in Computer Science, Software Engineering, or equivalent practical industry experience. Experience designing 'defensive' system architectures—building for error handling, auditability, and observability, specifically for scenarios where manual operational intervention is required. A collaborative mindset that values the speed of low-code platforms while maintaining the rigor ...

Remote Head of DevOps

Hiring Organisation
1inch
Location
Tranent, East Lothian, UK
secrets management and infrastructure provisioning. Service Reliability: Define and implement SLIs, SLOs, and SLAs to ensure the stability and performance of critical services. Observability & Incident Management: Oversee the full monitoring stack and establish formal incident response and post-mortem processes. Security & Compliance: Enforce "security by design" and maintain strict adherence … GitOps practices. Cloud Architecture: Proficiency in managing multi-cloud environments (AWS, Hetzner, GCP). Technical Depth: Strong understanding of service mesh and microservices architecture. Observability: Solid experience with observability tools for metrics, logging, and tracing. Security Mindset: Practical experience implementing security best practices within CI/CD. Communication: Ability ...

Remote Head of Engineering, POS Application Platform

Hiring Organisation
Moniepoint Inc
Location
Tranent, East Lothian, UK
that runs on our POS devices. Build and govern the four core platform pillars across the organisation: developer experience (tools, libraries, SDKs), platform reliability (observability, uptime, quality standards), foundational frameworks (architecture, coding standards, testing and release pipelines), and squad adoption (driving uptake of platform tooling across all embedded POS engineers … Payments, Onboarding, VAS, Savings, and Loans squads - setting the bar for engineering quality and providing the platform layer above all of them. Own POS observability end-to-end: monitoring, alerting, and incident response frameworks that ensure platform health across all devices and squads. Drive the architecture for how Moniepoint supports ...

Remote Senior Director, Engineering- X-Ops Platform

Hiring Organisation
grabjobs
Location
Gullane, East Lothian, UK
vision, strategy, and operating model for the X-Ops Platform organisation (e.g., Delivery Roadmap, AI First Developer Experience, CI/CD, Observability, SRE, Cloud & Infrastructure Enablement) Lead, coach, and develop engineering leaders (directors, managers and senior ICs), building high-performing teams with clear ownership and strong engineering culture Own platform … record of building reliable, secure platforms and improving developer productivity through pragmatic, outcome-driven investment Experience establishing operational excellence practices (incident response, on-call, observability, SLOs, post-incident learning) and driving continuous improvement Ability to set strategy and translate it into execution via roadmaps, prioritisation, and clear measures of success ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Haddington, East Lothian, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Software Engineer Infra Agent Systems UK

Hiring Organisation
Together Ai
Location
Dunbar, East Lothian, UK
infrastructure agents. Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions. Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs. Own services end to end, including architecture, implementation, testing … deployment, observability, and production operations. Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops. Turn what agents learn in production into reliable, reviewed software and automation. Requirements 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms. Strong systems design skills and experience ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
grabjobs
Location
Longniddry, East Lothian, UK
post-incident reviews in a "you build it, you run it" environment. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Partner with Product Management and Design to translate business requirements into scalable technical solutions. Minimum Qualifications Bachelor's degree in Computer Science … HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Kubernetes, Docker, and Helm. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Experience in enterprise SaaS organisations, particularly HR Tech or regulated ...

Remote DevOps Team Lead

Hiring Organisation
grabjobs
Location
Dunbar, East Lothian, UK
operation of Runware’s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...