951 to 975 of 3,664 Remote/Hybrid Observability Jobs

Remote Senior Data Engineer

Location
North Berwick, East Lothian, United Kingdom
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Remote Senior Data Engineer

Location
Sutton, South West London, United Kingdom
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Sr. Snowflake Architect

Location
City of Edinburgh, Scotland, United Kingdom
level policies, Tri-Secret/external key management, network policies, Private Link, SSO/OAuth/MFA, audit trails, and access history.Embed data quality, observability, and lineage using Snowflake Horizon Catalog and native observability features; define SLAs/SLOs, reconciliation frameworks, and incident/RCAs for data reliability.Architect cross-region ...

Staff Machine Learning Engineer, Platform (MLOPS)

Hiring Organisation
Credit Acceptance Corporation
Location
United States
Employment Type
Permanent
Salary
USD 226,042 Annual
drift detection against pinned baselines, sampling and judge pipelines, and the release-gate mechanics that stop a regression from shipping. Build and maintain the observability and evaluation substrate other teams depend on: trace and telemetry capture including multi-turn and multi-step agent traces, logging standards, evaluation pipeline plumbing … because it is mandated. Partner with Cloud Engineering, Data Engineering, Security and SRE so the ML platform sits inside enterprise governance, identity and observability rather than beside it. Respond to AI-specific production incidents and drive them to a closed corrective action: prompt injection attempts, rogue-agent cost spikes, data ...

Staff Software Engineer, Compute Platform

Hiring Organisation
GitHub
Location
United Kingdom
Salary
£ 70 K
define safe migration paths, operational standards, and long-term platform direction.Debug and resolve complex production issues across Kubernetes, cloud infrastructure, Linux, containers, networking, observability, and distributed systems.Build automation and guardrails that improve workload quality, capacity efficiency, cluster lifecycle operations, and incident response.Mentor engineers and raise the engineering bar through design … service style systems.Experience with Kubernetes controllers, operators, CRDs, admission control, scheduling, autoscaling, service mesh, or multi-cluster networking.Experience improving reliability through SLOs, observability, incident response, automation, and operational readiness practices.Experience leading Staff-level technical projects with high ambiguity and broad organizational impact. GitHub valuesCustomer-obsessedShip to learnGrowth mindsetOwn the outcomeBetter ...

Agentic Platform Engineer · Manchester, UK ·

Location
Manchester, England, United Kingdom
Protocol (MCP). Build production-grade agent services using Python, cloud-native architectures, event-driven design, automation and Infrastructure as Code. Implement robust evaluation, observability and continuous improvement capabilities, including testing, tracing, telemetry and performance optimisation. Embed security, governance and responsible AI principles through least-privilege access, policy enforcement, auditability … workflow state, retrieval-augmented generation and human-in-the-loop patterns. Experience building secure integrations with enterprise APIs, repositories, cloud services, CI/CD, observability or ITSM platforms. Experience creating evaluation frameworks for agent quality, task completion, safety, reliability, latency and cost. Strong understanding of agent security, including workload identity ...

Software Engineering Manager

Location
Greater London, England, United Kingdom
appropriate technical documentation. Champion strong engineering practices, including collaborative programming, testing approaches such as TDD, CI/CD and production ownership. Ensure strong observability and operational health, with useful code‐quality and Production metrics, appropriate KPIs, well‐calibrated alarms and clear ownership of actions following incidents and PIRs. Build relationships … Data/Data Science, CRM, Security, Legal, Privacy and other engineering teams. Experience establishing strong operational ownership of production software, including CI/CD, observability, support practices and continuous improvement. Commitment to inclusive leadership, frequent feedback and creating an environment where engineers can grow, challenge ideas and take meaningful ownership. ...

Senior Ai Engineer

Hiring Organisation
Typeform
Location
United Kingdom
Salary
£ 70 K
information in more conversational and personalised ways.The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance.You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … building, evaluating, and releasing AI systems.Help teams make informed decisions about models, frameworks, infrastructure, performance, and cost.Apply strong engineering practices across testing, security, observability, version control, and deployment.Share technical knowledge and support the development of other engineers.Keep up with relevant AI research, tools, and engineering practices, applying what is useful ...

Senior Ai Engineer

Hiring Organisation
Typeform
Location
United Kingdom, UK
Employment Type
Full-time
more conversational and personalised ways. The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance. You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … evaluating, and releasing AI systems. Help teams make informed decisions about models, frameworks, infrastructure, performance, and cost. Apply strong engineering practices across testing, security, observability, version control, and deployment. Share technical knowledge and support the development of other engineers. Keep up with relevant AI research, tools, and engineering practices, applying ...

Software Development Engineer in Test

Location
Greater London, England, United Kingdom
quality gates, and make production failures easier to identify and diagnose. You will work hands-on across C#, TypeScript, APIs, data, pipelines, and observability rather than focusing solely on UI testing. Your work will help Athos reduce manual checking, catch failures earlier, and create stronger evidence behind every release. … quality gates into Azure DevOps CI/CD pipelines. Investigate failures across application logs, APIs, databases, telemetry, test environments, and production systems. Improve observability and production quality through monitoring, alerting, and better visibility into feed and syndication failures. Contribute to test strategy, including determining what should be tested, at which ...

Senior Engineer

Location
Greater London, England, United Kingdom
turn ideas into working product quickly Contribute to product decisions and pragmatic technical trade-offs Diagnose and fix bugs quickly Improve testing, monitoring, and observability Maintain data integrity, system stability, and security Work closely with technical leadership Collaborate with our senior technical advisor on architecture and technical direction Implement technical … processes Design and maintain data-compliant systems with privacy, security, and regulatory standards embedded by default (e.g. UK GDPR, NHS DSPT, MHRA guidance) Implement observability and metrics to understand product performance, user behaviour, and system health Ensure our systems and features are audit-ready and aligned with healthcare regulatory requirements ...

Software Engineering Manager

Location
Bristol, England, United Kingdom
risks early and transparently. Establish engineering guardrails across scope, quality, and non-functional requirements, enabling teams to design optimal solutions within them. Champion observability and operational excellence, ensuring system health, SLOs, and alerting are visible and actively managed. Partner with Tech Leads and Architects on system design and evolution, bringing … architecture, including API design and integration, performance optimisation, security, and microservice or event‐driven patterns. Experience with CI/CD, modern development workflows, and observability practices. Proven track record of leading high‐performing teams in fast‐paced, complex or regulated environments. Passion for mentoring and developing engineers through coaching, feedback ...

Platform Engineer

Location
City of Edinburgh, Scotland, United Kingdom
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort.We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration* Cloud platforms, virtualisation and containers* Infrastructure as code and configuration management* Programming and scripting, such as Python, Bash or Go* Monitoring, observability and SRE practices* Infrastructure, networking and performance troubleshooting* Secure, resilient and scalable system design* Technical documentation and operational guidance* Technical leadership and mentoringBeneficial skills include:* Strong ...

Engineering Manager (Remote - UK)

Hiring Organisation
Reonomy
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
building and evolving AWS-native and cloud-based platforms, alongside legacy systemsSet and uphold strong engineering standards across code quality, testing, CI/CD, observability and documentationStay close to technical decisions through design reviews, architecture discussions and hands-on coachingBalance new feature delivery with technical debt, reliability, security and long … distributed systemsProficiency in at least one modern programming languageStrong grasp of system design and software engineering fundamentalsExperience with Infrastructure as Code, CI/CD, observability and secure production systemsAble to communicate technical ideas clearly to both technical and non-technical stakeholdersUnlock your Altus Experience!If you’re looking to advance ...

Senior ML Ops Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
services. Build and maintain orchestration pipelines using tools such as Dagster, Airflow or Prefect. Deploy, manage and optimise workloads in Kubernetes environments. Improve monitoring, observability and platform reliability. Implement infrastructure-as-code and CI/CD best practices. Partner with software engineers, ML engineers and data scientists to deliver scalable ...

Senior Engineering Manager - 9-10 month FTC

Location
Manchester, England, United Kingdom
practices, including test‐driven development and automated testing, helping teams build quality into the development process. Support strong operational ownership through CI/CD, observability, production support and the principle that teams own the systems they build and run. Help teams identify and address technical debt sustainably while maintaining appropriate … outcomes and translating these into clear engineering priorities. Strong understanding of modern software engineering practices, including automated testing and TDD principles, CI/CD, observability and production ownership. Experience facilitating technical decisions, bringing the right people and evidence together and constructively challenging thinking when needed. Comfortable balancing feature delivery ...

Senior Engineering Manager - 9-10 month FTC

Location
Greater London, England, United Kingdom
practices, including test‐driven development and automated testing, helping teams build quality into the development process. Support strong operational ownership through CI/CD, observability, production support and the principle that teams own the systems they build and run. Help teams identify and address technical debt sustainably while maintaining appropriate … outcomes and translating these into clear engineering priorities. Strong understanding of modern software engineering practices, including automated testing and TDD principles, CI/CD, observability and production ownership. Experience facilitating technical decisions, bringing the right people and evidence together and constructively challenging thinking when needed. Comfortable balancing feature delivery ...

Forward Deployed Engineer

Hiring Organisation
Janus Henderson
Location
London, UK
Employment Type
Full-time
architecture and design decisions for your solutions, keeping them secure, scalable, and aligned to the enterprise core stack. Take solutions through evaluation, observability, and our AI governance checkpoints, and keep them healthy afterwards. Build investment capability on NexusBuild advanced investment capability on Nexus and make it available across the firm. … portfolio construction. Snowflake, Microsoft Fabric/OneLake, or comparable enterprise data platforms. Azure AI Foundry, model gateways, inference routing, or AI evaluation and observability in production. TypeScript or a second production language, and front-end experience for user-facing AI applications. Supervisory responsibilitiesNo. This is an individual-contributor role. Forward ...

Lead DevOps Engineer

Hiring Organisation
Shortlist Recruitment
Location
Chester, Cheshire, UK
Employment Type
Full-time
deployment solutions using IaC and GitOps practicesBuild and maintain CI/CD pipelines to enable development teams to deploy applications quickly and reliablyImprove monitoring, observability and incident response to enhance system reliability and resilienceContribute to the ongoing replatforming of services towards AWS and Kubernetes, identifying opportunities to automate, improve performance ...

Engineering & Science - Senior Software Engineering

Location
Crawley, England, United Kingdom
infrastructure automation. Understanding of software design patterns, asynchronous programming, and event‐driven systems. Experience in regulated industries (healthcare preferred), cybersecurity best practices, and observability tools desired. What You’ll Get Work pattern: required to work from the Crawley HQ 4 days a week with 1 day WFH. ...

Infrastructure Python Developer

Hiring Organisation
Hays
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
Up to £400.0 per day + Inside IR35
/CD pipelines and software delivery processes. Experience working within Agile environments and using tools such as JIRA. Desirable skills and experience Knowledge of observability tools such as Prometheus, Grafana and OpenTelemetry. Experience with enterprise tooling including Control-M, TrueSight, Guardium, Tenable Nessus or Delinea. Exposure to HashiCorp Vault ...

Infrastructure Python Developer

Hiring Organisation
Hays Specialist Recruitment Limited
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£400.00 per day
/CD pipelines and software delivery processes. Experience working within Agile environments and using tools such as JIRA. Desirable skills and experience Knowledge of observability tools such as Prometheus, Grafana and OpenTelemetry. Experience with enterprise tooling including Control-M, TrueSight, Guardium, Tenable Nessus or Delinea. Exposure to HashiCorp Vault ...

DevOps Engineer

Hiring Organisation
Richmond Square Consulting Limited
Location
Hereford, Herefordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
with secure cloud migrations and modern platform design Familiarity with containerisation, platform engineering, and secure CI/CD environments Strong understanding of cloud governance, observability, and operational security practice Please note: candidates must hold active MOD SC clearance and be willing to undergo DV clearance. This role also requires regular ...

Dev Ops Systems Administrator

Hiring Organisation
Proactive Appointments
Location
Woking, Surrey, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
environments Automate infrastructure and configuration using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement and improve monitoring and observability using Grafana, Prometheus and CloudWatch Drive improvements in system reliability, performance, scalability and security Manage IAM, networking, firewalls, VPNs and cloud security Take ownership ...

Platform/Infrastructure Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
configuration management and automation (Ansible or similar) for consistent, repeatable infrastructure changes. Support containerised workloads (Docker, Kubernetes or similar). Maintain logging, monitoring, and observability tooling (ELK stack or similar). Manage secure file transfer services (MFT/SFTP or similar) for partner and customer data exchange. Configure and troubleshoot ...