1,851 to 1,875 of 2,348 Remote/Hybrid Observability Jobs

Remote Senior Data Engineer

Hiring Organisation
Codat
Location
Newcastle upon Tyne, UK
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Principal Software Engineer (Hybrid)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
other teams shift faster and more reliably Taken AI and agentic systems from prototype to production, with a strong understanding of infrastructure, authentication, and observability Delivered on high-stakes consulting engagements across multiple language paradigms, stacks, ecosystems, and client industries Built high-quality, maintainable software collaboratively, incrementally, and through … legacy systems with short and long-term business needs Led and delivered solutions to architecture-level problems including scalability, security, reliability, performance, maintainability, and observability Facilitated alignment across technical and non-technical stakeholders to move initiatives forward through ambiguity and complexity Provided mentorship and team support at scale, while sharing ...

Remote Senior Data Engineer

Hiring Organisation
Codat
Location
Chester Le Street, County Durham, UK
choices you are making and why. Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code. Make AI your default way of working, and find opportunities to apply it across our products and pipelines where … query and reason over. What You'll Bring Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence. A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring ...

Senior Data Engineer

Location
Northam, England, United Kingdom
using techniques such as partitioning, indexing and cachingDeveloping and maintaining data quality processes to ensure data is accurate, complete and consistent Improving the reliability, observability and performance of data solutions Working collaboratively with product managers, developers and stakeholders to define data requirements and understand how data is used Diagnosing … warehouse such as Snowflake, BigQuery or Redshift Experience with automated testing, version control and CI/CD Strong understanding of data quality, reliability and observability Experience diagnosing performance and production issues Ability to make, document and communicate technical design decisions Experience mentoring or supporting other engineers Excellent written and verbal ...

Staff Machine Learning Engineer, Platform (MLOPS)

Hiring Organisation
Credit Acceptance Corporation
Location
United States
Employment Type
Permanent
Salary
USD 226,042 Annual
drift detection against pinned baselines, sampling and judge pipelines, and the release-gate mechanics that stop a regression from shipping. Build and maintain the observability and evaluation substrate other teams depend on: trace and telemetry capture including multi-turn and multi-step agent traces, logging standards, evaluation pipeline plumbing … because it is mandated. Partner with Cloud Engineering, Data Engineering, Security and SRE so the ML platform sits inside enterprise governance, identity and observability rather than beside it. Respond to AI-specific production incidents and drive them to a closed corrective action: prompt injection attempts, rogue-agent cost spikes, data ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Swindon, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Manchester, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Belfast, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Amersham, Buckinghamshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Clophill, Bedfordshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Kidwelly, Carmarthenshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Taunton, Somerset, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Kingsbury, Warwickshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Alnwick, Northumberland, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Stevenston, North Ayrshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Linlithgow, West Lothian, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Beverley, East Yorkshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Strathaven, South Lanarkshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Uckfield, East Sussex, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Brigg, North Lincolnshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Senior Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Chester Le Street, County Durham, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Agentic Platform Engineer · Manchester, UK ·

Location
Manchester, England, United Kingdom
Protocol (MCP). Build production-grade agent services using Python, cloud-native architectures, event-driven design, automation and Infrastructure as Code. Implement robust evaluation, observability and continuous improvement capabilities, including testing, tracing, telemetry and performance optimisation. Embed security, governance and responsible AI principles through least-privilege access, policy enforcement, auditability … workflow state, retrieval-augmented generation and human-in-the-loop patterns. Experience building secure integrations with enterprise APIs, repositories, cloud services, CI/CD, observability or ITSM platforms. Experience creating evaluation frameworks for agent quality, task completion, safety, reliability, latency and cost. Strong understanding of agent security, including workload identity ...

Software Engineering Manager

Location
Greater London, England, United Kingdom
appropriate technical documentation. Champion strong engineering practices, including collaborative programming, testing approaches such as TDD, CI/CD and production ownership. Ensure strong observability and operational health, with useful code‐quality and Production metrics, appropriate KPIs, well‐calibrated alarms and clear ownership of actions following incidents and PIRs. Build relationships … Data/Data Science, CRM, Security, Legal, Privacy and other engineering teams. Experience establishing strong operational ownership of production software, including CI/CD, observability, support practices and continuous improvement. Commitment to inclusive leadership, frequent feedback and creating an environment where engineers can grow, challenge ideas and take meaningful ownership. ...

Senior Ai Engineer

Hiring Organisation
Typeform
Location
United Kingdom, UK
Employment Type
Full-time
more conversational and personalised ways. The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance. You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … evaluating, and releasing AI systems. Help teams make informed decisions about models, frameworks, infrastructure, performance, and cost. Apply strong engineering practices across testing, security, observability, version control, and deployment. Share technical knowledge and support the development of other engineers. Keep up with relevant AI research, tools, and engineering practices, applying ...

Senior Engineer

Location
Greater London, England, United Kingdom
turn ideas into working product quickly Contribute to product decisions and pragmatic technical trade-offs Diagnose and fix bugs quickly Improve testing, monitoring, and observability Maintain data integrity, system stability, and security Work closely with technical leadership Collaborate with our senior technical advisor on architecture and technical direction Implement technical … processes Design and maintain data-compliant systems with privacy, security, and regulatory standards embedded by default (e.g. UK GDPR, NHS DSPT, MHRA guidance) Implement observability and metrics to understand product performance, user behaviour, and system health Ensure our systems and features are audit-ready and aligned with healthcare regulatory requirements ...