1,126 to 1,150 of 3,664 Remote/Hybrid Observability Jobs

Remote Senior AI Software Engineer

Location
Welshpool, Radnorshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Location
Kirkliston, Midlothian, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Location
Harrogate, West Yorkshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Location
Downpatrick, County Down, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Location
Stockton-on-Tees, Durham, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Hiring Organisation
Learning Technologies Group
Location
United Kingdom
Salary
£ 70 K
code.This is a hands-on infrastructure role. You'll work across the full stack — from Terraform modules and Puppet manifests to Python automation and observability pipelines. The platform is not containerised — there is no Kubernetes here — so we're looking for someone who understands Linux systems deeply and can reason … service discovery and configuration management (etcd)Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache)Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotationsEvaluating and implementing distributed storage solutions as the platform evolvesImproving deployment workflows and release ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Location
United Kingdom
This is a hands-on infrastructure role. You'll work across the full stack - from Terraform modules and Puppet manifests to Python automation and observability pipelines. The platform is not containerised - there is no Kubernetes here - so we're looking for someone who understands Linux systems deeply and can reason … service discovery and configuration management (etcd) Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache) Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotations Evaluating and implementing distributed storage solutions as the platform evolves Improving deployment workflows ...

Staff Analytics Platform Engineer

Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Site Reliability Engineer

Location
Fenny Stratford, England, United Kingdom
hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Reliability testing and capacity planning Toil reduction We offer a hybrid working arrangement with one day per week in our Milton Keynes office. Key Responsibilities: Support … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands‐on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Site Reliability Engineer

Hiring Organisation
Connells Limited
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
hands-on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post-incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer a hybrid working arrangement with one day per week … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands-on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health ...

Forward Deployed Agentic AI Engineer

Location
Greater London, England, United Kingdom
solutions that solve real-world business challenges. You will bring deep expertise across modern full-stack technologies, distributed systems, cloud-native architectures, and observability, combined with hands‐on experience developing enterprise‐grade AI applications. You will design, build, and deploy intelligent AI agents, copilots, and automation solutions using Anthropic Claude … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Principal Agentic Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
Architect, you will define the architecture that enables AI agents to reason, automate, and operate at enterprise scale while ensuring security, governance, reliability, and observability remain foundational. You will establish the long-term vision for how AI agents, human operators, and digital platforms work together to create a highly automated … patterns, and governance frameworks for AI agents and multi-agent systems. Design the integration architecture connecting agents to enterprise systems including ServiceNow, cloud platforms, observability platforms, CMDB, developer platforms, and security tooling. Develop reference architectures for AI-powered incident response, service management, platform operations, FinOps, cybersecurity, and disaster recovery. Partner ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...

Staff AI Engineer - EU

Hiring Organisation
Typeform
Location
United Kingdom
Salary
£ 70 K
questions, and turn responses into useful insights.The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance.You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … difficult technical decisions and make trade-offs explicit.Mentor engineers and support other technical leads in growing their ownership and judgement.Improve engineering practices across testing, observability, security, incident response, and deployment.Build alignment around technical decisions through clear proposals, constructive discussion, and evidence.Evaluate relevant AI research and emerging tools, and help teams ...

Senior Full Stack Engineer - Lyst Shop (11 Month FTC - Maternity Cover)

Location
Greater London, England, United Kingdom
rely heavily on experimentation to validate ideas and guide decisions. Technical Excellence: You will help maintain a high bar for code quality, testing, observability, and system reliability. You’ll contribute to architectural discussions, improve developer experience, and proactively address technical debt where needed. Team Contribution: You will mentor and support … working relationships across Product, Design, QA, Analytics, and Engineering teams while actively participating in team ceremonies and technical discussions. Technical Impact: Improve the stability, observability, and maintainability of our systems through better monitoring, resilient code, and thoughtful testing practices. Growth & Ownership: Gain confidence working across our platform and infrastructure while ...

Staff AI Engineer - EU

Hiring Organisation
Typeform
Location
United Kingdom, UK
Employment Type
Full-time
turn responses into useful insights. The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance. You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … decisions and make trade-offs explicit. Mentor engineers and support other technical leads in growing their ownership and judgement. Improve engineering practices across testing, observability, security, incident response, and deployment. Build alignment around technical decisions through clear proposals, constructive discussion, and evidence. Evaluate relevant AI research and emerging tools ...

Senior AI / Agentic Engineer London

Location
Greater London, England, United Kingdom
Agentic ecosystem, responsible for the high-level design choices that define how agents run at PhysicsX. You will cover topics such as: Agent Observability: Own the implementation to enforce deep tracing, granular cost tracking, and observability across the lifecycle. Agent Deployment: Deliver an intuitive deployment lifecycle which simplifies questions around … behalf of users in a regulated enterprise environment. The Tech Stack Core Platform: Python (Primary), Go or TypeScript (Secondary), Kubernetes, Docker, Terraform. Observability & Evals: OTel, LangSmith, Arize, Braintrust. Who You Are An Architect at Heart: You have strong, reasoned opinions on Durable Execution vs. Standard Async, Vector Search vs. Keyword ...

Senior AI Engineer - Agentic AI

Location
Greater London, England, United Kingdom
that allow complex AI workflows to operate securely and efficiently at scale. You will be responsible for developing advanced orchestration capabilities, implementing evaluation and observability tooling and embedding enterprise controls for compliance and safety. If you are passionate about innovating with AI in real-world applications and scaling intelligent systems … ensuring graceful degradation and retries Apply enterprise security and governance practices including RBAC, prompt safety checks, traceability and secrets management Implement evaluation pipelines and observability frameworks using tools such as Langfuse, Arize or OpenTelemetry Contribute to architectural design decisions, code reviews and engineering standards for platform development Requirements Bachelor ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Junior DevOps Engineer

Location
Greater London, England, United Kingdom
/CD processes, automate manual processes to create a self-service environment for our developers, and maintain platform uptime SLAs by improving our observability stack. We use infrastructure-as-code to maintain our platform on AWS, so familiarity with common AWS Services (RDS, S3, ECS, EC2 etc.) and Terraform …/CD pipelines through Jenkins/Github Actions and other technology Support the Development and AI Engineering Teams - help troubleshoot their issues Bring observability through dashboards, alerting and log aggregation Run incident analysis and post mortems Environment management: ephe...dev/staging/prod What you’ll bring A strong foundation ...

Mid-Level Data Engineer (Python/ AWS)

Location
Belfast, Northern Ireland, United Kingdom
consumers Contribute to cloud-based data platform development in AWS Support lightweight frontend work (React/TypeScript) for data-focused tools where needed Implement observability practices (logging, monitoring, tracing) across data pipelines and APIs Improve reliability, performance, and failure handling across the platform Collaborate with engineers and analysts to deliver … skills and experience working in cross-functional teams Nice to Have Experience with TypeScript or JavaScript Experience contributing to frontend applications (React) Familiarity with observability tooling (logging, monitoring, tracing) Experience with data modelling and metadata management Exposure to CI/CD practices for data pipelines and services Experience working ...

Senior Data Engineer

Location
Epsom, England, United Kingdom
Engineering team operates a distributed leadership model. Each Senior Data Engineer owns a defined functional area, such as ingestion and integration, incident management and observability, governance, security and cost management, or CI/CD and DevOps, and is accountable for the standards and resilience of that area so that nothing … with structured support. Excellent communication skills, including documenting technical design proposals and translating complex technical concepts for technical and non-technical audiences. Desirable: Observability and incident management tooling, such as New Relic and ServiceNow. CI/CD, DevOps and infrastructure as code, such as Terraform and Azure DevOps. Data governance ...