1,901 to 1,925 of 2,349 Remote/Hybrid Observability Jobs

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

Principal Agentic Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
Architect, you will define the architecture that enables AI agents to reason, automate, and operate at enterprise scale while ensuring security, governance, reliability, and observability remain foundational. You will establish the long-term vision for how AI agents, human operators, and digital platforms work together to create a highly automated … patterns, and governance frameworks for AI agents and multi-agent systems. Design the integration architecture connecting agents to enterprise systems including ServiceNow, cloud platforms, observability platforms, CMDB, developer platforms, and security tooling. Develop reference architectures for AI-powered incident response, service management, platform operations, FinOps, cybersecurity, and disaster recovery. Partner ...

Forward Deployed Agentic AI Engineer

Location
Greater London, England, United Kingdom
solutions that solve real-world business challenges. You will bring deep expertise across modern full-stack technologies, distributed systems, cloud-native architectures, and observability, combined with hands‐on experience developing enterprise‐grade AI applications. You will design, build, and deploy intelligent AI agents, copilots, and automation solutions using Anthropic Claude … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...

Senior Full Stack Engineer - Lyst Shop (11 Month FTC - Maternity Cover)

Location
Greater London, England, United Kingdom
rely heavily on experimentation to validate ideas and guide decisions. Technical Excellence: You will help maintain a high bar for code quality, testing, observability, and system reliability. You’ll contribute to architectural discussions, improve developer experience, and proactively address technical debt where needed. Team Contribution: You will mentor and support … working relationships across Product, Design, QA, Analytics, and Engineering teams while actively participating in team ceremonies and technical discussions. Technical Impact: Improve the stability, observability, and maintainability of our systems through better monitoring, resilient code, and thoughtful testing practices. Growth & Ownership: Gain confidence working across our platform and infrastructure while ...

Staff AI Engineer - EU

Hiring Organisation
Typeform
Location
United Kingdom, UK
Employment Type
Full-time
turn responses into useful insights. The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance. You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … decisions and make trade-offs explicit. Mentor engineers and support other technical leads in growing their ownership and judgement. Improve engineering practices across testing, observability, security, incident response, and deployment. Build alignment around technical decisions through clear proposals, constructive discussion, and evidence. Evaluate relevant AI research and emerging tools ...

Junior DevOps Engineer

Location
Greater London, England, United Kingdom
/CD processes, automate manual processes to create a self-service environment for our developers, and maintain platform uptime SLAs by improving our observability stack. We use infrastructure-as-code to maintain our platform on AWS, so familiarity with common AWS Services (RDS, S3, ECS, EC2 etc.) and Terraform …/CD pipelines through Jenkins/Github Actions and other technology Support the Development and AI Engineering Teams - help troubleshoot their issues Bring observability through dashboards, alerting and log aggregation Run incident analysis and post mortems Environment management: ephe...dev/staging/prod What you’ll bring A strong foundation ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Senior AI Engineer - Agentic AI

Location
Greater London, England, United Kingdom
that allow complex AI workflows to operate securely and efficiently at scale. You will be responsible for developing advanced orchestration capabilities, implementing evaluation and observability tooling and embedding enterprise controls for compliance and safety. If you are passionate about innovating with AI in real-world applications and scaling intelligent systems … ensuring graceful degradation and retries Apply enterprise security and governance practices including RBAC, prompt safety checks, traceability and secrets management Implement evaluation pipelines and observability frameworks using tools such as Langfuse, Arize or OpenTelemetry Contribute to architectural design decisions, code reviews and engineering standards for platform development Requirements Bachelor ...

Mid-Level Data Engineer (Python/ AWS)

Location
Belfast, Northern Ireland, United Kingdom
consumers Contribute to cloud-based data platform development in AWS Support lightweight frontend work (React/TypeScript) for data-focused tools where needed Implement observability practices (logging, monitoring, tracing) across data pipelines and APIs Improve reliability, performance, and failure handling across the platform Collaborate with engineers and analysts to deliver … skills and experience working in cross-functional teams Nice to Have Experience with TypeScript or JavaScript Experience contributing to frontend applications (React) Familiarity with observability tooling (logging, monitoring, tracing) Experience with data modelling and metadata management Exposure to CI/CD practices for data pipelines and services Experience working ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Middlesbrough, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Bridgend, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Leicester, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Acrefair, Wrexham, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Dawlish, Devon, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Cambridge, Cambridgeshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Abingdon, Oxfordshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Hitchin, Hertfordshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Mauchline, East Ayrshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
East Horsley, Surrey, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Newcastle upon Tyne, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Lanark, South Lanarkshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Alexandria, West Dunbartonshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Remote Principal Platform Engineer (12 Month FTC)

Hiring Organisation
grabjobs
Location
Burton-on-Trent, Staffordshire, UK
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Senior Data Engineer

Location
Epsom, England, United Kingdom
Engineering team operates a distributed leadership model. Each Senior Data Engineer owns a defined functional area, such as ingestion and integration, incident management and observability, governance, security and cost management, or CI/CD and DevOps, and is accountable for the standards and resilience of that area so that nothing … with structured support. Excellent communication skills, including documenting technical design proposals and translating complex technical concepts for technical and non-technical audiences. Desirable: Observability and incident management tooling, such as New Relic and ServiceNow. CI/CD, DevOps and infrastructure as code, such as Terraform and Azure DevOps. Data governance ...