2,626 to 2,650 of 4,362 Permanent Observability Jobs

Senior TypeScript Engineer - AI, AWS & Microservices

Location
Greater London, England, United Kingdom
days in the office per week to collaborate with product teams. Expect modern tooling, AI‐assisted development, strong emphasis on reliability, security, and observability, plus a competitive salary and #J-18808-Ljbffr ...

Principal Platform Engineer

Location
Wales, United Kingdom
working on Mondays and Fridays Take technical ownership of the platforms that support Centerprise services and customer operations. You will lead improvements in reliability, observability and automation while remaining closely involved in complex engineering, major incidents and service recovery. Role Summary As Principal Platform Engineer, you will … hours escalation when required. Identify and address technical debt, operational risk and platform weaknesses. Ensure services remain supportable, recoverable and operationally efficient. Observability and service health Own monitoring and observability tooling, standards and operational dashboards. Develop service health metrics that provide clear and useful operational insight. Improve the quality ...

Backend Software Engineer – Infrastructure, Foundations

Location
Greater London, England, United Kingdom
compute workloads and efficiently schedule hundreds of thousands of containers hourly Design architecture and opinionated APIs that guide application developers Implement tracing and performance observability in high-scale distributed microservice architectures Build reliable, performant, and scalable systems for storage, authentication, and asset serving Automate deployment, management, and operations of distributed … keywords Distributed Systems Development Java Programming C++ Programming Python Programming Cloud Infrastructure Management ATS Optimization Keywords Hard Skills Software Engineering Data Processing Systems Performance Observability Microservice Architecture API Design Container Scheduling Open-Source Contribution High-Scale Systems Automation Prototyping Soft Skills Strong Communication Skills Team Collaboration Feedback Incorporation Quality Maintenance ...

Senior Backend Engineer - AI-Driven SaaS (London, Onsite)

Location
Greater London, England, United Kingdom
backend SaaS, strong Python with FastAPI, PostgreSQL, and AWS experience. The team collaborates across London and Bengaluru, with a focus on CI/CD, observability, and safe failure handling in high-stakes environments. #J-18808-Ljbffr ...

Data Engineer

Location
Greater London, England, United Kingdom
/CD and infrastructure as code. Create reusable components and maintain clear technical documentation. Quality & Governance (10%) : Implement robust data validation, testing, lineage and observability to ensure high-quality, trusted datasets. Support governance and privacy-conscious data handling. Collaboration & Enablement (10%) : Partner with Data Science, MLOps, Product and commercial teams … cloud environments (preferably AWS) Engineering Best Practice: Knowledge of CI/CD, testing, version control and infrastructure as code Data Quality & Governance: Understanding of observability, validation and maintaining reliable data systems Collaboration & Communication: Ability to translate business and data science needs into scalable solutions and communicate clearly with stakeholders Mindset ...

Software Engineer, Agentic AI Roku, Inc.

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development - from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Integration Architect - SAP SAAS Products

Location
Greater London, England, United Kingdom
4HANA Cloud, SuccessFactors, Ariba, Concur, Datasphere SAC) and SAP BTP. Define patterns, govern APIs/events, ensure secure, resilient data flows, and drive standardization, observability, and compliance. Core Responsibilities Architecture & Standards • Define canonical integration patterns (API-led, event-driven, batch/EDI) and reference architectures on SAP BTP. • Establish guidelines … SuccessFactors, Ariba, Concur, Datasphere, SAC, and S/4HANA Cloud integrations. • Strong security fundamentals (OAuth2, SAML, JWT, SCIM) and compliance awareness. • Hands-on with observability (Cloud ALM), performance tuning, and reliability engineering. • Experience with agile delivery, CI/CD (Git-based pipelines), and test automation for integrations. • Preferred Qualifications ...

Incident Manager

Location
Leeds, England, United Kingdom
engineering teams to fix the root causes generating non-actionable alerts. Partner with product and engineering teams to agree on monitoring thresholds and ensure observability is built into the development lifecycle. Qualifications 5+ years of experience in IT operations, site reliability engineering, observability engineering, or a senior technical monitoring role. … Deep hands‐on expertise with enterprise monitoring and observability platforms such as Zabbix, Splunk, DynaTrace, New Relic, Datadog, Grafana, or similar tools. Proven experience designing and implementing monitoring strategies for large-scale, distributed, customer-facing digital platforms. Strong understanding of alert engineering principles including threshold tuning, correlation rules, suppression logic ...

Software Engineer

Location
Greater London, England, United Kingdom
system. AI harnesses in production. Deploying agentic AI into real-world operational settings — acting on real money, tenancies and legal exposure, with the guardrails, observability and correctness that demands. Non-deterministic LLM working within compliant, secure deterministic software. Our stack We build on NestJS + TypeScript on GCP/… build Small, well-factored services. TDD and DDD as defaults. Trunk-based CI/CD — you ship to production and own it, with tests, observability and clean rollbacks. Lean frameworks, readable code, and we move fast because the tests and boundaries let us. What we're looking for #J ...

Sr Director, Enterprise AI

Location
Greater London, England, United Kingdom
vendors, platform integrations such as ServiceNow, Salesforce, and SAP, and custom-build trade-offs; bringing well-reasoned recommendations to executive leadership. Use Dynatrace's observability platform as a strategic advantage to monitor, measure, and continuously improve the reliability, performance, and business impact of internally deployed AI solutions. AI Governance …/partner decisions at the portfolio level, including vendor evaluation, contract negotiation support, and post-implementation performance tracking. Bonus: hands-on experience with observability or monitoring platforms - Dynatrace or similar - applied to AI system performance and reliability. Why you will love being a Dynatracer Dynatrace is a leader in unified ...

Staff Software Engineer - Customer Data Platform

Location
City of Westminster, England, United Kingdom
outcomes. Work with adjacent teams when needed to align on shared components and dependencies. Actively participate in on-call support, contributing to operational stability, observability, and performance. Coach and support other engineers through pairing, reviews, and mentoring, helping raise the team’s overall capability. Contribute to team OKRs and actively … that balance speed, maintainability, and long-term scalability. You actively identify technical debt, risks, or inefficiencies and take action to address them. You use observability and metrics to validate behavior, debug issues, and improve system health. You regularly support and unblock teammates, helping them deliver more effectively. You contribute ...

Platform Engineer, AI Enablement London, United Kingdom

Location
Greater London, England, United Kingdom
language models and AI agents. You’ll help create a governed model-access layer and a secure production environment for agentic workflows, with safety, observability, and operational excellence built in from the start. Your work will span the full lifecycle—from defining problems and designing systems to implementation, deployment … cost-effective access to AI models and tools. Build the runtime, services, and developer tooling used to run agentic workflows in production. Create observability across cost, performance, reliability, and usage through metrics, tracing, and logging. Implement governance and safety controls that make AI use secure, compliant, and auditable. Develop reusable ...

Head of Engineering

Location
Greater London, England, United Kingdom
down the business. Raise engineering quality Set clear standards for technical design, clean code, testing, code reviews, documentation, and release readiness. Improve automated testing, observability, production stability, and development workflows. Challenge weak technical decisions and help the team build better, more maintainable software. Introduce lightweight and predictable engineering processes without … perform deep code reviews, challenge architectural decisions, and solve complex production issues. Experience working with sensitive financial and personal data. Strong knowledge of testing, observability, incident management and software quality. Experience leading and developing engineering teams. Practical experience using AI across the software development lifecycle, beyond basic code generation. Strong ...

Senior Software Engineer II

Location
Greater London, England, United Kingdom
engineering standards, best practices and reusable patterns while partnering with Enterprise Architecture and influencing technical direction Drive engineering excellence by improving code quality, testing, observability, reliability and operational practices Support end-to-end delivery by guiding teams through complex technical challenges, improving decision-making, and contributing to planning and risk … data lakes/lakehouse architectures, Iceberg or similar table formats, as well as batch and streaming processing Knowledge of data quality, governance, cataloguing and observability tools (e.g. Datadog), with DBT or AI-assisted engineering practices as a plus Additional Information Your benefits We're a community here that cares ...

Senior Manager, Head of Engineering Standards & DevOps

Location
Eastleigh, England, United Kingdom
requirements for in-house and outsourced software and platform teams. Own the CI/CD pipeline strategy and platform tooling, integrating security, testing and observability into the developer experience. Govern technical quality across DTO while driving the adoption of modern engineering practices, automation and cloud-native capabilities. Position Requirements … standards, quality governance, DevOps practices and technical assurance across internal and third-party teams. Deep knowledge of modern software engineering, CI/CD, automation, observability and cloud-native delivery. A degree-level qualification in Computer Science, Software Engineering, Information Technology or a related discipline, or equivalent relevant experience. About ...

Software Engineer - Workflow Authoring & Deployment

Location
Greater London, England, United Kingdom
that let operators understand what robots are doing and why Build and maintain reliable, secure cloud infrastructure for running these services, including deployment pipelines, observability, and alerting Implement testing and validation pipelines to catch regressions across the UI, backend, and on-robot services - including simulation-based workflow validation before changes … Experience with cloud infrastructure and deployment: containerisation, CI/CD, infrastructure-as-code Comfort working across the stack from frontend to backend Experience with observability and alerting in systems where failures have real operational consequences Good instincts for where complexity should live and how to keep systems debuggable and understandable ...

NOC Engineer / SRE

Location
United Kingdom
Network Operations Center (NOC) responsibilities and engineering‐driven reliability practices. This role focuses on 24/7 service reliability, incident response, operational automation, and observability, while actively reducing operational toil through software and automation. Unlike a traditional NOC analyst, an SRE‐NOC is expected to engineer problems away, not just … Security, and Product teams Execute and improve runbooks, playbooks, and escalation paths Drive blameless post‐incident reviews (PIRs) and track corrective actions Monitoring, Alerting & Observability Own service health monitoring across infrastructure, applications, and dependencies Design and maintain alerting strategies that align with SLIs/SLOs Build dashboards using tools such ...

Senior Infrastructure Software Engineer

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … compute systems Integrate software and automation with hardware management and provisioning systems Improve tooling and workflows for managing infrastructure throughout its production lifecycle Reliability & Observability Build telemetry, logging, and observability capabilities that provide visibility into the health of our infrastructure Develop tools and automation to identify, diagnose, and respond ...

Senior Director of Software Engineering

Location
Greater London, England, United Kingdom
delivery across the SDLC, ensuring predictable execution and front-office responsiveness. Establish delivery rhythms and drive a production-first culture focused on observability, performance, and operational controls. Lead the design and delivery of integrated workflows, ensuring consistency and traceability across the trade lifecycle. Coordinate cross-team delivery for clean, aligned … experience with front-office UI/tooling (e.g., React/TypeScript) and automation for workflow-critical systems. Experience modernizing legacy estates and improving resilience, observability, and incident management. #J-18808-Ljbffr ...

Senior Associate Infrastructure Engineer (Windows)

Location
Tipton, England, United Kingdom
Executes increasing variety of user stories with different contexts across the different stages of the product lifecycle Uses automation, system tools, open-source solutions, observability and 'security first' principles in daily work Contributes to team agile ceremonies, leads demos and presentations, helps new engineers learn established norms Comfortably implements solutions … using automation, observability and security principles Engages in adoption of new technology Continues professional education; learns key product capabilities and has desire to learn higher levels of craftsmanship in day-to-day engineering tasks Will be part of a 24x7 on call rotation Minimum Qualifications At a minimum,here’swhat ...

Senior Platform Architect

Location
Greater London, England, United Kingdom
ensure on-time, high-quality delivery Proactively manage technical risk, migrations, and architectural debt Ensure Long-Term Technical Health Set and validate reliability, observability, and operational standards Monitor health signals and intervene before the customer knows they're happening Lead technical reviews, readiness checkpoints, and maturity assessments Certify teams … Willingness to travel occasionally for customer engagements and team events Nice to Haves Prior experience with Temporal or similar workflow orchestration platforms Experience with observability tools and practices for distributed systems Background in professional services, technical consulting, or solutions engineering roles Experience with Kubernetes, cloud platforms, and infrastructure-as-code ...

Remote DevOps Team Lead

Hiring Organisation
Runware
Location
Remote, UK
operation of Runware’s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Cloud Advisory Architecture Consultant

Location
Greater London, England, United Kingdom
where GenAI and Agentic play a role. Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps. Ensure Robustness & Security: Own the design of enterprise-wide applications that are highly available, fault-tolerant ...

Remote DevOps Team Lead

Hiring Organisation
Runware
Location
Wrexham, Wales, UK
operation of Runware’s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...

Remote DevOps Team Lead

Location
Stirling, Perthshire, United Kingdom
operation of Runware s infrastructure and orchestration systems Build automation and tooling to streamline model deployments, scaling, and hardware utilisation across distributed nodes Drive observability, alerting, and reliability practices to detect and resolve issues quickly and proactively Collaborate with engineers to optimise throughput, latency, and platform performance at every layer … similar languages Understand container runtimes like Docker and containerd, and have built or worked with orchestration systems beyond Kubernetes Are fluent in observability and debugging practices across distributed systems, using logs, metrics, traces, and profiling to drive insight and reliability Care deeply about reliability, efficiency, and engineering quality, and know ...