1,451 to 1,475 of 4,209 Observability Jobs

Senior Software Engineer, ML Infrastructure Roku, Inc.

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Platform Engineer - Common Platform

Location
Greater London, England, United Kingdom
develop automation and platform tooling Designing and implementing AWS-native infrastructure and services Building and maintaining standard CI/CD pipelines and workflows Supporting observability and monitoring capabilities across the platform Managing Kubernetes upgrades and platform improvements Supporting production incidents and resolving complex infrastructure issues Reviewing and approving infrastructure … design Enterprise security and governance Compliance frameworks and organisational policies Implementing platform changes across distributed teams Cloud cost optimisation Helm and Kubernetes management tooling Observability and monitoring You’ll be someone who enjoys building platforms and solving engineering problems rather than simply maintaining existing infrastructure. Strong communication and collaboration skills ...

AI-Driven Cloud & Platform Engineer for Digital Factory

Location
Manchester, England, United Kingdom
implement CI/CD automation, infrastructure-as-code, and self-service tooling, leveraging Docker, Kubernetes, and multiple cloud providers. You’ll define observability and reliability practices, collaborate with architects, and mentor teammates. You should bring hands-on experience in cloud, DevOps, and SRE disciplines, with a willingness to work across ...

AI-Driven Cloud & Platform Engineer for Digital Factory

Location
West of England, England, United Kingdom
implement CI/CD automation, infrastructure-as-code, and self-service tooling, leveraging Docker, Kubernetes, and multiple cloud providers. You’ll define observability and reliability practices, collaborate with architects, and mentor teammates. You should bring hands-on experience in cloud, DevOps, and SRE disciplines, with a willingness to work across ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Location
United Kingdom
This is a hands-on infrastructure role. You'll work across the full stack - from Terraform modules and Puppet manifests to Python automation and observability pipelines. The platform is not containerised - there is no Kubernetes here - so we're looking for someone who understands Linux systems deeply and can reason … service discovery and configuration management (etcd) Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache) Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotations Evaluating and implementing distributed storage solutions as the platform evolves Improving deployment workflows ...

AI-Driven Cloud & Platform Engineer for Digital Factory

Location
Newcastle upon Tyne, England, United Kingdom
implement CI/CD automation, infrastructure-as-code, and self-service tooling, leveraging Docker, Kubernetes, and multiple cloud providers. You’ll define observability and reliability practices, collaborate with architects, and mentor teammates. You should bring hands-on experience in cloud, DevOps, and SRE disciplines, with a willingness to work across ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Senior Cloud Engineer - Contract

Hiring Organisation
Flagstone
Location
London, UK
Employment Type
Full-time
Engineering builds and runs the new Azure cloud that everything else at Flagstone depends on. It's the foundation for our security tooling, our observability, and the hosting for our AI. The team works in infrastructure-as-code (Terraform and Bicep), owns the landing zones and hub-and-spoke networking … roll out golden paths that cut delivery cost and speed up engineering squads. Stand up our AI platform foundations: an AI gateway, an observability stack, and hosting for AI tools including Flagstone Concierge, with model access through AWS Bedrock. Build and maintain infrastructure pipelines with security scanning, plan validation ...

Vice President, Production Services Application Support

Location
Manchester, England, United Kingdom
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe ...

Full Stack Developer - React & Node.js

Hiring Organisation
Shree Narayani Networking Solutions Pvt Ltd
Location
San Mateo, California, United States
Employment Type
Permanent
Salary
USD Annual
party systems. Comfort with event-driven or streaming systems feeding real-time consumers, and the reliability mindset that comes with owning a live service (observability, retries, graceful degradation). Ability to own a service across the stack - stand it up, deploy it, monitor it - without needing hand-holding on infra. … modeling - designing canonical schemas and reconciling inconsistent inputs across sources. Real-time/event-driven systems - message queues or streaming (e.g., Kafka), and the observability tooling to run them. Cloud & operations - deploying and monitoring services in a cloud environment (AWS/GCP), containers, CI/CD. Strong signals (nice ...

Staff Analytics Platform Engineer

Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Site Reliability Engineer

Location
Fenny Stratford, England, United Kingdom
hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Reliability testing and capacity planning Toil reduction We offer a hybrid working arrangement with one day per week in our Milton Keynes office. Key Responsibilities: Support … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands‐on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Site Reliability Engineer

Hiring Organisation
Connells Limited
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
hands-on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post-incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer a hybrid working arrangement with one day per week … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands-on experience with Azure Monitoring (Application Insights, Alerts, Action ...

AI Engineer

Location
Greater London, England, United Kingdom
products. As a Senior/Lead AI Engineer, you will own and drive the development of our core agentic frameworks, evaluation pipelines, and observability tooling to ensure they operate with safety, trust, and intelligence at scale. You won't just be integrating AI into a product, you will lead … CrewAI is a plus. Production Python engineering : Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability : Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations : Distributed systems, scalable data pipelines (Kafka ...

Security Platform Engineer

Location
Farnborough, England, United Kingdom
agents Hands‐on experience with Kubernetes Experience managing and administering SIEM tooling (e.g. Splunk) Experience deploying vulnerability scanning and analysis tooling (e.g. Nessus) Kubernetes observability (e.g. fluentbit) Familiarity with container security principles and tools Scripting or automation skills (e.g. Python, Bash, or similar) Understanding of security frameworks and best practices … Knowledge of configuring SIEM tooling Basic understanding of threat frameworks, such as ATT&CK Experience with additional SIEM or observability platforms Experience with Microsoft Defender Experience with DevSecOps practices and pipeline security Certifications such as CKA, CKAD, CISSP, CEH, or similar Exposure to threat modelling and security architecture design Just ...

Network Automation Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
Python, Ansible, Terraform and Jinja2Integrating network automation into CI/CD pipelines for reliable, repeatable deploymentsCreating APIs and self-service tooling for engineering teamsImplementing observability and telemetry solutions for performance and reliabilityPartnering with network, platform and security teams to deliver resilient, scalable systemsContributing to incident response and production reliabilityOn-call … tools such as Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude CodeFamiliarity with Docker and KubernetesExposure to monitoring, observability or telemetry in distributed systemsPragmatic problem solver who can operate in ambiguity and take ownershipComfortable working in collaborative, fast-paced engineering teamsDeep understanding of networking ...

Senior Software Engineer

Hiring Organisation
Ask4.com
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
Salary
£60,000
Kubernetes configurations for production and non-production environments Integrate with network management systems, message brokers, and third-party APIs Instrument and monitor applications using observability tooling (metrics, logs, and traces Grafana, Prometheus, or similar) Provide technical mentorship to mid-level engineers and act as an escalation point for complex problems … Experience with message brokers and event-driven systems (NATS, RabbitMQ, Kafka, or similar) Exposure to OpenWiFi, OpenWrt or similar open-source network controller frameworks Observability experience. Use of metrics, logs, and traces using tools such as Grafana, Prometheus and Sentry Experience with AI agent development or LLM integration Agile/ ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health ...

Senior Software Engineer-AI

Hiring Organisation
Moodys
Location
London, UK
Employment Type
Full-time
ongoing operation with limited supervisionHands-on experience with cloud-native technologies, serverless applications, event-driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelinesSolid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade-offsExperience mentoring engineers through code … integrationContribute to technical designs, participate in design reviews, and identify risks, constraints, trade-offs, and alternative approachesMaintain engineering excellence through automated testing, code reviews, observability, monitoring, alerting, operational readiness, and participation in on-call supportApply machine learning operations practices, including prompt versioning, automated evaluation, deployment pipelines, monitoring, and production issue ...

Network Automation Engineer

Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self‐service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with network, platform and security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude Code Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Deep understanding ...

Developer – Scala

Location
Newcastle upon Tyne, England, United Kingdom
frontend engineers, QA, product owners, solution designers, and other backend developers to deliver high-quality product increments. Support production stability by investigating issues, improving observability, and continuously reducing technical debt. Required Skills and Experience: Professional backend development experience, ideally in enterprise, SaaS, or cloud-based product environments. Strong hands … problems. Nice to Have: Experience with Squeryl, Doobie, or similar Scala data access libraries. Familiarity with Grafana or Kibana dashboards, alerting, logging, and production observability practices. Experience with large distributed systems, horizontal scaling, resilient service design, or high-throughput Play/Pekko applications. Previous experience in Payroll ...

Tivoli Netcool OMNIbus SME

Hiring Organisation
Proactive Appointments
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£750.00 - £815.00 per day
effective monitoring coverage. Define platform standards, monitoring policies and best practices. Maintain technical documentation, architecture diagrams and support procedures. Contribute to the monitoring and observability roadmap and identify opportunities for platform modernisation. Provide technical mentoring and knowledge sharing across engineering and operational teams. Essential Experience Significant hands-on experience administering … integration technologies. Understanding of monitoring across AWS, Azure and Google Cloud . Experience with container and Kubernetes monitoring . Knowledge of enterprise observability frameworks and modern SRE practices . Tivoli Netcool OMNIbus SME Due to the volume of applications received for positions, it will not be possible to respond ...

Forward Deployed Agentic AI Engineer

Location
Greater London, England, United Kingdom
solutions that solve real-world business challenges. You will bring deep expertise across modern full-stack technologies, distributed systems, cloud-native architectures, and observability, combined with hands‐on experience developing enterprise‐grade AI applications. You will design, build, and deploy intelligent AI agents, copilots, and automation solutions using Anthropic Claude … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Principal Agentic Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
Architect, you will define the architecture that enables AI agents to reason, automate, and operate at enterprise scale while ensuring security, governance, reliability, and observability remain foundational. You will establish the long-term vision for how AI agents, human operators, and digital platforms work together to create a highly automated … patterns, and governance frameworks for AI agents and multi-agent systems. Design the integration architecture connecting agents to enterprise systems including ServiceNow, cloud platforms, observability platforms, CMDB, developer platforms, and security tooling. Develop reference architectures for AI-powered incident response, service management, platform operations, FinOps, cybersecurity, and disaster recovery. Partner ...