401 to 425 of 4,013 Observability Jobs

Senior Data Engineer (Data Platform)

Location
Greater London, England, United Kingdom
other parts of Teya Guaranteeing proper data governance following best practices, while not adding 100 steps of bureaucracy Improving data reliability, quality, and observability across key datasets, while ensuring data pipelines are simple to set up. Building, maintaining, and improving data models for large and complex datasets, while ensuring they … experience provisioning and managing cloud infrastructure with Terraform Experience with CI/CD pipelines, automated testing and Git-based development workflows Familiarity with observability practices, including logging, metrics, alerting and production troubleshooting. Strong grasp of software engineering principles and best practices Experience contributing to or leading data warehouse architecture ...

Automation & Platform Engineer

Location
Newbury, England, United Kingdom
services, agent services, APIs, and microservices.* Implement infrastructure-as-code for platform environments and network automation resources.* Ensure automation platforms meet security, compliance, availability, observability, and operational resilience requirements.## **Your Profile*** Experience in network automation, platform engineering, DevOps, cloud engineering, or telecom automation.* Strong hands-on experience with Ansible, Terraform … vendor APIs.* API management and orchestration.* Intent-to-configuration workflows.* Data pipelines.* Vector databases and graph APIs.* MCP integration.* Security frameworks and compliance controls.* Observability and logging.**Preferred Certifications*** Kubernetes CKA/CKAD.* Terraform Associate.* Red Hat Ansible certification.* Google Cloud, AWS, or Azure certification.* Cisco, Juniper, Nokia, or Ericsson ...

Senior Site Reliability Engineer

Hiring Organisation
Imanage
Location
London, United Kingdom
Salary
£ 70 K
cloud platform, collaborate across teams to promote standardization and resiliency, and participate in on-call rotations. You’ll be a key voice in observability, change management, and service scalability, providing guidance during complex technical decisions and high impact events. iManage is experiencing explosive growth in its flagship cloud product. … support our Kubernetes-based ecosystem. Maintaining the freshness and utility of platform services. Improving the security posture of our products. Designing automation, orchestration, observability, and disaster readiness into our products. Participating in production support and on-call rotations, providing senior-level guidance during critical events. Leading incident management and post ...

Deployed Architect, Professional Services (London)

Location
Greater London, England, United Kingdom
together. LangChain is a place where your contributions can shape how this technology shows up in the real world. Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. … backup strategies, and sizing Experience designing high-availability and disaster recovery solutions Strong understanding of networking, security (SSO/RBAC, TLS, secrets management), and observability (Prometheus, Grafana, Datadog) Experience with CI/CD pipelines for infrastructure and applications Agent Engineering & Development: 1+ years of experience building production AI/ ...

Senior Python Developer

Hiring Organisation
Tech 4
Location
City of London, London, United Kingdom
Employment Type
Permanent
integrations, and human-in-the-loop controls evolve across the stack. Continuously identify and exploit opportunities to improve performance, reliability, and user experience, using observability and analysis to find signals in noisy systems. Navigate confidently across legacy and greenfield contexts, applying AI tooling pragmatically to modernise where it matters most. … evolve architecture pragmatically. Strong system design fundamentals across scalability, performance, and distributed systems, including API design (REST, GraphQL). Hands-on experience with observability tooling (Datadog, Grafana, or similar) and a data-informed approach to system health and reliability. Solid SQL and data management skills, with an appreciation ...

Platform Lead - ML Ops

Location
Greater London, England, United Kingdom
scaling, spot instances, and down-scaling policies to eliminate waste, while providing full visibility into the unit economics of training and serving LLM models. Observability & Incident Response: Establish 24/7 incident response, telemetry, and observability metrics to monitor system performance, model drift, and data pipelines. Data Governance & Security: Enforce ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Hiring Organisation
Goldman Sachs
Location
Birmingham, West Midlands (County), United Kingdom
Salary
£ 70 K
operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. Define and maintain service level indicators (SLIs), service level objectives (SLOs), and error budgets to quantify and manage service reliability. Engineer automation … development lifecycle concepts, and developing applications in a Linux environment. Strong understanding of algorithms, data structures, software design, and distributed systems fundamentals. Experience with observability platforms, including distributed tracing, logging, metrics, and tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience with site reliability engineering practices, relational databases, Hadoop ...

Principal Engineer

Location
Greater London, England, United Kingdom
changes, unblocking delivery and coaching engineers through real work. Drive engineering excellence through test automation, trunk-based development, CI/CD optimisation, secure coding, observability, service-level objectives and effective production incident response. Embed AI-assisted engineering practices into day-to-day delivery, including drafting, review support, test generation … data and infrastructure. Deep experience with cloud-native platforms, Kubernetes or equivalent container orchestration, infrastructure-as-code, CI/CD pipelines, progressive delivery, observability and production operations. Confidence using AI-enabled engineering tools to improve quality, speed and traceability, while maintaining strong human oversight. Proven ability to build ...

Staff Software Engineer - AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Staff Software Engineer - AI

Location
London, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Lead .Net Software Engineer

Location
Cheltenham, England, United Kingdom
technical risks, bottlenecks, dependencies, and opportunities for improvement Contribute to AWS‐based architecture and engineering practices, including environments using Lambda, ECS, and EC2 Improve observability, reliability, performance, security, and operational readiness across the platform Contribute to Jenkins pipelines and CI/CD practices to improve consistency and delivery efficiency Help … with DevOps practices and infrastructure‐aware development Experience with messaging, event streaming, or related event‐driven technologies Experience with Blazor and MudBlazor Familiarity with observability and monitoring tooling in distributed systems Experience defining governance, controls, or operating models for AI agents or AI‐enabled internal tools Experience working ...

Staff Software Engineer - AI

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
technologies in production environments • Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines • Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency • Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence • Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement • Partner with ...

Data Operations DevOps Engineer

Location
United Kingdom
/CD and GitOps workflows to support deployment and improve the team’s release management processes. Drive operational excellence and security practices: improve observability, documentation and security practices across SRCNet services. Experience working in a professional DevOps engineering capacity for at least two years Postgraduate qualification in Astronomy, Physics, Computer … lifecycle management of containerised services. Demonstrable ability to diagnose and resolve complex operational issues across application, infrastructure and networking layers, supported by experience with observability and operational tooling (e.g. Prometheus, Grafana, Elastic stack), and the ability to manage workloads and targets in a dynamic and collaborative environment. Experience of building ...

Lead AI Engineer

Location
Sheffield, England, United Kingdom
Agentic AI integrations within enterprise environments. Collaborate with architects and engineering teams to ensure scalable, secure, and maintainable solutions. Define standards for AI observability, governance, security, and performance. Mentor engineers and provide technical leadership across AI development initiatives. Contribute hands‐on to solution design, development, code reviews, and production deployment. ...

Senior Software Development Engineer

Location
Reading, England, United Kingdom
agile delivery environments (e.g. Scrum) Desirable Experience with cloud platforms (e.g. Azure) Experience with CI/CD pipelines and release processes Experience with observability tools (monitoring, logging, alerting) Experience in technical leadership or mentoring roles Experience working in safety- or regulation-driven environments Technology Stack .NET (latest versions) Kubernetes & Docker … Azure (SQL, CosmosDB, cloud services) PostgreSQL TypeScript/modern web frameworks (e.g. Vue.js) Observability tooling (e.g. Azure Monitor, Prometheus, Grafana) Azure DevOps/CI-CD pipelines Holidays: 25 days per annum + 8 days bank holidays (options to buy/sell days) 37.5 hour working week Pension – 4% employee ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
Lead code reviews, mentor engineers, and ensure high engineering standards. Contribute to platform SDKs, internal libraries, and backend frameworks used across charters. Performance, Reliability & Observability Implement tracing, structured logging, metrics, dashboards, and alerting. Optimize services for latency, concurrency, throughput, and cost efficiency. Ensure system reliability through automated testing, load testing … Kubernetes, and cloud platforms (AWS or GCP). Strong CI/CD experience using GitHub Actions, Jenkins, Argo, or similar tools. Deep knowledge of observability practices — logs, metrics, traces, performance analysis. Proven ability to write clean, testable, well‐structured production code. Qualifications – Nice to Have Experience with streaming or queuing ...

ML Ops Lead

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
scaling, spot instances, and down-scaling policies to eliminate waste, while providing full visibility into the unit economics of training and serving LLM models. Observability & Incident Response: Establish 24/7 incident response, telemetry, and observability metrics to monitor system performance, model drift, and data pipelines. Data Governance & Security: Enforce ...

AI Engineer

Location
Leeds, England, United Kingdom
context control, and guardrails. Develop retrieval‐augmented workflows to enhance context, reliability, and performance. Perform quality assurance on AI outputs by implementing robust AI observability practices, including monitoring model behaviour, detecting anomalies, and ensuring visibility into AI performance and reliability. Contribute to ongoing research and development, staying current with emerging … chosen when they are safer, simpler, or more cost effective. Ensure AI-enabled solutions consider full total cost of ownership, including token consumption, performance, observability, and ongoing maintenance, with awareness of cost‐efficiency and model‐selection trade‐ Knowledge Sharing, Mentoring and Governance: Mentor and support both technical and non‐technical ...

Senior Software Engineer - Food (Distributed Systems)

Location
Greater London, England, United Kingdom
through clean, maintainable and well-tested code, promoting best practices through code reviews, pair programming, documentation and continuous improvement initiatives. Drive operational excellence and observability by designing effective monitoring and alerting, leveraging tools such as Dynatrace and participating in support activities to ensure critical supply chain and pricing data remains … uses a variety of technologies, including: Backend: Java, Spring, Spring Boot, Micronaut Frontend: React, Next.js, TypeScript, Angular Cloud & Infrastructure: Azure Cloud, Kubernetes Observability: Dynatrace Databases: SQL Server, MongoDB Caching & Performance: Ignite, Redis What's in it for you? Working at M&S means being part of something bigger - helping ...

Lead Azure Engineer

Hiring Organisation
Teksystems
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£400 - £550/day
building CI/CD pipelines with automated quality gates such as unit and integration tests, SAST/DAST where applicable, and dependency scanning. Drive observability and production excellence by implementing logging, metrics, and tracing, creating actionable alerts, defining SLOs and SLIs, and leading incident response and post-incident improvement activities. … Azure DevOps for CI/CD. The role involves close collaboration with product, architecture, UX, QA, and platform teams in a setting that emphasises observability, security, and production excellence. The engineering culture promotes inclusive teamwork, high code quality, and automation across the delivery lifecycle. Location Sheffield, UK Rate/Salary ...

DevOps Laravel Engineer

Location
Telford, England, United Kingdom
escalation point for production incidents alongside our 24/7 out-of-hours support partner, following clear triage and communication processes. Deploy centralised observability and alerting systems (e.g. Grafana, or Sentry etc) to monitor API response times, queue depths, memory thresholds, and database connection pooling in real time, escalating critical … code and architectural reviews for infrastructure changes. Desirable Familiarity with Laravel-specific tooling Experience working in a regulated or security-focused industry. Exposure to observability tooling (for example Grafana, Prometheus, New Relic or Datadog). Experience supporting applications used across multiple countries or time zones. AWS Certifications (e.g. AWS Certified ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. Define and maintain service level indicators (SLIs), service level objectives (SLOs), and error budgets to quantify and manage service reliability. Engineer automation … development lifecycle concepts, and developing applications in a Linux environment. Strong understanding of algorithms, data structures, software design, and distributed systems fundamentals. Experience with observability platforms, including distributed tracing, logging, metrics, and tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience with site reliability engineering practices, relational databases, Hadoop ...

Lead Software Engineer - AI-Native Applications

Hiring Organisation
IFS
Location
London, UK
Employment Type
Full-time
technologies such as Kubernetes, Kafka, Redpanda, PostgreSQL and MongoDB.Experience working with AWS and/or Azure. Strong understanding of CI/CD, automated testing, observability and production operations. Experience designing secure software, authentication and authorisation mechanisms and applying DevSecOps principles. Strong analytical and problem-solving skills with the ability … Experience building Enterprise SaaS or ERP products. Experience working with Model Context Protocol (MCP) or similar AI integration standards. Experience with vector databases, AI observability or AI evaluation frameworks. Experience designing and delivering enterprise-scale AI platforms or AI-powered products. QualificationsA degree in Computer Science, Software Engineering or Information ...

Data Engineer, Vice President

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
datasets, and extensible pipelines that support multiple Company Intelligence products and advanced analytics use cases. Ensure high standards of data quality, governance, lineage, and observability, proactively managing operational and compliance risks across enterprise-grade data products. Partner effectively with product, analytics, and business leaders, developing a deep understanding of strategic … control, testing, CI/CD). Experience building and operating data pipelines using workflow orchestration frameworks (e.g. Apache Airflow), with a focus on reliability, observability, dependency management, and operational resilience. Experience designing and operating cloud-native data platforms (AWS or Azure preferred) and enterprise data warehouses (Snowflake preferred), including performance ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
system integrations Design reliable APIs, data models, caching, synchronization, and recovery paths using PostgreSQL, Valkey, REST, and service integrations across the Cloudbeds platform Own observability, automated testing, release quality, performance, security, and production support for customer‐critical workflows Own changes across the POS application, backend services, infrastructure, and deployment configuration … asynchronous or event‐driven workflows Experience owning cloud‐native delivery and operations, including Terraform, Docker, Kubernetes, Helm, Argo CD, AWS, CI/CD, observability, and production incident response Experience building point‐of‐sale or adjacent real‐time transactional systems where money, inventory, and people meet at the moment of service ...