451 to 475 of 4,967 Permanent Observability Jobs

Staff Software Engineer - AI

Location
London, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Lead .Net Software Engineer

Location
Cheltenham, England, United Kingdom
technical risks, bottlenecks, dependencies, and opportunities for improvement Contribute to AWS‐based architecture and engineering practices, including environments using Lambda, ECS, and EC2 Improve observability, reliability, performance, security, and operational readiness across the platform Contribute to Jenkins pipelines and CI/CD practices to improve consistency and delivery efficiency Help … with DevOps practices and infrastructure‐aware development Experience with messaging, event streaming, or related event‐driven technologies Experience with Blazor and MudBlazor Familiarity with observability and monitoring tooling in distributed systems Experience defining governance, controls, or operating models for AI agents or AI‐enabled internal tools Experience working ...

Data Operations DevOps Engineer

Location
United Kingdom
/CD and GitOps workflows to support deployment and improve the team’s release management processes. Drive operational excellence and security practices: improve observability, documentation and security practices across SRCNet services. Experience working in a professional DevOps engineering capacity for at least two years Postgraduate qualification in Astronomy, Physics, Computer … lifecycle management of containerised services. Demonstrable ability to diagnose and resolve complex operational issues across application, infrastructure and networking layers, supported by experience with observability and operational tooling (e.g. Prometheus, Grafana, Elastic stack), and the ability to manage workloads and targets in a dynamic and collaborative environment. Experience of building ...

Lead AI Engineer

Location
Sheffield, England, United Kingdom
Agentic AI integrations within enterprise environments. Collaborate with architects and engineering teams to ensure scalable, secure, and maintainable solutions. Define standards for AI observability, governance, security, and performance. Mentor engineers and provide technical leadership across AI development initiatives. Contribute hands‐on to solution design, development, code reviews, and production deployment. ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
Lead code reviews, mentor engineers, and ensure high engineering standards. Contribute to platform SDKs, internal libraries, and backend frameworks used across charters. Performance, Reliability & Observability Implement tracing, structured logging, metrics, dashboards, and alerting. Optimize services for latency, concurrency, throughput, and cost efficiency. Ensure system reliability through automated testing, load testing … Kubernetes, and cloud platforms (AWS or GCP). Strong CI/CD experience using GitHub Actions, Jenkins, Argo, or similar tools. Deep knowledge of observability practices — logs, metrics, traces, performance analysis. Proven ability to write clean, testable, well‐structured production code. Qualifications – Nice to Have Experience with streaming or queuing ...

Senior Software Development Engineer

Location
Reading, England, United Kingdom
agile delivery environments (e.g. Scrum) Desirable Experience with cloud platforms (e.g. Azure) Experience with CI/CD pipelines and release processes Experience with observability tools (monitoring, logging, alerting) Experience in technical leadership or mentoring roles Experience working in safety- or regulation-driven environments Technology Stack .NET (latest versions) Kubernetes & Docker … Azure (SQL, CosmosDB, cloud services) PostgreSQL TypeScript/modern web frameworks (e.g. Vue.js) Observability tooling (e.g. Azure Monitor, Prometheus, Grafana) Azure DevOps/CI-CD pipelines Holidays: 25 days per annum + 8 days bank holidays (options to buy/sell days) 37.5 hour working week Pension – 4% employee ...

ML Ops Lead

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
scaling, spot instances, and down-scaling policies to eliminate waste, while providing full visibility into the unit economics of training and serving LLM models. Observability & Incident Response: Establish 24/7 incident response, telemetry, and observability metrics to monitor system performance, model drift, and data pipelines. Data Governance & Security: Enforce ...

AI Engineer

Location
Leeds, England, United Kingdom
context control, and guardrails. Develop retrieval‐augmented workflows to enhance context, reliability, and performance. Perform quality assurance on AI outputs by implementing robust AI observability practices, including monitoring model behaviour, detecting anomalies, and ensuring visibility into AI performance and reliability. Contribute to ongoing research and development, staying current with emerging … chosen when they are safer, simpler, or more cost effective. Ensure AI-enabled solutions consider full total cost of ownership, including token consumption, performance, observability, and ongoing maintenance, with awareness of cost‐efficiency and model‐selection trade‐ Knowledge Sharing, Mentoring and Governance: Mentor and support both technical and non‐technical ...

Senior Software Engineer - Food (Distributed Systems)

Location
Greater London, England, United Kingdom
through clean, maintainable and well-tested code, promoting best practices through code reviews, pair programming, documentation and continuous improvement initiatives. Drive operational excellence and observability by designing effective monitoring and alerting, leveraging tools such as Dynatrace and participating in support activities to ensure critical supply chain and pricing data remains … uses a variety of technologies, including: Backend: Java, Spring, Spring Boot, Micronaut Frontend: React, Next.js, TypeScript, Angular Cloud & Infrastructure: Azure Cloud, Kubernetes Observability: Dynatrace Databases: SQL Server, MongoDB Caching & Performance: Ignite, Redis What's in it for you? Working at M&S means being part of something bigger - helping ...

DevOps Laravel Engineer

Location
Telford, England, United Kingdom
escalation point for production incidents alongside our 24/7 out-of-hours support partner, following clear triage and communication processes. Deploy centralised observability and alerting systems (e.g. Grafana, or Sentry etc) to monitor API response times, queue depths, memory thresholds, and database connection pooling in real time, escalating critical … code and architectural reviews for infrastructure changes. Desirable Familiarity with Laravel-specific tooling Experience working in a regulated or security-focused industry. Exposure to observability tooling (for example Grafana, Prometheus, New Relic or Datadog). Experience supporting applications used across multiple countries or time zones. AWS Certifications (e.g. AWS Certified ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. Define and maintain service level indicators (SLIs), service level objectives (SLOs), and error budgets to quantify and manage service reliability. Engineer automation … development lifecycle concepts, and developing applications in a Linux environment. Strong understanding of algorithms, data structures, software design, and distributed systems fundamentals. Experience with observability platforms, including distributed tracing, logging, metrics, and tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience with site reliability engineering practices, relational databases, Hadoop ...

Lead Software Engineer - AI-Native Applications

Hiring Organisation
IFS
Location
London, UK
Employment Type
Full-time
technologies such as Kubernetes, Kafka, Redpanda, PostgreSQL and MongoDB.Experience working with AWS and/or Azure. Strong understanding of CI/CD, automated testing, observability and production operations. Experience designing secure software, authentication and authorisation mechanisms and applying DevSecOps principles. Strong analytical and problem-solving skills with the ability … Experience building Enterprise SaaS or ERP products. Experience working with Model Context Protocol (MCP) or similar AI integration standards. Experience with vector databases, AI observability or AI evaluation frameworks. Experience designing and delivering enterprise-scale AI platforms or AI-powered products. QualificationsA degree in Computer Science, Software Engineering or Information ...

Data Engineer, Vice President

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
datasets, and extensible pipelines that support multiple Company Intelligence products and advanced analytics use cases. Ensure high standards of data quality, governance, lineage, and observability, proactively managing operational and compliance risks across enterprise-grade data products. Partner effectively with product, analytics, and business leaders, developing a deep understanding of strategic … control, testing, CI/CD). Experience building and operating data pipelines using workflow orchestration frameworks (e.g. Apache Airflow), with a focus on reliability, observability, dependency management, and operational resilience. Experience designing and operating cloud-native data platforms (AWS or Azure preferred) and enterprise data warehouses (Snowflake preferred), including performance ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
system integrations Design reliable APIs, data models, caching, synchronization, and recovery paths using PostgreSQL, Valkey, REST, and service integrations across the Cloudbeds platform Own observability, automated testing, release quality, performance, security, and production support for customer‐critical workflows Own changes across the POS application, backend services, infrastructure, and deployment configuration … asynchronous or event‐driven workflows Experience owning cloud‐native delivery and operations, including Terraform, Docker, Kubernetes, Helm, Argo CD, AWS, CI/CD, observability, and production incident response Experience building point‐of‐sale or adjacent real‐time transactional systems where money, inventory, and people meet at the moment of service ...

Production AI Engineer - Vice President

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
engineering techniques to integrate large language models (LLMs) into operational tooling, incident response pipelines, and developer productivity platforms. Leads the development of AI-native observability solutions — leveraging intelligent agents to detect anomalies, predict failures, and automate remediation before issues impact end users. Writes clean, well-tested, and well-documented code … Kanban).Operational experience of using middleware technologies (MQ, Apache Kafka, etc.) to run services at scale is desirable. Strong experience with end-to-end observability stacks (Datadog, AppDynamics, Dynatrace, etc.) is desirable. Degree in Computer Science, Mathematics, Physics, or a related technical subject is desirable. Experience of senior stakeholder management. ...

Senior AI Architect| London

Hiring Organisation
Infosys Technologies
Location
London, United Kingdom
Salary
£ 70 K
hallucination in production.• LLMOps, Evaluation & Responsible AI: Experience operationalizing LLM and agentic systems at scale—evaluation harnesses and metrics for quality, groundedness, and safety; observability, tracing, and monitoring (e.g., LangSmith, LangFuse); guardrails and red-teaming; and continuous optimization of accuracy, cost, and latency. Understanding of AI governance, security, privacy, bias … standards for agentic AI—agent orchestration, MCP-based tool/data integration, shared skills and connectors, memory and state management, guardrails, human oversight, and observability—to enable safe, reliable, and scalable production deployment across teams.• Solution Implementation: Collaborate with data scientists and engineers to implement Generative AI solutions, ensuring ...

Principal Architect

Hiring Organisation
Virtusa
Location
London, UK
Employment Type
Full-time
Code Proficiency in one or more technologies Java, .NET, Python, Go, Node.js Strong understanding of Security architecture Data architecture High availability & disaster recovery Observability & monitoring Leadership Skills Strong stakeholder management and communication skills. Ability to influence cross-functional teams and executive leadership. Experience leading large engineering transformations. Strategic thinking with ...

DevOps Engineer - UK

Location
Greater London, England, United Kingdom
/CD pipelines Strong AWS architecture knowledge Experience with large-scale distributed systems Hands-on experience with containerization & orchestration Experience with monitoring and observability tools Solid understanding of infrastructure security British Citizen or right to work in UK (no visa sponsorship available) Strong communication & problem-solving skills Experience working with ...

Site Reliability Engineer

Hiring Organisation
E-Solutions IT Services UK Ltd
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£280.00 - £300.00 per day
similar tools. • Troubleshoot production issues, conduct root cause analysis, and implement preventive measures. • Automate operational tasks using Python or other scripting languages. • Contribute to observability and monitoring improvements using modern tools and best practices. • Participate in on-call rotations and incident response processes. Required Skills and Experience: • 5–9 years ...

Technology Integration Specialist

Hiring Organisation
Ncounter
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£130,000 - £150,000 per annum
/CD knowledge across tooling such as Jenkins, GitLab CI or ArgoCD Strong Linux knowledge and experience working across AWS, GCP or Azure Observability experience with technologies such as Prometheus or Grafana Scripting or software development capability, ideally Python or Go Experience leading complex technical projects involving multiple engineering teams ...

cloud engineer in cloud platforms

Location
Greater London, England, United Kingdom
cloud platforms and services with Architects and engineering teams Support Azure networking, identity, security, compute, storage and platform services Implement monitoring, logging, alerting and observability across cloud environments Troubleshoot complex platform, deployment and infrastructure issues Improve reliability, scalability and operational performance through automation and engineering best practice Support containerised workloads ...

Java Software Engineer - VP

Hiring Organisation
Henderson Scott
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Data: Proficiency with Docker, Kubernetes, and relational databases (MS SQL, Sybase). Delivery Automation & Monitoring: Strong focus on automated testing, automated release pipelines, and observability tools (Grafana, Prometheus). Mindset: Delivery-focused problem solver with a hands-on approach and strong stakeholder communication skills. Desirable Experience Background in Equity Swaps ...

C# Developer

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
banking).Cloud experience (Azure or AWS), containers (Docker) and orchestration (Kubernetes).Messaging/event streaming (Kafka, RabbitMQ, Azure Service Bus) and distributed systems patterns. Observability tooling (logging, metrics and tracing) and production support experience. What You'll Bring A pragmatic approach to delivery with attention to quality and detail. Ability ...

Junior Data Engineer

Location
Greater London, England, United Kingdom
primary), some GCP* **Warehouse & Storage:** Snowflake, S3/Parquet* **Data & ETL:** dbt, Fivetran* **Platform & Infra:** Kubernetes, Kafka, RabbitMQ, Argo, GitHub Actions, HashiCorp Vault* **Observability:** Datadog, Grafana* **Dashboarding:** Preset* **Other**: Claude### **What we’re looking for**Strong fundamentals and the ability to apply them pragmatically:* Solid programming ability (Python or similar ...

Senior Software Engineer - Backend & Distributed Systems Engineer

Location
Greater London, England, United Kingdom
operations team. Automate Delivery and Infrastructure: Manage infrastructure using Pulumi or Terraform and improve automated testing and deployment through GitHub Actions. Improve Reliability and Observability: Build effective monitoring, logging, tracing and alerting across our data, service, model and agent infrastructure. What We’re Looking For 7+ years of professional experience ...