1,126 to 1,150 of 4,987 Permanent Observability Jobs

Director of AI

Location
Tendring, England, United Kingdom
driven development practices, including test datasets, error analysis, deterministic checks, model-based evaluation, and business outcome measurement. Ensure solutions meet production requirements for reliability, observability, scalability, latency, maintainability, security, and cost. Guide practices for CI/CD, model and prompt versioning, monitoring, tracing, regression testing, and ongoing optimization. Provide hands … cost, security, and user experience. Experience with cloud-native architecture and production deployment on GCP, AWS, or Azure. Familiarity with containers, CI/CD, observability, and MLOps or LLMOps practices. Demonstrated success in a consulting, professional-services, or complex client-facing environment. Ability to move comfortably between executive conversations, product ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
across our systems. Improve developer experience through automation, self-serve tooling and infrastructure-as-code practices that increase engineering velocity. Own and evolve our observability, incident response and reliability practices to minimise downtime and improve system performance. Partner closely with product and engineering teams to support new services, migrations … production environments. Strong experience operating and scaling infrastructure components such as Kafka, Redis and managed relational databases like RDS. Strong understanding of networking, security, observability and distributed systems fundamentals. Experience improving CI/CD pipelines and deployment automation in fast-moving engineering teams. Comfortable debugging production issues and participating ...

Lead Software Engineer - Application Owner

Location
Bournemouth, England, United Kingdom
after-action reviews, and closure of follow-up actions. Define and continuously improve production readiness standards, including release safety and rollback strategy, dependency awareness, observability requirements, and operational runbooks. Contribute hands‐on to design and delivery, including system design, code reviews, automation, and complex troubleshooting, with secure‐by‐design … with strong operational accountability, including controls, resiliency and recovery, and remediation tracking. Strong system design fundamentals and cloud‐native operational patterns, including scalability, reliability, observability, and dependency management. Hands‐on experience with Go‐based services and modern CI/CD practices. Experience operating workloads on AWS and Kubernetes ...

Senior Cloud Engineer

Location
Greater London, England, United Kingdom
hours on-call rotas as required. Support the full lifecycle of application workloads, including maintenance, patching, backups, upgrades, and controlled decommissioning. Reliability, SRE & Observability Drive improvements in reliability, resilience, performance, and efficiency through strong SRE practices. Own and enhance observability and monitoring, ensuring meaningful alerting, clear operational dashboards, and rapid … with an Enterprise wide leadership and team structure. Skills & Experience – Essential Extensive experience designing, building, and operating distributed AWS environments, including networking, IAM, security, observability, and monitoring. Strong hands-on experience with cloud security, compliance, and governance in enterprise or regulated environments. Proficiency in infrastructure-as-code (e.g. Terraform ...

Senior Data Platform Engineer

Location
Greater London, England, United Kingdom
policy-as-code, and action logging. Build platform-, tool-, and agent-level kill switches, alongside dry-run/safe-testing modes for agent workflows. Observability & FinOps Implement AI system observability, including prompt logging, output monitoring, quality scoring, and drift detection. Establish operational monitoring for gateway usage, latency, error rates … plain English. And to really stand out from the crowd Experience with AI gateways, LLM proxies, or model-routing patterns. Exposure to AI observability tools, vector databases (Databricks AI Search), RAG pipelines, or non-deterministic system monitoring. Hands-on experience implementing policy-as-code, action logging, or AI kill switches. ...

Senior Software Engineer- Python

Location
Manchester, England, United Kingdom
specialists (traders, data scientists, platform) without needing to be an expert in their field on day one Raising the bar across software (typing, testing, observability, tooling, and thoughtful refactors), tooling, and ways of working What you'll do Design, build and operate the Python backend services behind Asset Backed Trading … modelling, optimisation. Event driven or serverless architectures at meaningful scale Infrastructure as Code (CDK, CloudFormation, or Terraform) Experience with Pydantic, strict typing, and mypy Observability in production (CloudWatch, Grafana) Exposure to energy, trading, optimisation, or other constraint heavy domains Working in a monorepo with shared libraries and multiple deployable services ...

Core AI Engineer

Location
Greater London, England, United Kingdom
running centralised MCP servers, ensuring secure, reliable access to enterprise tools and data Owning platform reliability, performance and scalability across Kubernetes‐based infrastructure, including observability, capacity planning and incident response Building self‐service tooling and APIs to enable teams to provision and consume AI infrastructure independently Integrating platform services with … services Familiarity with sandboxing and workload isolation technologies Experience in quantitative finance or low‐latency systems AWS experience particularly in hybrid environments Experience with observability tooling such as Prometheus, Grafana or OpenTelemetry Contributions to open‐source projects in relevant domains Why join us? Highly competitive compensation plus annual discretionary bonus ...

AI Engineers London

Location
Greater London, England, United Kingdom
Experience implementing, maintaining and evaluating retrieval systems (vector/graph databases, ingestion pipelines, chunking strategies, retrieval techniques such as HyDE) Implement feedback loops and observability to continuously improve system performance Craft effective prompts and optimize for latency, cost, and quality across different model providers and configurations Required Skills and Experience … fine‐tuning is preferable to prompt engineering or RAG Experience with real‐time streaming, multimodal models, or search technologies like Elasticsearch Familiarity with model observability tools (LangSmith, Weights & Biases) and cost optimization strategies Experience in specialized verticals (financial services, energy, healthcare, legal, retail) with understanding of compliance, security, and responsible ...

Senior Software Engineer, ML Infrastructure

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Software Engineer III - Full Stack, Global Banking Tech

Location
Glasgow, Scotland, United Kingdom
analysis, analyzing diverse datasets, logs, and telemetry to identify patterns, build visualizations and reporting, and reduce repeat incidents via preventative controls, automation, and enhanced observability Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit … agentic frameworks such as Google ADK or LangChain Familiarity with monitoring, tracing, and troubleshooting tools such as log aggregation platforms, API testing tools, and observability dashboards #J-18808-Ljbffr ...

DevOps Engineer - SC Cleared - Hybrid - Inside IR35

Location
City Of London, England, United Kingdom
Experience of working within a production environment, with strong on-prem Kubernetes/OpenShift deployments IAC using Terraform with CI/CD pipelines (Jenkins) Observability tools to include design and operate end to end logging, metrics, tracing, dashboards to include alerting systems using ELK, Splunk, Grafana Cloud platforms to include ...

Senior Software Engineer I

Location
City Of London, England, United Kingdom
maintain generative AI services and reusable components using mostly Python and a little bit Java. Defining and promote best practices in engineering, including scalability, observability, testing, and CI/CD. Contributing to system designs spanning multiple services and modules, aligning with architectural best practices. Collaborating with product, platform, and research … work collaboratively across functions in an Agile or Kanban environment. Nice to Have Experience operationalizing LLMs or building internal AI platforms. Familiarity with observability practices (metrics, logging, alerts). Exposure to knowledge graphs or semantic search systems. U.S. National Base Pay Range U.S. National Base Pay Range ...

Senior Machine Learning Engineer (ML Platform)

Location
Greater London, England, United Kingdom
infrastructure. Improve automation across the ML lifecycle, including model packaging, deployment, versioning, monitoring, and release processes. Maintain and improve the reliability and observability of the ML platform and model-serving services, including logging, metrics, and alerting. Participate in the team's on-call rotation, investigate production incidents, and contribute … experience provisioning and managing cloud infrastructure with Terraform. Experience with CI/CD pipelines, automated testing, and Git-based development workflows. Familiarity with observability practices, including logging, metrics, alerting, and production troubleshooting. Strong grasp of software engineering principles and best practices. Experience contributing to or leading data warehouse architecture ...

Innovation Software Developer

Hiring Organisation
EOS IT Solutions
Location
Armagh, United Kingdom
Salary
£ 70 K
real-time and high-volume sensor streams.Design Modern APIs and ServicesDevelop secure, scalable REST and gRPC services using FastAPI, Flask, and Node.js.Implement authentication, authorization, observability, testing, and performance monitoring.Ensure reliability, maintainability, and best-practice API design.Deliver Digital Twin ExperiencesIntegrate real-time operational data with advanced 3D and 4D visualizations.Work with … modern machine learning frameworks including PyTorch and Scikit-learn.Support DevOps and Delivery ExcellenceDeploy containerized applications using Docker and Kubernetes.Implement CI/CD pipelines and observability tooling.Produce high-quality technical documentation and operational runbooks.What We're Looking ForEssential Skills & Experience5+ years of software development experience in complex environments.Strong expertise in Python ...

Vice President, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
engineering teams; set clear objectives, coach talent, and foster succession planning. Own end-to-end delivery for critical software: requirements, architecture, implementation, testing, deployment, observability, and reliability. Raise engineering excellence and resilience: best practices and automation across code, testing, microservices/APIs, performance, and infrastructure; secure-by-design with threat … scalable, observable, testable systems; strong API design. Strong DevOps practices: CI/CD (e.g., GitLab), automated testing (JUnit/Spock), code reviews, telemetry/observability (Splunk, AppDynamics), containers (Docker), and cloud. Hands-on AI development using modern tools and IDEs (e.g., Windsurf) and experience integrating AI into product workflows. Excellent ...

Innovation Software Developer

Hiring Organisation
EOS IT Solutions
Location
Armagh, Co. Armagh, UK
Employment Type
Full-time
high-volume sensor streams. Design Modern APIs and ServicesDevelop secure, scalable REST and gRPC services using FastAPI, Flask, and Node.js. Implement authentication, authorization, observability, testing, and performance monitoring. Ensure reliability, maintainability, and best-practice API design. Deliver Digital Twin ExperiencesIntegrate real-time operational data with advanced 3D and 4D visualizations. … learning frameworks including PyTorch and Scikit-learn. Support DevOps and Delivery ExcellenceDeploy containerized applications using Docker and Kubernetes. Implement CI/CD pipelines and observability tooling. Produce high-quality technical documentation and operational runbooks. What We're Looking ForEssential Skills & Experience5+ years of software development experience in complex environments. Strong ...

Vice President, Full-Stack Engineer Opportunities

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
engineering teams; set clear objectives, coach talent, and foster succession planning.. Own end-to-end delivery for critical software: requirements, architecture, implementation, testing, deployment, observability, and reliability. Raise engineering excellence and resilience: best practices and automation across code, testing, microservices/APIs, performance, and infrastructure; secure-by-design with threat … scalable, observable, testable systems; strong API design. Strong DevOps practices: CI/CD (e.g., GitLab), automated testing (JUnit/Spock), code reviews, telemetry/observability (Splunk, AppDynamics), containers (Docker), and cloud Hands-on AI development using modern tools and IDEs (e.g., Windsurf) and experience integrating AI into product workflows Excellent ...

Senior UI Architect , Vice President

Location
Greater London, England, United Kingdom
centric digital experiences that align with business objectives. Engineering Excellence Guide development teams through architectural implementation and technical challenges. Establish CI/CD, testing, observability, and quality engineering practices for UI applications. Lead performance optimisation initiatives, including Core Web Vitals improvements. Ensure maintainability, scalability, and resilience of front-end solutions. … REST APIs GraphQL OAuth2/OpenID Connect Enterprise SSO solutions Quality & Performance Web accessibility (WCAG) Core Web Vitals Automated testing frameworks Performance monitoring and observability Preferred Experience Large-scale financial services, banking, asset management, or regulated industry experience. Experience with Adobe Experience Manager (AEM) Headless CMS. Experience delivering enterprise digital ...

Data Lead

Hiring Organisation
UBS
Location
London, UK
Employment Type
Full-time
orchestration (Airflow, Prefect, or similar) and reliability patterns such as retries, backfills, and exactly-once semantics. Deep experience with Azure, including AKS, Helm, EventHub, observability tools (e.g. Log Analytics) and infrastructure-as-code (e.g. Terraform). Advanced experience operating Delta tables at scale, covering transaction semantics, schema evolution, partitioning … Marimo. Batch and streaming processing proficiency, including late data handling, watermarking, backfills, and explicit throughput vs. latency trade-offs. Production data reliability and observability mindset, covering lineage, data quality checks, SLAs, CI/CD for data pipelines, and close collaboration with quants, risk managers, and traders to deliver trustworthy data ...

Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, United Kingdom
Salary
£ 70 K
platform reliability.Proficiency in Linux/Unix environments and Bash/Shell scripting.Good understanding of CI/CD principles and automated deployment pipelines.Experience with monitoring, observability, alerting, and operational telemetry platforms.Good troubleshooting and problem-solving skills with the ability to perform root cause analysis and incident remediation.Experience working in Agile teams … availability.Some other highly valued skills may help:Experience administering enterprise database platforms such as Oracle.Knowledge of Oracle Database architecture, PL/SQL.Experience with enterprise observability platforms such as Splunk, Elastic, Observe, Grafana, Prometheus.Familiarity with Infrastructure as Code and platform engineering practices.You may be assessed on the key critical skills relevant ...

Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, UK
Employment Type
Full-time
Linux/Unix environments and Bash/Shell scripting. Good understanding of CI/CD principles and automated deployment pipelines. Experience with monitoring, observability, alerting, and operational telemetry platforms. Good troubleshooting and problem-solving skills with the ability to perform root cause analysis and incident remediation. Experience working in Agile … other highly valued skills may help: Experience administering enterprise database platforms such as Oracle. Knowledge of Oracle Database architecture, PL/SQL.Experience with enterprise observability platforms such as Splunk, Elastic, Observe, Grafana, Prometheus. Familiarity with Infrastructure as Code and platform engineering practices. You may be assessed on the key critical ...

Technical Lead (AWS Serverless & Microservices)

Location
Greater London, England, United Kingdom
through automation, testing practices and governance frameworks Promoting secure-by-design principles and embedding security best practice across teams and platforms Establishing and evolving observability capabilities including logging, metrics, tracing and alerting Championing metrics-led engineering to improve delivery performance, reliability and customer outcomes Providing technical guidance on complex engineering … Expertise in quality engineering, automated testing strategies and CI/CD practices Strong understanding of secure software development principles and security frameworks Experience implementing observability strategies using logging, metrics, tracing and alerting tools Proven ability to influence stakeholders and communicate complex technical concepts clearly Experience coaching, mentoring and developing high ...

Principal DevOps Engineer

Location
Greater London, England, United Kingdom
templates, shared steps, release promotion, rollback) Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable) Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards) Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning … Ansible (architecture, reusable components, secure operations) Strong deployment/release engineering experience with Octopus Deploy and GitHub (release governance, environment promotion, rollback) Monitoring/observability expertise with VictoriaMetrics and/or Prometheus (alerting strategy, metrics design, operational readiness) Production experience running Redis , RabbitMQ , Nginx (HA, tuning, troubleshooting) Strong understanding ...

Credit Principal Algo Data Lead

Hiring Organisation
UBS
Location
London, UK
Employment Type
Full-time
orchestration (Airflow, Prefect, or similar) and reliability patterns such as retries, backfills, and exactly-once semantics. Deep experience with Azure, including AKS, Helm, EventHub, observability tools (e.g. Log Analytics) and infrastructure-as-code (e.g. Terraform). Advanced experience operating Delta tables at scale, covering transaction semantics, schema evolution, partitioning … Marimo. Batch and streaming processing proficiency, including late data handling, watermarking, backfills, and explicit throughput vs. latency trade-offs. Production data reliability and observability mindset, covering lineage, data quality checks, SLAs, CI/CD for data pipelines, and close collaboration with quants, risk managers, and traders to deliver trustworthy data ...

Software Engineer, Enterprise

Location
Greater London, England, United Kingdom
applications can operate at enterprise scale across hybrid and multi-cloud environments. Manage and evolve cloud infrastructure (AWS, Azure, or GCP), driving automation, observability, and security for large-scale AI deployments. Collaborate with ML and product teams to bring cutting-edge GenAI models into production through efficient APIs, model serving … experience with GenAI applications, model integration, or AI agent systems—understanding how to deploy, evaluate, and scale AI workloads in production. Strong understanding of observability, CI/CD , and security best practices for running services in enterprise or multi-tenant environments. Ability to balance rapid iteration with production-grade quality ...