526 to 550 of 4,012 Observability Jobs

Senior Tech Consultant

Location
Leeds, England, United Kingdom
direction, identifying opportunities to improve a client’s product or technology beyond the immediate brief. Implement and improve CI/CD, cloud infrastructure and observability, primarily using GitHub Actions, AWS and Azure. Promote strong engineering practices around code quality, automated testing, peer review and security, while mentoring and supporting other … building production‐grade AI‐native products, integrating frontier models and designing reliable agentic systems, with a strong understanding of evaluation, context engineering, tool use, observability, security, latency, cost and non‐deterministic behaviour. Strong problem‐solving skills, with the ability to take complex or ambiguous problems and turn them into well ...

Senior DevOps & Cloud SRE Lead (AWS / K8s)

Location
Greater London, England, United Kingdom
production EKS/Kubernetes clusters, ingress controllers, and service meshes Build automated CI/CD pipelines using GitHub Actions, Docker, and Helm Set up observability stack (Prometheus, Grafana, Datadog) and incident management Requirements 4+ years in DevOps, Site Reliability Engineering, or Cloud Architecture Deep expertise in AWS, Kubernetes, Terraform, Docker ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer – Associate - London London · United Kingdom · Associate

Location
Greater London, England, United Kingdom
build, run and continuously improve this business-critical service. The role combines software engineering, systems engineering and production expertise to improve the reliability, scalability, observability, incident response and operational efficiency of the CTL platform. Required Skills/Experience Strong programming ability in one or more modern languages – Java … building maintainable automation beyond simple scripts. Good understanding of networking, messaging, distributed systems, data structures, algorithms and software design fundamentals. Hands-on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes ...

Analytics Services Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
usability, scalability and the developer experienceCollaborating with research, data and engineering teams to accelerate time-to-insight through modern analytics solutionsDriving improvements in automation, observability and resilience across analytics servicesEvaluating and adopting emerging technologies such as AI assistants, data mesh and cloud-native analytics solutionsDefining SLAs, KPIs and monitoring strategies … code, using Terraform or AnsibleDeep understanding of AWS analytics technologies including EMR, MSK, Athena, Redshift, Glue and MWAAExperience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetryStrong problem-solving skills and a systematic approach to diagnosing and resolving issuesHighly desirable skillsExperience with streaming frameworks ...

Applied AI ML Lead - Python & Agentic AI

Location
Auchentibber, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...

Software Engineer

Location
Greater London, England, United Kingdom
Python, C/C++, Go, Typescript, or Java. CI/CD and DevOps: Knowledge of building pipelines, automated testing and deployment strategies. Monitoring and observability: Familiarity with tools like Grafana, Prometheus or other observability platforms. Containerisation and orchestration: Docker, Kubernetes. We are open to a wide range of backgrounds ...

Staff Software Engineer-AI

Location
London, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Senior Platform Engineer

Location
Welwyn Garden City, England, United Kingdom
GitOps workflows to enable safe, fast, and repeatable delivery. Championing DevSecOps principles, embedding security and compliance into the software delivery lifecycle. Establishing and improving observability, monitoring, and incident response practices, including vulnerability management and remediation. Mentoring engineers and contributing to a strong engineering culture through knowledge sharing, documentation, and technical … would be great if you have the following Experience with Helm, Kustomize, and Kubernetes ecosystem tooling. Familiarity with Azure and Azure DevOps. Experience with observability platforms and Kubernetes policy enforcement tools. Proficiency in scripting or programming (e.g. Bash, Python, PowerShell, C#). Experience designing multi‐region or highly available systems. ...

Full Stack Engineer - Must be DV Cleared

Hiring Organisation
VIQU IT Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
platform engineers so models, pipelines and infrastructure land as one product rather than three. Improving the day to day: CI/CD, observability, security hardening, documentation. Sitting in front of customers and partners and turning what they need into something buildable. What You Will Need Python. Serious commercial experience building … secrets management. Putting AI or ML models into production, LLM features included. Infrastructure as code with Terraform, Ansible or similar, plus configuration management. Observability and incident response: metrics, logging, tracing. Message queues, event-driven architectures or data pipelines. Also Welcome Kubernetes. Deploying, running or debugging workloads, including K3s, RKE2 ...

Full Stack Engineer - Contract - Long Term

Hiring Organisation
VIQU IT
Location
London, United Kingdom
Employment Type
Contract
platform engineers so models, pipelines and infrastructure land as one product rather than three. Improving the day to day: CI/CD, observability, security hardening, documentation. Sitting in front of customers and partners and turning what they need into something buildable. What You Will Need Python. Serious commercial experience building … secrets management. Putting AI or ML models into production, LLM features included. Infrastructure as code with Terraform, Ansible or similar, plus configuration management. Observability and incident response: metrics, logging, tracing. Message queues, event-driven architectures or data pipelines. Also Welcome Kubernetes. Deploying, running or debugging workloads, including K3s, RKE2 ...

Lead GCP Engineer

Location
City Of London, England, United Kingdom
practices across the platform. Drive adoption and maturity of CI/CD pipelines, Infrastructure as Code and automated deployment processes. Implement and improve platform observability, monitoring, alerting and operational support capabilities. Develop and maintain reusable automation patterns and engineering standards. Leverage Terraform and Infrastructure as Code principles to ensure repeatable … Infrastructure as Code expertise. Strong experience with CI/CD tooling and deployment automation. Experience building and operating resilient cloud platforms. Strong understanding of observability, monitoring, incident management and operational excellence. Data Engineering & Integration Experience designing secure data integration solutions. Knowledge of APIs, databases, data pipelines and cloud‐based data ...

Senior Platform Engineering Manager at Prolific

Location
United Kingdom
operational excellence, and innovation. Champion SRE Culture: Own availability and embed SRE principles across the organization, including defining SLOs, SLAs, error budgets, and enhancing observability and incident remediation. Platform & Developer Experience: Own the developer-facing platform — golden paths, self-service infrastructure, and internal tooling — so teams can provision, deploy … Infrastructure-as-Code (Terraform/Terragrunt, Crossplane), with GitOps workflows using tools like ArgoCD. Reliability & Architecture: Solid understanding of complex infrastructure and application architecture, observability principles, and incident management. Stability & Velocity: Experience balancing the need for platform stability and reliability with the goal of increasing developer productivity and velocity. Bridge ...

AI Engineer Python

Location
Bracknell, England, United Kingdom
Terraform. Establish strong testing, code quality, documentation-as-code and secure software development practices. Improve application performance, scalability and reliability through profiling, benchmarking, observability and architectural improvements. Implement monitoring, alerting and operational standards for platform services. Conduct thorough code reviews, mentor engineers and raise engineering standards across the team. Partner … including automated testing, code reviews, CI/CD and release engineering. Practical experience with Docker, Kubernetes and Infrastructure as Code, ideally Terraform. Experience with observability, monitoring, alerting, performance profiling and benchmarking. Solid understanding of secure coding practices, vulnerability management and security considerations in software delivery. Applied understanding of Generative ...

Principal Software Architect (UK)

Location
Greater London, England, United Kingdom
containerization frameworks (Kubernetes, Docker, Helm). Experience with infrastructure‐as‐code, configuration management (Ansible, Terraform), and network automation protocols (NetConf, YANG). Distributed Systems & Observability: Deep familiarity with modern distributed systems tooling, including RESTful APIs, gRPC, message brokers (Kafka, NATS), and data serialization formats (JSON, YAML). Demonstrated expertise … system observability, telemetry, and monitoring stacks (Prometheus, Grafana). High‐Performance Networking: Knowledge of high‐performance packet processing techniques (e.g., DPDK, eBPF) and low‐latency systems tuning. Proficiency in network protocols and architecture, including TCP/IP, UDP, routing, and network performance optimization. Preferred Qualifications Master's degree in computer ...

Senior Manager- Software Engineering

Location
Greater London, England, United Kingdom
microservices, including REST and event‐driven patterns, and enforce best practices for versioning, contracts, and backward compatibility. Advance operational excellence by defining SLOs, improving observability and alerting, hardening on‐call procedures and runbooks, and leading incident response and post‐mortems. Solve complex distributed system challenges (such as throughput, latency, consistency … with microservices, API design, and event‐driven systems, as well as experience with containerization and orchestration (Docker, Kubernetes). Strong operational mindset: expert in observability, monitoring, incident response, performance engineering, and adherence to security best practices. Familiarity with CI/CD pipelines, automated testing strategies (unit, integration, e2e), and modern ...

Staff Software Engineer - AI

Location
Greater London, England, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event‐driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands‐on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Senior Data Platform Engineer (Python & Databricks)

Location
City of Edinburgh, Scotland, United Kingdom
ingesting, transforming, and delivering financial data. Design data models that support enterprise reporting, analytics, and downstream integrations. Improve the performance, reliability, maintainability, and observability of existing data workflows. Troubleshoot complex issues across data-processing and application layers. Python and API Development Design, build, and maintain production-grade applications and services … services or other data-intensive industries. Experience with cloud-based data architectures. Familiarity with infrastructure-as-code tools such as Terraform. Experience improving the observability and operational reliability of data pipelines. Familiarity with modern data governance, access-control, and data-quality practices. Experience working in an enterprise environment with strict ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US Our client ...

Safety Engineer - Free Tier Abuse

Location
Greater London, England, United Kingdom
models into production systems Architect robust APIs, data pipelines, and service architectures supporting real‐time and batch moderation workflows Implement comprehensive monitoring, alerting, and observability systems; establish SLIs, SLOs, and performance benchmarks Partner with ML engineers to translate research models into production‐ready systems and integrate them across our product … pipelines, and Python expertise (asynchronous Python, backend frameworks) Infrastructure & DevOps proficiency: cloud platforms (AWS/GCP), containerization (Docker/K8s), CI/CD pipelines Observability mindset with experience in monitoring tools (Prometheus, Grafana) and building observable systems Track record of taking products or systems from 0→1 with measurable impact ...

Lead DevOps Engineer

Location
Manchester, England, United Kingdom
lifecycle. Automation and reliability engineering Proficient in Python with solid scripting in PowerShell and Bash to drive automation and operational efficiency, combined with strong observability practice (for example Azure Monitor, Sentinel, Application Insights and Log Analytics) using SLIs, SLOs and distributed tracing to sustain and improve system reliability. Autonomy … Code and CI/CD. This is a hands‐on leadership role where you'll influence key technical decisions, champion DevSecOps, resilience and observability, and work closely with engineering teams to create high‐quality services that make a real impact. You'll use technologies including Azure, Azure DevOps and Terraform ...

Senior Service Reliability Engineer

Location
Knutsford, England, United Kingdom
reliability and customer experience expectations. Working across Engineering, Infrastructure, Security, and Product teams, you will identify opportunities to automate manual processes, improve monitoring and observability, strengthen resilience, and reduce operational risk. You will also contribute to service reviews, support governance and control activities, and provide clear communication to stakeholders during … reliability for critical business services. Cloud, Platform & Engineering Expertise –Strong hands-on knowledge of AWS cloud technologies, microservices, APIs, containerized platforms (OpenShift/Kubernetes), observability tooling, automation, CI/CD, Infrastructure as Code, and modern platform engineering practices. Senior Stakeholder & Operational Leadership–Proven ability to lead critical incidents, drive root ...

SRE Engineer

Location
Greater London, England, United Kingdom
engineering standards Driving infrastructure‐as‐code and automation across Azure and on‐prem Improving image bakery pipelines for secure, repeatable server builds Embedding observability using metrics, logs, traces, and effective alerting Ensuring all practices align with ISO 27001 and internal security frameworks Managing automated patching, vulnerability remediation and configuration compliance … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions) Experience with monitoring and observability stacks Solid understanding of OS fundamentals (Windows/Linux), security, networking Background in scripting or software development (PowerShell, Python, Go) Experience with containers and orchestration (Docker ...

Senior Service Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, United Kingdom
Salary
£ 70 K
meet reliability and customer experience expectations.Working across Engineering, Infrastructure, Security, and Product teams, you will identify opportunities to automate manual processes, improve monitoring and observability, strengthen resilience, and reduce operational risk. You will also contribute to service reviews, support governance and control activities, and provide clear communication to stakeholders during … service reliability for critical business services.Cloud, Platform & Engineering Expertise –Strong hands-on knowledge of AWS cloud technologies, microservices, APIs, containerized platforms (OpenShift/Kubernetes), observability tooling, automation, CI/CD, Infrastructure as Code, and modern platform engineering practices.Senior Stakeholder & Operational Leadership–Proven ability to lead critical incidents, drive root cause ...