501 to 525 of 4,012 Observability Jobs

Staff Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
role for someone who can design and build scalable systems, not just manage cloud infrastructure. You'll work across infrastructure, automation, CI/CD, observability and production reliability, while writing high-quality software in Go and helping set the technical standard for how engineering teams build and operate services …/CD pipelines, making builds, tests and deployments faster and safer Build automation and internal tooling that removes manual toil for engineering teams Improve observability across metrics, logging, tracing and production monitoring Define and manage SLOs across latency, availability and error rates Lead incident response, resolve complex production issues ...

Senior Forward Deployment Engineer

Location
Slough, England, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

AI Native DevOps Platform Engineer

Location
Greater London, England, United Kingdom
using AI‐first engineering practices. Working closely with Product Engineering Teams and Technical Leadership, you will build the cloud platforms, infrastructure, deployment pipelines, automation, observability frameworks, and engineering tooling that enable the rapid delivery of both AI‐powered and traditional cloud‐native applications. We are building an AI‐native engineering … platform tooling. Drive engineering productivity through AI, automation, self‐service capabilities and platform standardisation. Improve release processes and operational excellence across teams. Reliability & Observability Implement monitoring, logging, tracing, and alerting solutions. Establish platform SRE principles and operational standards. Proactively identify and resolve reliability, security, and performance issues. Lead incident response ...

Senior Golang Engineer - Contract

Hiring Organisation
Spencer Rose Ltd
Location
Bristol, Somerset, United Kingdom
Employment Type
Contract
Contract Rate
GBP 400 Daily
deployment strategies Designing Back End services capable of meeting demanding performance and non-functional requirements Performance tuning and optimising Back End applications Working with observability and monitoring tools to maintain reliability and performance Collaborating closely with other engineers while also taking ownership of individual technical areas Contributing to coding standards … best practices and Agile development Experience optimising and fine-tuning Back End applications against demanding NFRs Strong analytical and problem-solving skills Experience with observability and monitoring tools such as Splunk or Dynatrace Comfortable working collaboratively within a team while also taking ownership and working independently Desirable experience Experience developing ...

Lead AI Software Engineer

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle. What you will do Lead build execution within a squad, platform capability, or engineering domain. Translate … generated and engineer‐written code to ensure correctness, maintainability, security, and alignment with specifications. Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification. Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership. Ensure AI‐generated outputs are explainable, traceable, secure ...

Lead Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, Nottinghamshire, United Kingdom
Salary
£ 70 K
Role Profile:We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a Technical Lead SRE, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. … projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design.Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.Design and evolve monitoring and alerting solutions that ...

Lead Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, UK
Employment Type
Full-time
Role Profile: We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division. As a Technical Lead SRE, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing … projects building environments, monitoring, alerting, and ensuring operational readiness from day one. Collaborate with Architecture and Engineering teams to embed reliability, scalability, security, and observability into system design. Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs. Design and evolve monitoring ...

Senior Backend Engineer | AI Platform

Location
Greater London, England, United Kingdom
high degree of autonomy and ownership, as you'll be responsible for designing scalable solutions that empower multiple engineering teams while ensuring reliability, observability, and cost efficiency. What are we looking for: 5+ years of experience in Software Engineering, Backend Engineering, or Platform Engineering. Strong experience building and maintaining backend … LangChain, LangGraph, CrewAI, or similar. Experience working with cloud platforms such as Google Cloud Platform (preferred), AWS, or Azure. Strong understanding of system reliability, observability, monitoring, and incident management. Experience with Infrastructure as Code and cloud-native architectures. Previous experience working within a Platform Engineering team is a strong plus. ...

AWS DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£600/day
/CD pipelines utilising GitHub Actions and Jenkins. Deploy, manage and troubleshoot containerised applications with Docker and Kubernetes (EKS). Implement monitoring, logging and observability solutions to improve platform visibility. Automate operational and deployment processes using Python and Bash scripting. Ensure platform reliability, scalability and high availability. Maintain security, compliance … Python automation experience. Solid understanding of IAM and cloud security principles. Experience designing and implementing CI/CD pipelines. Exposure to monitoring, logging and observability tools. Experience working within Agile delivery environments. Contract Overview £600 per day, Outside IR35. Initial 6-month contract with strong extension potential. Hybrid working model ...

Senior Cloud SRE Lead: AWS & Kubernetes, Infra & CI/CD

Location
Greater London, England, United Kingdom
will build and maintain CI/CD pipelines using GitHub Actions, Docker, and Helm, and set up Prometheus/Grafana/Datadog for observability and incident response. Strong scripting in Python/Bash/Go is essential. #J-18808-Ljbffr ...

Lead AI Software Engineer London, United Kingdom Value Stream Engineering Posted 12 hours ago

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle.## **What you will do*** Lead build execution within a squad, platform capability, or engineering domain.* Translate … generated and engineer-written code to ensure correctness, maintainability, security, and alignment with specifications.* Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification.* Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership.* Ensure AI-generated outputs are explainable, traceable, secure ...

Staff AI Engineer, Payments Intelligence

Location
Greater London, England, United Kingdom
70+ languages, adapting to context and intent. Evaluation, safety, and guardrails — define how the unit measures agent quality and safety, with rigorous evaluation, observability, and guardrails so agents behave reliably in a compliance‐sensitive, multi‐market environment. Intelligence and data products — architect the client‐intelligence systems that turn payment data … systems at scale. Model serving and LLM inference, with the cloud infrastructure to run AI workloads in production (AWS or similar). Evaluation and observability for AI — hands‐on with tools like LangFuse, LangSmith, Braintrust, or MLflow — plus solid automated testing and CI/CD. Proven technical leadership — mentoring engineers ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to take ownership of the cloud infrastructure and DevOps practices powering our BIM Platform, working across Azure, Kubernetes, CI/CD, observability, security and developer tooling. The Role This is a hands‐on senior engineering role with broad ownership across our cloud and platform infrastructure. … Code using Terraform, ARM templates and Helm Design and improve CI/CD pipelines using GitHub Actions, enabling fast, safe and repeatable deployments Build observability across our infrastructure and services through monitoring, logging, alerting and distributed tracing Improve platform reliability, scalability and performance through capacity planning, autoscaling and resource optimisation ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Senior Platform Engineer Engineering · Manchester ·

Location
Manchester, England, United Kingdom
over £50 billion in loans annually, you will embed automated security scanning and compliance-as-code into delivery pipelines, spearhead cloud cost-optimization and observability initiatives, and provide technical guidance and mentorship to mid-level engineers across the team. About you: GCP Cloud Expertise: Proven track record managing, scaling … Paths," creating internal self-service tooling, writing developer documentation, and conducting code reviews to drive a strong "you build it, you run it" culture. Observability & Cost Optimization: Strong background in system resilience, latency reduction, observability implementation, and cloud cost-optimization initiatives (e.g., resource tagging and footprint reduction). What will ...

Senior DevOps Engineer

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
performance, and availability Investigate and resolve production incidents in a timely manner Perform root cause analysis and implement preventative measures Enhance logging, alerting, and observability across the platform Security & Governance Implement Azure security best practices and policies Manage identity and access controls in line with governance standards Ensure compliance with … modern development workflows Desirable Knowledge of scripting languages (PowerShell, Bash, or Python) Experience working with Azure Front Door, WAF, or CDN technologies Exposure to observability tooling (distributed tracing, metrics platforms) Experience with cost management and optimisation in Azure Understanding of DevOps principles and Agile delivery practices Personal Attributes Strong problem ...

Private Cloud Architect

Hiring Organisation
Experis
Location
England, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£411/day
engineering platforms used across security and data-focused teams. DevOps & SRE Build, maintain, and optimise CI/CD pipelines. Implement monitoring, alerting, logging, and observability solutions. Troubleshoot production incidents and contribute to operational support activities. Continuously improve service reliability, security, scalability, and performance. Essential Skills & Experience Strong software engineering background … Lake technologies. Experience supporting Data Engineering platforms. Understanding of API security concepts including OAuth2, JWT, and mTLS. Knowledge of Site Reliability Engineering (SRE) and observability practices. If you receive suspicious outreach claiming to be from us, please contact us via the ManpowerGroup website. ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, UK
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

DevOps / SRE Engineer (Sheffield)

Location
Sheffield, England, United Kingdom
DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud‐native platform. You’ll work across the business to improve how we build, deploy and monitor our systems , while playing a key role … platforms Solid experience with Terraform and IaC automation Experience participating in or managing production incidents and on‐call Strong grasp of monitoring, alerting, and observability principles Ability to diagnose and fix complex distributed systems issues Demonstrated use of GenAI tools (ChatGPT, GitHub Copilot, Claude) in engineering workflows Excellent communication ...

DevOps / SRE Engineer (London)

Location
Greater London, England, United Kingdom
DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud‐native platform. You’ll work across the business to improve how we build, deploy and monitor our systems , while playing a key role … platforms Solid experience with Terraform and IaC automation Experience participating in or managing production incidents and on‐call Strong grasp of monitoring, alerting, and observability principles Ability to diagnose and fix complex distributed systems issues Demonstrated use of GenAI tools (ChatGPT, GitHub Copilot, Claude) in engineering workflows Excellent communication ...

Lead Agentic AI Architect

Hiring Organisation
CBSbutler Holdings Limited trading as CBSbutler
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£710/day inside ir35
memory, orchestration and tool integration Production-grade LLM and RAG architectures Intelligent search, knowledge assistants, workflow automation and decision-support systems Enterprise LLMOps, MLOps, observability and AI governance AI security, guardrails, evaluation and Responsible AI Cloud AI platforms and hybrid/multi-cloud architecture You'll also provide technical leadership … Hybrid/Multi-Cloud Engineering & Platform Python | Java | Scala | TypeScript | SQL | Kubernetes | Docker | Terraform | Bicep | CI/CD | Infrastructure as Code AI Governance & Observability LLMOps | MLOps | Model Governance | AI Security & Guardrails | LangSmith | Arize | Monitoring & Evaluation Frameworks Qualifications Relevant qualifications may include: Master's degree in Computer Science, AI, Data Science ...

Lead Engineer, Site Reliability Engineering

Location
Greater London, England, United Kingdom
Write automation to scale systems sustainably, prevent service issues, or when they occur, quickly recover service. Partner with development teams to improve system reliability, observability, and release velocity. Participate in on-call rotations, incident response, postmortems, and root cause analysis and resolution. Be a vocal advocate of strong/sound … Infrastructure as Code using Terraform.* Hands on Experience with one of the following cloud platforms: Azure, AWS, or GCP.* Knowledge on Docker and Kubernetes.* Observability tools like Datadog, Dynatrace or similar.* Implement and maintain CI/CD pipelines.* Incident response and running blameless post-mortems.* Proficient in Git workflows.## ## ...

Software Engineer Backend

Location
West of England, England, United Kingdom
other runtime conditions. Design appropriate controls for authentication, authorization, permissions, and secure access to tools, models, services, and data. Design for scalability, reliability, performance, observability, and fault tolerance across services. Write high-quality, readable, testable code that follows established engineering standards and best practices. Participate in code and design reviews … strong engineering and quality practices across the team. Troubleshoot and resolve complex issues across backend services, workflows and integrations. Build monitoring, logging, testing, and observability into production services to support reliable agent execution and product operations. Collaborate with DevOps, and infrastructure teams to support reliable deployment and operation. Stay current ...

Analytics Services Platform Engineer

Location
Greater London, England, United Kingdom
developer experience Collaborating with research, data and engineering teams to accelerate time‐to‐insight through modern analytics solutions Driving improvements in automation, observability and resilience across analytics services Evaluating and adopting emerging technologies such as AI assistants, data mesh and cloud‐native analytics solutions Defining SLAs, KPIs and monitoring strategies … using Terraform or Ansible Deep understanding of AWS analytics technologies including EMR, MSK, Athena, Redshift, Glue and MWAA Experience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetry Strong problem‐solving skills and a systematic approach to diagnosing and resolving issues Highly Desirable Skills ...