1,401 to 1,425 of 6,401 Observability Jobs

Software Engineer, Enterprise

Location
Greater London, England, United Kingdom
applications can operate at enterprise scale across hybrid and multi-cloud environments. Manage and evolve cloud infrastructure (AWS, Azure, or GCP), driving automation, observability, and security for large-scale AI deployments. Collaborate with ML and product teams to bring cutting-edge GenAI models into production through efficient APIs, model serving … experience with GenAI applications, model integration, or AI agent systems—understanding how to deploy, evaluate, and scale AI workloads in production. Strong understanding of observability, CI/CD , and security best practices for running services in enterprise or multi-tenant environments. Ability to balance rapid iteration with production-grade quality ...

Security & Network Engineer - 12 months Fixed term

Hiring Organisation
Techtronic Industries - Europe HQ
Location
Maidenhead, Berkshire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

Engineering Manager

Location
United Kingdom
clear priorities, managing dependencies, identifying risks, and removing obstacles that impact team execution. Ensure solutions meet appropriate standards for quality, scalability, reliability, performance, security, observability, and maintainability. Establish and reinforce strong software engineering practices, including code reviews, automated testing, CI/CD, documentation, monitoring, and production readiness. Foster effective collaboration … partner effectively with Product Management and translate product and business objectives into engineering plans. Strong understanding of software quality, testing, CI/CD, observability, reliability, and operational excellence. Working knowledge of cloud platforms such as AWS or Azure and modern cloud-based application architectures. Strong problem-solving and decision-making ...

Principal AI Platform Engineer

Location
City of Edinburgh, Scotland, United Kingdom
security, performance and availability. Using technologies such as containers, Kubernetes, vLLM and AI gateway platforms, you will deploy and operate scalable inference services, improve observability and performance, and investigate complex technical issues across the platform. You will also help establish engineering standards for operating AI services within a secure enterprise … such as LiteLLM, Bifrost or similar GPU workloads, including performance, utilisation and resource management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Secure, resilient and scalable service design Authentication, access control, rate limiting and service integration Technical documentation and operational guidance Security Clearance ...

Context Plane Python Engineer

Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end‐to‐end — from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands‐on experience using enterprise-authorized AI‐assisted software development ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
create, maintain and debug reproducible multi-language CI pipelines, and optimize CI performance across large compute clusters. Build and maintain infrastructure observability, alerting, runbooks, and incident response workflows for Fractile's infrastructure. IaC TODO Scale and maintain Fractile's Bazel monorepo as we continue growing across Python, C++, Rust, SystemVerilog … high performance and extensibility, such as Bazel, Buck, Pants, Please, etc. Experience with infrastructure as code (Terraform, OpenTofu, or Pulumi) Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on‐call practice Strong proficiency ...

Senior AI Engineer| London

Hiring Organisation
Infosys Technologies
Location
London, UK
Employment Type
Full-time
Agentic AI, classic ML and automation space Experience and good understanding of GenAI prompt engineering, RAG pipelines, Supervised/unsupervised ML and AI observability Experience in Enterprise-grade RAG-based solutions with LLMs (OpenAI, Hugging Face, LLaMA, etc.) and vector databases (Pinecone, Weaviate, FAISS). Experience in designing and scaling … Knowledge of AgentOps and OpenTelemetry Understanding of Network Security Concepts, Network Telemetry and Analytics Understanding of Cloud computing and Virtualization Exposure to APM/Observability tools (Dynatrace, AppDynamics, Datadog, Splunk etc) Exposure to onshore-offshore model working with professionals spread across the globePersonalBesides the professional qualifications, we respect and place ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Senior Software Engineer: Agentic Development Enablement

Location
Greater London, England, United Kingdom
Claude Code and GitHub Copilot Design and implement practical guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Help define how controls should work consistently across local development environments and CI/CD pipelines Explore changes to the development environment … background as a software engineer Broad technical understanding across several of the following: developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end‐user application delivery Ability to work in ambiguous ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Senior Software Engineer - Backend

Location
Manchester, England, United Kingdom
caching and data-access strategies Building reliable transactional workflows Developing applications and services within AWS Deploying and operating containerised applications Improving monitoring, alerting and observability Exploring AI-assisted engineering and modern development tooling At Senior level, you’ll be expected to understand the wider system rather than only the individual … traffic or data volumes Performance optimisation, caching and latency reduction Designing for resilience and failure Containerisation and orchestration technologies such as Docker and Kubernetes Observability, monitoring and operating production systems Automated testing and modern engineering practices We don’t expect candidates to have worked with every technology in our stack. ...

Senior Platform Engineer

Location
Warminster, England, United Kingdom
Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This is not a pure cloud or greenfield platform role. … failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. Understanding of Distributed Systems in production. ...

Software Engineering Specialist

Location
Belfast City District, Northern Ireland, United Kingdom
identify dependency, migration, and simplification opportunities. Cloud Platform, Reliability & Operations · Ensure billing services are deployed and operated effectively in AWS and EKS, with strong observability, logging, alerting, and operational controls. · Lead production readiness, resilience planning, performance tuning, and root‐cause analysis for customer and revenue‐impacting incidents. · Improve CI/… microservices, and API‐first architectures. Cloud Platforms (AWS/EKS): Proven expertise deploying and operating cloud‐native applications in AWS and EKS, including containerisation, observability, resilience, and CI/CD practices. Technical Leadership & Delivery: Experience leading end‐to‐end engineering delivery, making architectural decisions, mentoring engineers, and driving operational excellence ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Dunfermline, Fife, UK
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Livingston, West Lothian, UK
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Remote Data Engineering Manager

Location
United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...

Remote Data Engineering Manager

Location
Durham, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Brynteg, Denbighshire, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Halesworth, Suffolk, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Neath, Glamorgan, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Chesham, Buckinghamshire, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...

Remote Data Engineering Manager

Location
Lechlade, Gloucestershire, United Kingdom
workflow from reactive problem-solving to structured, agile delivery. Oversee the maintenance and optimization of high-performance data pipelines, implementing CI/CD automation, observability frameworks, and strict data quality gates. Roll up your sleeves when necessary to assist with complex code reviews, Python/Scala development, or unblocking … Python (and/or Scala) and advanced SQL across relational and non-relational databases. Experience implementing CI/CD, Infrastructure as Code, and observability/monitoring for data pipelines. Bachelor's degree in Computer Science, Engineering, Statistics, Information Systems, or a related quantitative field. Experience implementing data contracts, data catalogues ...