1,076 to 1,100 of 4,792 Observability Jobs

Software Engineer III - Full Stack, Global Banking Tech

Location
Glasgow, Scotland, United Kingdom
analysis, analyzing diverse datasets, logs, and telemetry to identify patterns, build visualizations and reporting, and reduce repeat incidents via preventative controls, automation, and enhanced observability Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit … agentic frameworks such as Google ADK or LangChain Familiarity with monitoring, tracing, and troubleshooting tools such as log aggregation platforms, API testing tools, and observability dashboards #J-18808-Ljbffr ...

DevOps Engineer - SC Cleared - Hybrid - Inside IR35

Location
City Of London, England, United Kingdom
Experience of working within a production environment, with strong on-prem Kubernetes/OpenShift deployments IAC using Terraform with CI/CD pipelines (Jenkins) Observability tools to include design and operate end to end logging, metrics, tracing, dashboards to include alerting systems using ELK, Splunk, Grafana Cloud platforms to include ...

Senior Software Engineer I

Location
City Of London, England, United Kingdom
maintain generative AI services and reusable components using mostly Python and a little bit Java. Defining and promote best practices in engineering, including scalability, observability, testing, and CI/CD. Contributing to system designs spanning multiple services and modules, aligning with architectural best practices. Collaborating with product, platform, and research … work collaboratively across functions in an Agile or Kanban environment. Nice to Have Experience operationalizing LLMs or building internal AI platforms. Familiarity with observability practices (metrics, logging, alerts). Exposure to knowledge graphs or semantic search systems. U.S. National Base Pay Range U.S. National Base Pay Range ...

Senior Machine Learning Engineer (ML Platform)

Location
Greater London, England, United Kingdom
infrastructure. Improve automation across the ML lifecycle, including model packaging, deployment, versioning, monitoring, and release processes. Maintain and improve the reliability and observability of the ML platform and model-serving services, including logging, metrics, and alerting. Participate in the team's on-call rotation, investigate production incidents, and contribute … experience provisioning and managing cloud infrastructure with Terraform. Experience with CI/CD pipelines, automated testing, and Git-based development workflows. Familiarity with observability practices, including logging, metrics, alerting, and production troubleshooting. Strong grasp of software engineering principles and best practices. Experience contributing to or leading data warehouse architecture ...

Innovation Software Developer

Hiring Organisation
EOS IT Solutions
Location
Armagh, United Kingdom
Salary
£ 70 K
real-time and high-volume sensor streams.Design Modern APIs and ServicesDevelop secure, scalable REST and gRPC services using FastAPI, Flask, and Node.js.Implement authentication, authorization, observability, testing, and performance monitoring.Ensure reliability, maintainability, and best-practice API design.Deliver Digital Twin ExperiencesIntegrate real-time operational data with advanced 3D and 4D visualizations.Work with … modern machine learning frameworks including PyTorch and Scikit-learn.Support DevOps and Delivery ExcellenceDeploy containerized applications using Docker and Kubernetes.Implement CI/CD pipelines and observability tooling.Produce high-quality technical documentation and operational runbooks.What We're Looking ForEssential Skills & Experience5+ years of software development experience in complex environments.Strong expertise in Python ...

Vice President, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
engineering teams; set clear objectives, coach talent, and foster succession planning. Own end-to-end delivery for critical software: requirements, architecture, implementation, testing, deployment, observability, and reliability. Raise engineering excellence and resilience: best practices and automation across code, testing, microservices/APIs, performance, and infrastructure; secure-by-design with threat … scalable, observable, testable systems; strong API design. Strong DevOps practices: CI/CD (e.g., GitLab), automated testing (JUnit/Spock), code reviews, telemetry/observability (Splunk, AppDynamics), containers (Docker), and cloud. Hands-on AI development using modern tools and IDEs (e.g., Windsurf) and experience integrating AI into product workflows. Excellent ...

Vice President, Full-Stack Engineer Opportunities

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
engineering teams; set clear objectives, coach talent, and foster succession planning.. Own end-to-end delivery for critical software: requirements, architecture, implementation, testing, deployment, observability, and reliability. Raise engineering excellence and resilience: best practices and automation across code, testing, microservices/APIs, performance, and infrastructure; secure-by-design with threat … scalable, observable, testable systems; strong API design. Strong DevOps practices: CI/CD (e.g., GitLab), automated testing (JUnit/Spock), code reviews, telemetry/observability (Splunk, AppDynamics), containers (Docker), and cloud Hands-on AI development using modern tools and IDEs (e.g., Windsurf) and experience integrating AI into product workflows Excellent ...

Senior UI Architect , Vice President

Location
Greater London, England, United Kingdom
centric digital experiences that align with business objectives. Engineering Excellence Guide development teams through architectural implementation and technical challenges. Establish CI/CD, testing, observability, and quality engineering practices for UI applications. Lead performance optimisation initiatives, including Core Web Vitals improvements. Ensure maintainability, scalability, and resilience of front-end solutions. … REST APIs GraphQL OAuth2/OpenID Connect Enterprise SSO solutions Quality & Performance Web accessibility (WCAG) Core Web Vitals Automated testing frameworks Performance monitoring and observability Preferred Experience Large-scale financial services, banking, asset management, or regulated industry experience. Experience with Adobe Experience Manager (AEM) Headless CMS. Experience delivering enterprise digital ...

Site Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, United Kingdom
Salary
£ 70 K
platform reliability.Proficiency in Linux/Unix environments and Bash/Shell scripting.Good understanding of CI/CD principles and automated deployment pipelines.Experience with monitoring, observability, alerting, and operational telemetry platforms.Good troubleshooting and problem-solving skills with the ability to perform root cause analysis and incident remediation.Experience working in Agile teams … availability.Some other highly valued skills may help:Experience administering enterprise database platforms such as Oracle.Knowledge of Oracle Database architecture, PL/SQL.Experience with enterprise observability platforms such as Splunk, Elastic, Observe, Grafana, Prometheus.Familiarity with Infrastructure as Code and platform engineering practices.You may be assessed on the key critical skills relevant ...

Technical Lead (AWS Serverless & Microservices)

Location
Greater London, England, United Kingdom
through automation, testing practices and governance frameworks Promoting secure-by-design principles and embedding security best practice across teams and platforms Establishing and evolving observability capabilities including logging, metrics, tracing and alerting Championing metrics-led engineering to improve delivery performance, reliability and customer outcomes Providing technical guidance on complex engineering … Expertise in quality engineering, automated testing strategies and CI/CD practices Strong understanding of secure software development principles and security frameworks Experience implementing observability strategies using logging, metrics, tracing and alerting tools Proven ability to influence stakeholders and communicate complex technical concepts clearly Experience coaching, mentoring and developing high ...

Principal DevOps Engineer

Location
Greater London, England, United Kingdom
templates, shared steps, release promotion, rollback) Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable) Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards) Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning … Ansible (architecture, reusable components, secure operations) Strong deployment/release engineering experience with Octopus Deploy and GitHub (release governance, environment promotion, rollback) Monitoring/observability expertise with VictoriaMetrics and/or Prometheus (alerting strategy, metrics design, operational readiness) Production experience running Redis , RabbitMQ , Nginx (HA, tuning, troubleshooting) Strong understanding ...

Software Engineer, Enterprise

Location
Greater London, England, United Kingdom
applications can operate at enterprise scale across hybrid and multi-cloud environments. Manage and evolve cloud infrastructure (AWS, Azure, or GCP), driving automation, observability, and security for large-scale AI deployments. Collaborate with ML and product teams to bring cutting-edge GenAI models into production through efficient APIs, model serving … experience with GenAI applications, model integration, or AI agent systems—understanding how to deploy, evaluate, and scale AI workloads in production. Strong understanding of observability, CI/CD , and security best practices for running services in enterprise or multi-tenant environments. Ability to balance rapid iteration with production-grade quality ...

Security & Network Engineer - 12 months Fixed term

Hiring Organisation
Techtronic Industries - Europe HQ
Location
Maidenhead, Berkshire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

Engineering Manager

Location
United Kingdom
clear priorities, managing dependencies, identifying risks, and removing obstacles that impact team execution. Ensure solutions meet appropriate standards for quality, scalability, reliability, performance, security, observability, and maintainability. Establish and reinforce strong software engineering practices, including code reviews, automated testing, CI/CD, documentation, monitoring, and production readiness. Foster effective collaboration … partner effectively with Product Management and translate product and business objectives into engineering plans. Strong understanding of software quality, testing, CI/CD, observability, reliability, and operational excellence. Working knowledge of cloud platforms such as AWS or Azure and modern cloud-based application architectures. Strong problem-solving and decision-making ...

Principal AI Platform Engineer

Location
City of Edinburgh, Scotland, United Kingdom
security, performance and availability. Using technologies such as containers, Kubernetes, vLLM and AI gateway platforms, you will deploy and operate scalable inference services, improve observability and performance, and investigate complex technical issues across the platform. You will also help establish engineering standards for operating AI services within a secure enterprise … such as LiteLLM, Bifrost or similar GPU workloads, including performance, utilisation and resource management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Secure, resilient and scalable service design Authentication, access control, rate limiting and service integration Technical documentation and operational guidance Security Clearance ...

Context Plane Python Engineer

Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end‐to‐end — from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands‐on experience using enterprise-authorized AI‐assisted software development ...

Infrastructure Software Engineering – Platform & Build

Location
Greater London, England, United Kingdom
create, maintain and debug reproducible multi-language CI pipelines, and optimize CI performance across large compute clusters. Build and maintain infrastructure observability, alerting, runbooks, and incident response workflows for Fractile's infrastructure. IaC TODO Scale and maintain Fractile's Bazel monorepo as we continue growing across Python, C++, Rust, SystemVerilog … high performance and extensibility, such as Bazel, Buck, Pants, Please, etc. Experience with infrastructure as code (Terraform, OpenTofu, or Pulumi) Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on‐call practice Strong proficiency ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast‐paced front‐office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Senior Software Engineer: Agentic Development Enablement

Location
Greater London, England, United Kingdom
Claude Code and GitHub Copilot Design and implement practical guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Help define how controls should work consistently across local development environments and CI/CD pipelines Explore changes to the development environment … background as a software engineer Broad technical understanding across several of the following: developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end‐user application delivery Ability to work in ambiguous ...

Senior Software Engineer - Backend

Location
Manchester, England, United Kingdom
caching and data-access strategies Building reliable transactional workflows Developing applications and services within AWS Deploying and operating containerised applications Improving monitoring, alerting and observability Exploring AI-assisted engineering and modern development tooling At Senior level, you’ll be expected to understand the wider system rather than only the individual … traffic or data volumes Performance optimisation, caching and latency reduction Designing for resilience and failure Containerisation and orchestration technologies such as Docker and Kubernetes Observability, monitoring and operating production systems Automated testing and modern engineering practices We don’t expect candidates to have worked with every technology in our stack. ...

Senior Platform Engineer

Location
Warminster, England, United Kingdom
Support and maintain existing simulation and training systems, as well as existing deployment and virtualisation tools. Apply SRE practices to improve system reliability, including observability (metrics, logs, tracing), incident response, and root cause analysis. What We Are Looking For: This is not a pure cloud or greenfield platform role. … failures Pragmatic and delivery-focused, with a bias toward keeping systems running. Strong collaborator across engineering disciplines Adopts an SRE mindset, focusing on reliability, observability, and continuous improvement of running systems. Key Technical Proficiencies: Expert working knowledge of Kubernetes, Helm, Teraform, Ansible, and Docker. Understanding of Distributed Systems in production. ...

Software Engineering Specialist

Location
Belfast City District, Northern Ireland, United Kingdom
identify dependency, migration, and simplification opportunities. Cloud Platform, Reliability & Operations · Ensure billing services are deployed and operated effectively in AWS and EKS, with strong observability, logging, alerting, and operational controls. · Lead production readiness, resilience planning, performance tuning, and root‐cause analysis for customer and revenue‐impacting incidents. · Improve CI/… microservices, and API‐first architectures. Cloud Platforms (AWS/EKS): Proven expertise deploying and operating cloud‐native applications in AWS and EKS, including containerisation, observability, resilience, and CI/CD practices. Technical Leadership & Delivery: Experience leading end‐to‐end engineering delivery, making architectural decisions, mentoring engineers, and driving operational excellence ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...