751 to 775 of 4,110 Observability Jobs

SRE Engineer

Hiring Organisation
Ricoh
Location
London, UK
Employment Type
Full-time
practices, tooling, and engineering standardsDriving infrastructure‐as‐code and automation across Azure and on‐premImproving image bakery pipelines for secure, repeatable server buildsEmbedding observability using metrics, logs, traces, and effective alertingEnsuring all practices align with ISO 27001 and internal security frameworksManaging automated patching, vulnerability remediation and configuration complianceBuilding dashboards … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions)Experience with monitoring and observability stacksSolid understanding of OS fundamentals (Windows/Linux), security, networkingBackground in scripting or software development (PowerShell, Python, Go)Experience with containers and orchestration (Docker, Kubernetes ...

Senior AI Engineer

Location
Greater London, England, United Kingdom
passthrough and logic centre for enterprise datasets, consuming other MCP servers and presenting them through one governed interface; and the orchestration, evaluation, and observability services beneath Libros and PRISM, two of the projects we are delivering with Percepta. As a Senior Applied AI Engineer, reporting to the Principal AI Engineer … own. Platforms built jointly with Percepta transfer into our ownership with maintainable designs and a team that can extend them without external help. Evaluation, observability, and control are built into what you ship rather than bolted on before release. Your responsibilities Build and run our AI products and platform Design ...

Platform Engineer - Edinburgh

Location
City of Edinburgh, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post-incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills Strong experience with the AWS cloud platform and core services. Hands ...

Lead Java Developer

Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc. Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

Corporate KYC Sr Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Lead SRE

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Lead SRE - Chase UK

Location
London, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Lead SRE

Location
Westminster, West End, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Senior Software Engineer - Commercial Trading

Location
City of Westminster, England, United Kingdom
business. You’ll play a key role in shaping technical direction, improving engineering standards and ensuring our products and platforms are built with strong observability, operational excellence and best-in-class engineering practices. Due to high interest, this role may close earlier than advertised. We recommend applying as soon … team uses a variety of modern technologies, including: Backend: Java, Spring, Spring Boot, Micronaut Frontend: React, Next.js, TypeScript, Angular Cloud & Infrastructure: Azure Cloud, Kubernetes Observability: Dynatrace Databases: SQL Server, MongoDB Caching & Performance: Ignite, Redis What’s in it for you? Working at M&S means being part of something bigger ...

Lead Java Developer

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience: Strong experience in Core Java, J2EE, Spring FrameworkExposure to Python scripting and data analysisExperience in fast moving Capital Markets Front … messaging technologies such as Kafka, JMS, gRPC etcProficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuningExperience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed cachesHands ...

SC / NPPV3 DevOps Engineer - Azure

Location
Greater London, England, United Kingdom
using Docker, Kubernetes/AKS and Helm. Develop secure cloud landing zones, governance and infrastructure aligned with security and compliance requirements. Implement monitoring and observability using Azure Monitor, Log Analytics, Application Insights, Prometheus and Grafana. Automate infrastructure and operational processes using Python, PowerShell and Bash. Troubleshoot complex production issues ...

Cloud Network Engineer - Active UK Government Security Clearance Required

Location
Greater London, England, United Kingdom
Support routing, switching, firewalling, VPN, and load-balancing technologies Contribute to network upgrades, migrations, optimisation, and continuous improvement Monitor network health and performance using observability and monitoring tooling Participate in root-cause analysis and problem management Maintain technical documentation, configuration records, and operational procedures Collaborate with engineering, service management … networking BGP/OSPF VLANs/VRFs MPLS Firewalls such as Palo Alto or Check Point VPN/IPSec Load balancing Network monitoring and observability Enterprise networking in production environments Nice to have: Cisco ACI/Nexus Cisco SD-WAN/SDA AWS networking - VPC, Transit Gateway, Route 53 Azure ...

Senior Database Administrator - Core Infrastructure

Hiring Organisation
Kraken
Location
Moffat, Dumfries & Galloway, UK
Employment Type
Full-time
sound design and regularly validated procedures. Reduce manual operational work by building automation, improving process consistency, and enabling safe, low-friction database workflows. Improve observability and alert quality by championing meaningful metrics, reducing noise, and ensuring operational clarity across environments. Enhance database security through robust access controls, disciplined patching … orchestration platforms (building container images and managing Kubernetes workloads at scale).Strong security instincts around access control, upgrade processes, and safe operational workflows. Observability expertise: monitoring, alerting hygiene, and readiness for incident response. Strong communication and collaboration skills with the ability to partner with stakeholders, negotiate long-term plans, write ...

DevOps / GenAI Engineer

Hiring Organisation
Tenth Revolution Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£400.00 - £450.00 per day
controls into the deployment lifecycle Use AWS Kiro and AI agents to automate DevOps workflows Collaborate with development, infrastructure, and security teams Improve monitoring, observability, and deployment reliability Required Skillset Experience with AWS Kiro (L300 certification required or willingness to complete the certifications) Proven background building CI/CD pipelines ...

Senior Infrastructure & Security Engineer

Location
Swindon, England, United Kingdom
recovery and business-continuity capability: design it, test it regularly, and automate backup/restore and failover wherever possible.* Build and maintain alerting and observability so the right people can see what they need — consolidating and improving across Grafana/Loki (application logs), ManageEngine log analytics (server logs) and PRTG … hire.* Hands-on Oracle Cloud Infrastructure (OCI).* Direct experience with the specific security tooling: CrowdStrike, ESET, Rapid7.* Octopus Deploy.* ManageEngine.* MySQL administration.* Observability stacks: Grafana/Loki, PRTG.* Experience migrating manually-built environments to infrastructure-as-code.* Direct experience with ISO 27001, Cyber Essentials Plus and MOD frameworks ...

Senior or Staff Software Engineer, SRE/ Platform Team

Location
Greater London, England, United Kingdom
code with Kubernetes and Terraform. You'll be at the forefront of shaping our foundational architecture, ensuring it’s both resilient and scalable. Drive Observability and Monitoring: Establish and maintain a state‐of‐the‐art observability and monitoring stack. Your insights will enable us to stay ahead of potential issues ...

Senior AI Platform Engineer

Hiring Organisation
9fin
Location
London, UK
Employment Type
Full-time
developer tooling that enable self-service AI development across engineering teams. Design secure, scalable deployment pipelines for AI models and applications. Build AI observability capabilities including monitoring, tracing, evaluation, cost optimisation, and production quality measurement. Collaborate closely with AI Engineers, Backend Engineers and Engineering Leadership to define platform architecture … monitor, and operate AI services in production. AI Operations & ObservabilityHave experience implementing monitoring, tracing, evaluation, and cost optimisation for AI systems. Have experience with observability solutions such as Arize Phoenix, Langfuse, or LangsmithUnderstand the operational challenges of deploying LLM-powered applications, including latency, reliability, hallucination monitoring, and model quality evaluation. ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices. Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks. Linux & Automation: Strong Linux and Git fundamentals with Bash ...

Staff Software Engineer, Liquidity Management (C#/.NET) New London, UK

Location
Greater London, England, United Kingdom
work that distributes effectively across the team Establish coding standards, review practices, and testing strategies that improve overall code quality Champion engineering methodologies including observability, incident response, and production excellence Share your enthusiasm for tech trends, explore and learn new technologies, engage with tech communities, mentor fellow engineers, and lead … cloud platform (preferably Azure) Experience with containerization (Docker) and orchestration (Kubernetes) Understanding of infrastructure as code and CI/CD pipeline design Knowledge of observability: logging, metrics, tracing, and alerting strategies Experience with microservices deployment patterns and service mesh concepts Core Technologies: Deep expertise with C#/.NET Expert ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
features and platform capabilities end to end, from design through build, test and deployment into production. Contribute to production readiness: documentation, runbooks, monitoring and observability – ensuring nothing ships until it’s genuinely ready. Maintain and improve existing platform components, addressing technical debt and performance issues with clear prioritisation. Standards & Quality … with data pipelines, data quality processes and the infrastructure that supports AI/ML at scale. Experience with DevOps, CI/CD pipelines and observability tooling. Exposure to telecom or large-scale consumer-facing platforms. SKILLS & BEHAVIOURAL COMPETENCIES Technical Depth & Judgement Consistently makes sound technical decisions, balancing short-term delivery ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real-time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost-control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Senior Software Engineer - Customer & Claims

Location
Greater London, England, United Kingdom
Create an environment where less experienced engineers can grow and do their best work. Continuous Improvement Drive improvements to engineering practices, testing, deployment and observability within the domain. Continuously grow your own capability and that of the engineers around you. Contribute to Homeprotect's wider engineering standards as they mature. … oriented language), and familiarity with a modern frontend framework. Deep, hands‐on understanding of modern software engineering practices: continuous integration and deployment, automated testing, observability and cloud‐native development. Experience operating production systems, in a public cloud environment such as Azure or GCP, that need to be reliable, secure ...

Software Development Coach

Hiring Organisation
McGregor Boyall Associates Limited
Location
Scotland, United Kingdom
Employment Type
Contract
Bring Significant commercial experience in software engineering and technical leadership. Strong knowledge of cloud-native development, CI/CD, infrastructure as code, automated testing, observability and secure software development. Practical experience with AWS, GitLab or comparable platforms, APIs, containers and modern programming languages. Experience coaching and mentoring teams with different ...

AI Platform Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, Castle Baynard, United Kingdom
Employment Type
Permanent
Salary
£80000 - £100000/annum
platform components. Deploying infrastructure using Terraform and supporting containerised applications. Building and maintaining CI/CD pipelines using GitHub and Azure DevOps. Improving observability, monitoring and platform resilience. Supporting vector search, embedding pipelines and knowledge ingestion. Applying security and governance best practice across the AI platform. What we're looking ...

Software Development Coach

Location
United Kingdom
Bring Significant commercial experience in software engineering and technical leadership. Strong knowledge of cloud-native development, CI/CD, infrastructure as code, automated testing, observability and secure software development. Practical experience with AWS, GitLab or comparable platforms, APIs, containers and modern programming languages. Experience coaching and mentoring teams with different ...