976 to 1,000 of 6,285 Permanent Observability Jobs

Corporate KYC Sr Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Paisley, Renfrewshire, UK
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Lead Java Developer

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience: Strong experience in Core Java, J2EE, Spring FrameworkExposure to Python scripting and data analysisExperience in fast moving Capital Markets Front … messaging technologies such as Kafka, JMS, gRPC etcProficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuningExperience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed cachesHands ...

Software Engineer III - Full Stack, Global Banking Tech

Location
Glasgow, Scotland, United Kingdom
analysis, analyzing diverse datasets, logs, and telemetry to identify patterns, build visualizations and reporting, and reduce repeat incidents via preventative controls, automation, and enhanced observability Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit … agentic frameworks such as Google ADK or LangChain Familiarity with monitoring, tracing, and troubleshooting tools such as log aggregation platforms, API testing tools, and observability dashboards J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals ...

SC / NPPV3 DevOps Engineer - Azure

Location
Greater London, England, United Kingdom
using Docker, Kubernetes/AKS and Helm. Develop secure cloud landing zones, governance and infrastructure aligned with security and compliance requirements. Implement monitoring and observability using Azure Monitor, Log Analytics, Application Insights, Prometheus and Grafana. Automate infrastructure and operational processes using Python, PowerShell and Bash. Troubleshoot complex production issues ...

Cloud Network Engineer - Active UK Government Security Clearance Required

Location
Greater London, England, United Kingdom
Support routing, switching, firewalling, VPN, and load-balancing technologies Contribute to network upgrades, migrations, optimisation, and continuous improvement Monitor network health and performance using observability and monitoring tooling Participate in root-cause analysis and problem management Maintain technical documentation, configuration records, and operational procedures Collaborate with engineering, service management … networking BGP/OSPF VLANs/VRFs MPLS Firewalls such as Palo Alto or Check Point VPN/IPSec Load balancing Network monitoring and observability Enterprise networking in production environments Nice to have: Cisco ACI/Nexus Cisco SD-WAN/SDA AWS networking - VPC, Transit Gateway, Route 53 Azure ...

DevOps / GenAI Engineer

Hiring Organisation
Tenth Revolution Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£400.00 - £450.00 per day
controls into the deployment lifecycle Use AWS Kiro and AI agents to automate DevOps workflows Collaborate with development, infrastructure, and security teams Improve monitoring, observability, and deployment reliability Required Skillset Experience with AWS Kiro (L300 certification required or willingness to complete the certifications) Proven background building CI/CD pipelines ...

Senior Infrastructure & Security Engineer

Location
Swindon, England, United Kingdom
recovery and business-continuity capability: design it, test it regularly, and automate backup/restore and failover wherever possible.* Build and maintain alerting and observability so the right people can see what they need — consolidating and improving across Grafana/Loki (application logs), ManageEngine log analytics (server logs) and PRTG … hire.* Hands-on Oracle Cloud Infrastructure (OCI).* Direct experience with the specific security tooling: CrowdStrike, ESET, Rapid7.* Octopus Deploy.* ManageEngine.* MySQL administration.* Observability stacks: Grafana/Loki, PRTG.* Experience migrating manually-built environments to infrastructure-as-code.* Direct experience with ISO 27001, Cyber Essentials Plus and MOD frameworks ...

Senior or Staff Software Engineer, SRE/ Platform Team

Location
Greater London, England, United Kingdom
code with Kubernetes and Terraform. You'll be at the forefront of shaping our foundational architecture, ensuring it’s both resilient and scalable. Drive Observability and Monitoring: Establish and maintain a state‐of‐the‐art observability and monitoring stack. Your insights will enable us to stay ahead of potential issues ...

Senior AI Platform Engineer

Hiring Organisation
9fin
Location
London, UK
Employment Type
Full-time
developer tooling that enable self-service AI development across engineering teams. Design secure, scalable deployment pipelines for AI models and applications. Build AI observability capabilities including monitoring, tracing, evaluation, cost optimisation, and production quality measurement. Collaborate closely with AI Engineers, Backend Engineers and Engineering Leadership to define platform architecture … monitor, and operate AI services in production. AI Operations & ObservabilityHave experience implementing monitoring, tracing, evaluation, and cost optimisation for AI systems. Have experience with observability solutions such as Arize Phoenix, Langfuse, or LangsmithUnderstand the operational challenges of deploying LLM-powered applications, including latency, reliability, hallucination monitoring, and model quality evaluation. ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices. Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks. Linux & Automation: Strong Linux and Git fundamentals with Bash ...

Staff Software Engineer, Liquidity Management (C#/.NET) New London, UK

Location
Greater London, England, United Kingdom
work that distributes effectively across the team Establish coding standards, review practices, and testing strategies that improve overall code quality Champion engineering methodologies including observability, incident response, and production excellence Share your enthusiasm for tech trends, explore and learn new technologies, engage with tech communities, mentor fellow engineers, and lead … cloud platform (preferably Azure) Experience with containerization (Docker) and orchestration (Kubernetes) Understanding of infrastructure as code and CI/CD pipeline design Knowledge of observability: logging, metrics, tracing, and alerting strategies Experience with microservices deployment patterns and service mesh concepts Core Technologies: Deep expertise with C#/.NET Expert ...

Senior Data Architect

Hiring Organisation
Eneco
Location
Rotterdam, Zuid-Holland, Netherlands
Employment Type
Permanent
Salary
EUR Annual
responsible for Defining and owning the target architecture and evolution roadmap for Data & GenAI Platforms, covering ingestion, transformation, orchestration, storage, serving, observability, and governance. Architecting and scaling a Lakehouse-based design interoperating with Databricks and Snowflake for analytics, ML, and serving use cases. Building and formalizing the GenAI ecosystem, including … product teams to rapidly deploy AI environments and workflows. Industrializing the modern data stack with standardized dbt modeling practices and Airflow orchestration conventions. Implementing observability and FinOps monitoring for pipelines, models, cost control, quality, compliance, and performance. Operationalizing data governance through Collibra (or equivalent), embedding GDPR-compliant privacy-by-design ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
features and platform capabilities end to end, from design through build, test and deployment into production. Contribute to production readiness: documentation, runbooks, monitoring and observability – ensuring nothing ships until it’s genuinely ready. Maintain and improve existing platform components, addressing technical debt and performance issues with clear prioritisation. Standards & Quality … with data pipelines, data quality processes and the infrastructure that supports AI/ML at scale. Experience with DevOps, CI/CD pipelines and observability tooling. Exposure to telecom or large-scale consumer-facing platforms. SKILLS & BEHAVIOURAL COMPETENCIES Technical Depth & Judgement Consistently makes sound technical decisions, balancing short-term delivery ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real-time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost-control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Senior Software Engineer - Customer & Claims

Location
Greater London, England, United Kingdom
Create an environment where less experienced engineers can grow and do their best work. Continuous Improvement Drive improvements to engineering practices, testing, deployment and observability within the domain. Continuously grow your own capability and that of the engineers around you. Contribute to Homeprotect's wider engineering standards as they mature. … oriented language), and familiarity with a modern frontend framework. Deep, hands‐on understanding of modern software engineering practices: continuous integration and deployment, automated testing, observability and cloud‐native development. Experience operating production systems, in a public cloud environment such as Azure or GCP, that need to be reliable, secure ...

AI Platform Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, Castle Baynard, United Kingdom
Employment Type
Permanent
Salary
£80000 - £100000/annum
platform components. Deploying infrastructure using Terraform and supporting containerised applications. Building and maintaining CI/CD pipelines using GitHub and Azure DevOps. Improving observability, monitoring and platform resilience. Supporting vector search, embedding pipelines and knowledge ingestion. Applying security and governance best practice across the AI platform. What we're looking ...

Software Development Coach

Location
United Kingdom
Bring Significant commercial experience in software engineering and technical leadership. Strong knowledge of cloud-native development, CI/CD, infrastructure as code, automated testing, observability and secure software development. Practical experience with AWS, GitLab or comparable platforms, APIs, containers and modern programming languages. Experience coaching and mentoring teams with different ...

Software Development Coach

Hiring Organisation
MCGREGOR BOYALL ASSOCIATES LIMITED
Location
London, UK
Bring Significant commercial experience in software engineering and technical leadership. Strong knowledge of cloud-native development, CI/CD, infrastructure as code, automated testing, observability and secure software development. Practical experience with AWS, GitLab or comparable platforms, APIs, containers and modern programming languages. Experience coaching and mentoring teams with different ...

Forward Deployed AI Engineer

Location
Birmingham, England, United Kingdom
frontier models where reasoning genuinely earns it), prompt and context budgeting, caching. Cost is a non-negotiable metric on every engagement. Ship it properly. Observability and tracing, CI/CD, IaC, monitoring the client can actually operate, security and compliance review. Transfer knowledge deliberately. We don't run a long … tool exposure, auth patterns. Claude Code as an autonomous SDLC agent: sub‐agents, hooks, custom skills, agent marketplaces, context sharing across a team. LLM observability and tracing tooling (LangFuse, LangSmith, Arize, Braintrust or similar). Regulated‐environment delivery: HIPAA, GDPR, SOC 2, FCA/PRA, PII handling, data residency, guardrails ...

DevOps Engineer

Hiring Organisation
SF Partners
Location
Nationwide, United Kingdom
Employment Type
Permanent
Salary
£75000 - £110000/annum excellent training & progression
most of the following key skills: - Deep technical ownership of IDP's for a large Developer user base - Knowledge of cloud native system design - Observability stack exposure - Grafana, Prometheus, Open Telemetry etc - Experience designing AWS and supporting AWS landing zones - IAC experience - Terraform, Ansible, Redhat etc - Strong experience building ...

Machine Learning Engineer (Applied AI ML)

Location
Greater London, England, United Kingdom
learning architectures such as transformers, CNNs, and autoencoders Specialism or well-researched interest in NLP Broad knowledge of MLOps tooling for versioning, reproducibility, and observability Experience monitoring, maintaining, and enhancing existing models over an extended period Extensive experience with PyTorch and related data science Python libraries such as pandas Experience ...

data engineer in data platforms

Location
Greater London, England, United Kingdom
technical advisor Set engineering standards, patterns, and best practices across teams Review designs and code, providing technical direction and mentorship Improve data quality, testing, observability, and operational excellence Требования Strong Python and SQL skills Deep experience with Spark and modern data platforms such as Databricks and Snowflake Solid understanding ...

Platform Engineer - Enterprise Platforms

Location
Greater London, England, United Kingdom
hosting Managing AWS services including S3, FSx, EBS and WorkSpaces/AppStream Developing configuration management solutions using tools such as Puppet Maintaining monitoring and observability capabilities Supporting enterprise backup and recovery platforms Building and maintaining CI/CD pipelines using GitHub Actions Supporting production incidents, problem management and platform requests ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
security, performance, maintainability, and testing. Contribute to technical standards and continuous improvement initiatives within the team. Operational Excellence & Collaboration: Maintain reliable production environments, improve observability and monitoring systems, and respond to proactive and reactive support alerts. Work collaboratively across the Technology department while mentoring and supporting other engineers. What ...

TypeScript Engineer

Location
Greater London, England, United Kingdom
these too, but support will be offered if not: Familiarity with infrastructure as code, for example CDK or Terraform. Exposure to observability tooling, incident response, or production monitoring practices. A broader understanding of LLM concepts such as tokens, embeddings, hallucinations, and safe use cases in software delivery. What ...