1 to 25 of 61 Permanent Observability Jobs in the City of London

Production Engineer

Location
City Of London, England, United Kingdom
identify patterns, and drive intelligent automation solutions Hands-on experience with containerization technologies such as Docker and Kubernetes, including cluster management, deployment, scaling, and observability Deep practical knowledge of algorithmic trading workflows, including the behaviour, lifecycle, and risk controls of execution algos used across the EMEA markets Experience designing ...

DevOps Engineer

Hiring Organisation
Anson McCade
Location
City of London, London, United Kingdom
/version control • CI/CD – Jenkins, GitHub Actions, GitLab, Bamboo, Bitbucket etc. • Docker/Kubernetes • Cloud networking, Linux, routing and security • Monitoring, observability and operational tooling • Agile delivery environments • Microservices and API development National Security & Clearance This is a National Security-focused role. You'll be working on projects ...

Senior Java Developer

Hiring Organisation
HCLTech
Location
City of London, London, United Kingdom
tools such as Maven or Gradle Containerization and Orchestration - Practical experience with Docker and Kubernetes, including deploying and managing containerized Java services Monitoring and Observability - Experience with monitoring and alerting tooling such as Geneos, Grafana and Prometheus or equivalent, with an appreciation of what good operational visibility looks like ...

Senior Software Engineer

Hiring Organisation
Method-Resourcing
Location
City of London, London, United Kingdom
Employment Type
Permanent
distributed systems where availability and performance are critical Identifying and solving complex performance and scalability challenges Contributing to architectural and technical decisions Improving resilience, observability and overall system reliability Working closely with engineers and technical stakeholders to deliver high-quality software Supporting infrastructure, deployment and continuous improvement The Technology ...

Senior AI Engineer

Location
City Of London, England, United Kingdom
agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems, identifying operational requirements for logging, instrumentation and alerting. Instrument solutions to capture outcome data against established baselines. Integrate AI capabilities ...

Machine Learning Engineering Lead

Hiring Organisation
Hackajob Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
maintainability, reliability, and production support. Strong understanding of ML engineering and MLOps practices, including model lifecycle management, CI/CD, testing, monitoring, release management, observability, and operational support. Practical experience with LLM-based capabilities, including retrieval-augmented generation, semantic search, embeddings, prompt design, evaluation, guardrails, and observability. Experience with agentic ...

Engineering Team Lead | Tech | London, UK

Location
City Of London, England, United Kingdom
other multi-agent/LLM orchestration frameworks Experience configuring Azure services (App Service, Container Registry, Blob Storage) or AWS equivalents Familiarity with LLM observability/evaluation tooling (e.g. Langfuse) and resilience patterns for third-party AI providers (e.g. circuit breakers, provider fallback) Awareness of data protection and AI transparency considerations ...

Lead Agentic AI Engineer

Location
City Of London, England, United Kingdom
implement evaluation frameworks for quality, grounding, task success, safety, latency, and cost Experience deploying AI solutions into production with focus on cost optimisation, scalability, observability, security, and governance Experienced in using Microsoft GenAI ecosystem, including M365, Copilot Studio, Azure AI Foundry, and Microsoft Agent Framework Proficient in Python and production ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
user experiences. ThousandEyes is deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform ...

Associate Director Lead AI Architect

Hiring Organisation
Anson Mccade
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
services . Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama . Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact. Optimise AI solutions for performance, scalability, security and cost-efficiency. Identify where ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
hands-on, Linux-focused DevOps role where you will take ownership of large-scale production environments, working across Linux systems, automation, Kubernetes, containerisation, networking, observability and cloud infrastructure. Location: London, hybrid working with a minimum of 2 days per week in the office Salary: Up to £100,000 per annum … networking, security and performance troubleshooting within Linux environments Hands-on experience operating containerised workloads using Docker and Kubernetes Experience with monitoring, logging and observability technologies such as Prometheus, Grafana, Loki, OpenTelemetry or the ELK Stack Experience with Infrastructure as Code and automation tooling such as Terraform and Ansible Experience building ...

Data Engineer

Hiring Organisation
The Travel Corporation
Location
City of London, London, United Kingdom
dbt. Build analytical models that give consistent views of customers, bookings, products and key entities. Implement automated testing, data quality controls, monitoring and observability across critical datasets. Apply modern engineering practice including Git, code review, automated testing and CI/CD. Monitor and optimise workloads for performance, scalability, resilience ...

Engineering Manager

Location
City Of London, England, United Kingdom
things down when something surfaces. Run the tribe's services in production. Client money moves through them, so incident response, on-call health and observability are yours. Reviews are blameless, and what we learn has to change something. Decide how the team uses AI. We use Claude Code every ...

Platform Engineer

Hiring Organisation
Next Ventures
Location
City of London, London, United Kingdom
Core Focus) Build secure, scalable Azure and GCP cloud platforms. Drive IaC adoption using Terraform and cloud‐native tooling. Implement automation, CI/CD, observability, and reliability engineering practices. Define standards for networking, identity, security, monitoring, and resilience. Improve developer experience through platform consistency, automation, and self‐service. 🔐 Governance, Security ...

Senior Data Engineer

Hiring Organisation
The Travel Corporation
Location
City of London, London, United Kingdom
version of the truth. Engineer solutions in Snowflake, SQL, Python and modern transformation tooling such as dbt. Raise the bar on testing, data quality, observability, version control and CI/CD, and make good practice the default. Tune workloads for performance, resilience, scalability and sensible cloud cost. Lead technically through ...

Senior Data Platform Engineer (Data Lake and Catalog)

Location
City Of London, England, United Kingdom
example in engineering quality through clean, well-tested code and strong coding practices. Promote an SLO-driven culture by contributing to reliability goals, observability, and incident learnings. Partner closely with engineers and other stakeholders to clarify requirements and deliver effectively. Break down complex problems into pragmatic technical solutions, estimates ...

Senior Product Manager

Hiring Organisation
Tata Consultancy Services
Location
City of London, London, United Kingdom
continuous improvement. Desirable Skills: Experience with data platforms such as Snowflake, Azure Data Lake, or similar cloud-based technologies. Understanding of model monitoring, observability, and performance tracking practices. Exposure to AI/ML-enabled products and data monetization initiatives. Rewards & Benefits TCS is consistently voted a Top Employer ...

Senior Data Platform Engineer

Location
City Of London, England, United Kingdom
data needs. Partner closely with internal teams to understand their data requirements, translating them into well-engineered, production-ready solutions. Set up and own observability for data workloads - monitoring, alerting, and logging - and lead incident response when things don't go to plan. Participate actively in code reviews, raising ...

Senior Frontend Engineer

Location
City Of London, England, United Kingdom
software works Make architectural calls on your domain, record down, present and defend them in front of the tribe. Build in security, reliability and observability from the start. Use AI tooling well. We use Claude Code daily and we expect judgement about where it earns its place and where ...

Site Reliability Engineer

Hiring Organisation
SR2 | Socially Responsible Recruitment | Certified B Corporation™
Location
City of London, London, United Kingdom
Site Reliability Engineer (SRE) DevSecOps | Cloud Engineering | Observability | Production Environments | London SR2 is supporting a major 3-year programme and looking for an experienced Site Reliability Engineer (SRE) to join the Production Engineering team. This function underpins the reliability, security, and performance of all live environments, from production systems … native infrastructure. Beyond supporting live systems, this team also acts as a centre of excellence, guiding project teams in adopting best practices across DevSecOps, observability, and cost optimisation. Key Responsibilities: Build, maintain, and support production and demo environments Automate infrastructure provisioning and deployment workflows (Terraform, GitHub Actions, GitOps) Package ...

Technical Product Manager

Location
City Of London, England, United Kingdom
technologies and third-party platforms to identify opportunities that accelerate innovation and product transformation. Ensure products are built with resiliency, scalability, security, reliability and observability at their core, enabling long-term platform success. Leverage customer insights, operational metrics and market intelligence to inform product strategy, prioritisation and continuous improvement. What ...

Senior SRE: Kubernetes, GitOps & AI-Driven Reliability

Location
City Of London, England, United Kingdom
canary deployments, and secure secret management within a hybrid London-based team. Working in the Webex Engineering Group, you will advance AI-assisted configuration, observability, and reliable delivery across dev, staging, and production environments. #J-18808-Ljbffr ...

Lead Data Engineer - Start-Up / Scale-up experience required

Hiring Organisation
FBI &TMT
Location
City of London, London, United Kingdom
Employment Type
Permanent
workloads Deliver batch and real-time/streaming pipelines Build and optimise vector search and retrieval pipelines that support AI features Implement monitoring, observability and data quality frameworks Drive security, reliability and governance across the data layer Partner closely with backend and AI engineers to support training, inference and product ...

Platform Engineer

Location
City Of London, England, United Kingdom
overhead. Support and enhance CI/CD platforms and engineering workflows, including the safe promotion of infrastructure and platform changes. Improve platform reliability, security, observability, performance and operational excellence. Partner with software engineers, quantitative developers, research teams, Technology Operations and security colleagues to deliver secure, scalable and resilient platforms. Education … experience with AWS and cloud‐native infrastructure. Experience with Kubernetes, including Amazon EKS, and Docker or other container runtimes. Experience with monitoring and observability tools such as Grafana, Prometheus, Loki or Datadog. Experience with artifact‐management platforms such as JFrog Artifactory. Knowledge of AWS Batch, AWS Step Functions, AWS Identity ...

Scala Developer - Remote Contract

Hiring Organisation
Stealth iT Consulting
Location
City of London, London, United Kingdom
environment (Scrum/Kanban). Participate in code reviews, architecture discussions, and pair programming sessions. Troubleshoot and resolve production issues; contribute to reliability and observability (logging, metrics, alerts). Assist in defining CI/CD pipelines and deployment processes (e.g., Jenkins, GitHub Actions, Concourse). Produce concise technical documentation ...