76 to 100 of 118 Observability Jobs in the City of London

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£780 - £830 per day
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone … shape how thousands of infrastructure and application events are detected, correlated and actioned across EMEA, and you'll have genuine scope to modernise observability capability rather than simply keep the lights on. This is a hands-on technical leadership role with no direct reports, so your influence comes from your ...

AI Metrics & Model Evaluation Lead

Location
City Of London, England, United Kingdom
metrics and model evaluation for cutting‐edge AI products used in highly regulated industries. You’ll define metrics, build dashboards, and ensure observability of product performance from user interaction to model outputs. You will partner with AI leadership, Product and Engineering to drive data‐informed decisions, create golden datasets ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...

Senior Software Engineer / Senior AI Engineer

Location
City Of London, England, United Kingdom
within a defined problem, building and testing tool use, retrieval pipelines, and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability, and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure … agentic AI solutions to production standards within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability, and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

ClickHouse Platform Architect - Greenfield, High-Scale

Location
City Of London, England, United Kingdom
Colehouse Group is seeking a ClickHouse Solutions Architect to build a greenfield, enterprise-scale observability platform for a global banking client. You will drive the architectural decisions from first principles to meet performance, multi-tenancy and retention requirements. The role covers discovery, data-modeling, ingestion, storage tiering and security, across ...

AI Metrics & Model Evaluation Analyst

Location
City Of London, England, United Kingdom
products are used, how models perform, and what changes drive business value in regulated environments. In this role you’ll define metrics, build observability, create dashboards, and partner with Product and Engineering to ensure meaningful, data-driven release decisions that advance AI capabilities. #J-18808-Ljbffr ...

Strategic Global Enterprise Account Director

Location
City Of London, England, United Kingdom
will act as a Strategic Hunter and Executive Orchestrator, rebuilding executive relationships and shaping multi-year pipelines across Cisco’s networking, security, cloud, observability, and services portfolio. Responsibilities include reactivating dormant accounts, architecting 12–24 month #J-18808-Ljbffr ...

Principal Engineer I, Prepurchase Platform (Remote)

Location
City Of London, England, United Kingdom
demand on-sales. You will write production code daily, influence technical direction, and collaborate across multiple teams within the Prepurchase domain. You will drive observability, resilience patterns, and AI-assisted enhancements while embedding across services or working horizontally. #J-18808-Ljbffr ...

Core Platform Developer

Location
City Of London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment. Key Responsibilities Design and build shared backend services, frameworks … developer tooling that support internal application and service development. Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities. Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams. Help ...

MLOps Engineer

Hiring Organisation
DGH Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
platform reliability. Key Responsibilities - Design, deploy, and manage AI platforms and agent infrastructure - Build and maintain CI/CD pipelines and DevOps workflows - Implement observability, monitoring, and logging solutions - Optimise performance, scalability, and cost efficiency - Support AI teams with infrastructure, deployment, and integration - Ensure platform security, compliance, and high availability …/CD, automation, and DevOps best practices - Experience with Kubernetes/containerisation technologies - Strong programming skills (e.g. Python, Go, Node.js) - Experience with observability tools (e.g. OpenTelemetry, Datadog) - Understanding of security, performance optimisation, and scalability Desirable Skills - Experience working on AI/ML platforms or deployments - Exposure to large-scale distributed ...

Technical Leader

Location
City Of London, England, United Kingdom
manage technical debt, prioritizing improvements that provide meaningful value to the team and the product. Promote engineering best practices around testing, CI/CD, observability, documentation, and operational excellence. Stay current with emerging technologies and industry practices, evaluating when new technologies can provide meaningful improvements to our products and engineering … computer science fundamentals, including data structures, algorithms, concurrency, and system design. Strong understanding of modern software development practices, including CI/CD, TDD, DevOps, observability, and automated quality gates. Confident communication and collaboration skills, with the ability to articulate technical decisions, challenge assumptions, and explain complex technical concepts to both ...

Cloud Platforms Engineer

Location
City Of London, England, United Kingdom
runs the internal platform that PEI’s engineering and data teams build on: the cloud accounts, the reusable infrastructure code, the delivery pipelines, the observability, and the security and cost guardrails that wrap around them. We treat that platform as a product with internal customers – the measure of our work … Desirable: experience with data platform infrastructure such as Databricks, or similar – prior Databricks experience is not required. Desirable: experience with Datadog, or another mature observability platform. Desirable: experience of high‐traffic, international, content‐heavy web platforms. Technical Skills Confident with Linux and containers, and able to debug from the command ...

DevOps Lead

Hiring Organisation
TurleyWay Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
manage and optimise containerised environments using Kubernetes. Oversee source control, CI/CD pipelines and development workflows within GitHub and build and maintain monitoring, observability and alerting solutions using Grafana. To be considered you be able to demonstrate proven experience in a DevOps Lead, Senior DevOps Engineer, or Platform Engineering … expertise in Kubernetes and container orchestration technologies, administering and managing GitHub and CI/CD pipelines. Solid knowledge of Grafana and modern monitoring/observability practices and extensive database exposure, including performance, administration, and optimisation of enterprise database environments. In return we offer a competitive basic salary plus bonus scheme ...

DevOps Manager

Hiring Organisation
TurleyWay Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
manage and optimise containerised environments using Kubernetes. Oversee source control, CI/CD pipelines and development workflows within GitHub and build and maintain monitoring, observability and alerting solutions using Grafana. To be considered you be able to demonstrate proven experience in a DevOps Lead, Senior DevOps Engineer, or Platform Engineering … expertise in Kubernetes and container orchestration technologies, administering and managing GitHub and CI/CD pipelines. Solid knowledge of Grafana and modern monitoring/observability practices and extensive database exposure, including performance, administration, and optimisation of enterprise database environments. In return we offer a competitive basic salary plus bonus scheme ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Platform Principal Engineer

Location
City Of London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

Agentic AI Enterprise Architect

Hiring Organisation
HCLTech
Location
City of London, London, United Kingdom
harnesses and test suites that measure agent correctness, safety and regression before anything ships. Own AgentOps/DevSecOps: CI/CD for agents, versioning, observability and telemetry, shift-left security, and Responsible AI governance baked in from day one. Run a continuous, adaptable feedback loop: feed production telemetry, evals … Eval-driven development: designing evaluation harnesses and measuring agent quality, safety and reliability. Standards-based integration and DevSecOps: APIs, secure auth, CI/CD, observability and AgentOps. Ability to conceptualize a business problem as an agent quickly, and operate effectively in ambiguous, customer-embedded settings. Client-facing maturity: translates fluidly ...

Senior UI Platform Engineer, London

Location
City Of London, England, United Kingdom
developer workflows Help define standards for application structure, testing, deployment, and supportability Partner with other engineering teams on shared concerns such as authentication, authorization, observability, and frontend/backend integration patterns Contribute to internal platform applications such as onboarding, access administration, and service discovery experiences Evaluate and maintain third-party … internal tooling Experience with Vite, Next.js, or both Experience with UI/component libraries such as Ant Design, AG Grid, or similar Experience with observability tooling such as Datadog, OpenTelemetry, or Grafana Experience with Microsoft Entra ID/Azure AD or similar identity platforms Experience publishing and maintaining internal ...

Senior Pre-Sales Consultant, Enterprise Payments

Location
City Of London, England, United Kingdom
workshops, both virtually and on-site. Lead and contribute to RFI and RFP responses, providing detailed technical input across areas including: Integration Security Resiliency Observability Compliance Identify risks, gaps, constraints, and assumptions while clearly articulating solution trade-offs. Plan, design, and support Proof of Concept (POC) activities, enabling customers … identity frameworks, including OpenID Connect and OAuth Cloud platforms such as: AWS Microsoft Azure Google Cloud Platform (GCP) IBM Cloud Oracle Cloud Monitoring and observability tools, including Splunk and SIEM solutions Pre-Sales and Customer-Facing Experience Proven experience in a customer-facing technical role, solution architecture role ...

Dataiku Solution Architect

Hiring Organisation
Everforth Quinnox
Location
City of London, London, United Kingdom
Employment Type
Permanent
dashboards use controlled, reconciled, and traceable data from the governed platform. Define access, refresh, performance, lineage, and reconciliation standards for reporting solutions. Support curve observability and the monitoring of data quality, source availability, processing status, and workflow completion. Nonfunctional Architecture Define nonfunctional requirements for performance, scalability, security, availability, resiliency, recoverability … observability, maintainability, and supportability. Design monitoring and alerting across Dataiku, Power Automate, PostgreSQL or Amazon RDS, integrations, WebApps, and reporting components. Establish recovery patterns for failed source deliveries, workflow errors, data-quality issues, integration failures, and interrupted processing. Define capacity and performance considerations for regional processing, historical replay, concurrent users ...

Backend Java Developer – Data Fabric / Platform Engineering

Location
City Of London, England, United Kingdom
platforms/query engines (e.g., Starburst or similar) Own API contracts with living documentation in CI/CD Build production-grade, testable pipelines Drive observability, reliability, and performance Contribute to architecture decisions (modularity, DI, extensibility) What You Bring (Must-Have) Strong hands-on experience in Java (17/21) + … backend performance Production-grade testing using JUnit 5, Mockito Experience with clean architecture, DI, modular design Comfortable owning CI/CD, code quality, observability Familiarity with Docker, Maven, Jenkins ⭐ Nice to Have Apache Calcite Starburst or federated query engines JVM performance tuning High-throughput service interfaces (REST/gRPC) Data ...