101 to 125 of 144 Observability Jobs in Central London

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Senior Frontend Engineer — Form-Heavy React/TS

Location
City Of London, England, United Kingdom
onboarding clients and accounts with KYC/KYB workflows. You will design, implement, and own features end-to-end, ensuring security, reliability, and observability from day one. Strong collaboration and clear communication with engineers and non-engineers are essential. #J-18808-Ljbffr ...

Data Platform Engineer

Location
City Of London, England, United Kingdom
consistency and reusability across environments. Build and optimize CI/CD pipelines using Azure DevOps and GitHub Actions to support rapid, reliable deployments. Implement observability practices including logging, metrics, and alerting using observability tools. Collaborate with the Lead Engineer and Architects to align implementation with platform standards and patterns. Provide … Fabric. Proven experience with infrastructure-as-code using Terraform and building CI/CD pipelines via Azure DevOps and GitHub Actions. Strong grasp of observability practices, including logging, metrics, alerting, and performance optimization. Deep understanding of cloud security, with experience applying secure-by-design principles in Azure and/ ...

Director of Software Engineering (AIOps) - Executive Director

Location
City Of London, England, United Kingdom
reins and drive impact, we’ve got an opportunity just for you. As a Director of Software Engineeringat JPMorgan Chase within theEngineer's Observability Platforms team, you lead a technical area and drive impact within teams, technologies, and projects firm wide. Utilize your in-depth knowledge of software, applications, technical … troubleshooting root cause analysis finding for application support teams. Job responsibilities Leads technology and process implementations to achieve functional technology objectives in the Observability Platforms space, providing essential services for Site Reliability Engineers, Operations and Engineers across the whole firm Innovates, designs and delivers technical solutions that can be leveraged ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

Lead Site Reliability Engineer

Location
Westminster, West End, United Kingdom
undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast paced front office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...

Principal Engineer I, Prepurchase Platform (Remote)

Location
City Of London, England, United Kingdom
demand on-sales. You will write production code daily, influence technical direction, and collaborate across multiple teams within the Prepurchase domain. You will drive observability, resilience patterns, and AI-assisted enhancements while embedding across services or working horizontally. #J-18808-Ljbffr ...

Site Reliability Engineer

Location
City Of London, England, United Kingdom
global trading and investment operations. Working closely with software engineers, quantitative researchers, traders, and infrastructure teams, you will be responsible for building automation, improving observability, and ensuring critical production systems operate at the highest levels of availability and efficiency. Key Responsibilities Design, build, and maintain highly reliable, scalable, and automated … infrastructure platforms. Drive improvements in system performance, monitoring, observability, and operational efficiency. Troubleshoot and resolve complex production incidents across distributed systems. Develop tools and automation to reduce operational overhead and improve platform resilience. Partner with engineering teams to improve system design, deployment processes, and operational readiness. Participate in incident management ...

ML Ops Engineer

Location
City of Westminster, England, United Kingdom
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services. You'll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams. This is not a research role. … model metadata and reproducibility across research and production. Build reusable tooling and platform capabilities that support multiple models and engineering teams. Model serving and observability Deploy and operate batch and online inference services in containerised cloud environments. Define and meet availability, latency, throughput and recovery objectives for ML services. Monitor ...

Core Platform Developer

Location
City Of London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment. Key Responsibilities Design and build shared backend services, frameworks … developer tooling that support internal application and service development. Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities. Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams. Help ...

MLOps Engineer

Hiring Organisation
DGH Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
platform reliability. Key Responsibilities - Design, deploy, and manage AI platforms and agent infrastructure - Build and maintain CI/CD pipelines and DevOps workflows - Implement observability, monitoring, and logging solutions - Optimise performance, scalability, and cost efficiency - Support AI teams with infrastructure, deployment, and integration - Ensure platform security, compliance, and high availability …/CD, automation, and DevOps best practices - Experience with Kubernetes/containerisation technologies - Strong programming skills (e.g. Python, Go, Node.js) - Experience with observability tools (e.g. OpenTelemetry, Datadog) - Understanding of security, performance optimisation, and scalability Desirable Skills - Experience working on AI/ML platforms or deployments - Exposure to large-scale distributed ...

Software Engineer - Cloud Compute Platform

Location
City of Westminster, England, United Kingdom
contribute to include: Workload orchestration across Kubernetes clusters Platform APIs and Kubernetes operators Cloud platform integrations Multi-tenant workload isolation and security Observability, health checks, and operational tooling Workload scheduling, disruption management, and rolling updates What You Will Do: Design and implement services, APIs, and Kubernetes controllers using Go. Contribute … engineers to understand requirements and deliver reliable solutions. Participate in technical design discussions and help evaluate implementation trade-offs. Improve the reliability, scalability, observability, and maintainability of existing systems. Write automated tests, documentation, and operational runbooks. Participate in code reviews and provide constructive feedback to teammates. Help investigate and resolve ...

Senior Associate, Full-Stack Engineer

Location
Westminster, West End, United Kingdom
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile (Scrum ...

Senior Software Engineer

Location
City Of London, England, United Kingdom
Make the architectural calls on your domain, write them down, and defend them in front of the tribe. Build in security, availability, reliability and observability from the start, and instrument your services so the answer to what's happening is already in a dashboard. Use AI tooling well. … awkward parts that come with them: idempotency, ordering, retries, poison messages, exactly-once as a promise nobody can keep. Security, availability, reliability and observability built in from the start, not bolted on at the end. You think about who can reach what, you know how your service behaves when ...

DevOps Lead

Hiring Organisation
TurleyWay Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
manage and optimise containerised environments using Kubernetes. Oversee source control, CI/CD pipelines and development workflows within GitHub and build and maintain monitoring, observability and alerting solutions using Grafana. To be considered you be able to demonstrate proven experience in a DevOps Lead, Senior DevOps Engineer, or Platform Engineering … expertise in Kubernetes and container orchestration technologies, administering and managing GitHub and CI/CD pipelines. Solid knowledge of Grafana and modern monitoring/observability practices and extensive database exposure, including performance, administration, and optimisation of enterprise database environments. In return we offer a competitive basic salary plus bonus scheme ...

DevOps Manager

Hiring Organisation
TurleyWay Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
manage and optimise containerised environments using Kubernetes. Oversee source control, CI/CD pipelines and development workflows within GitHub and build and maintain monitoring, observability and alerting solutions using Grafana. To be considered you be able to demonstrate proven experience in a DevOps Lead, Senior DevOps Engineer, or Platform Engineering … expertise in Kubernetes and container orchestration technologies, administering and managing GitHub and CI/CD pipelines. Solid knowledge of Grafana and modern monitoring/observability practices and extensive database exposure, including performance, administration, and optimisation of enterprise database environments. In return we offer a competitive basic salary plus bonus scheme ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Software Engineer, AI Product Engineering

Location
City of Westminster, England, United Kingdom
dates, restatements — and keep the model honest as the product grows. Own your services end to end: implementation, tests, CI/CD, deployment, observability, and the on‐call that comes with them. Collaboration and Technical Leadership Work directly with product leadership to shape what gets built, challenge assumptions, and turn … having shipped something and watched it behave in production. Comfort with modern delivery practice: automated testing, CI/CD, containers and Kubernetes, and observability as a first‐class concern. Strong business sense. You understand who uses what you build and why it matters to them, and you make trade‐offs ...

Product Associate - SRE Team - Chase UK

Location
Westminster, West End, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Platform Principal Engineer

Location
City Of London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

Agentic AI Enterprise Architect

Hiring Organisation
HCLTech
Location
City of London, London, United Kingdom
harnesses and test suites that measure agent correctness, safety and regression before anything ships. Own AgentOps/DevSecOps: CI/CD for agents, versioning, observability and telemetry, shift-left security, and Responsible AI governance baked in from day one. Run a continuous, adaptable feedback loop: feed production telemetry, evals … Eval-driven development: designing evaluation harnesses and measuring agent quality, safety and reliability. Standards-based integration and DevSecOps: APIs, secure auth, CI/CD, observability and AgentOps. Ability to conceptualize a business problem as an agent quickly, and operate effectively in ambiguous, customer-embedded settings. Client-facing maturity: translates fluidly ...

Platform Operations Director

Hiring Organisation
ClearCourse
Location
City of London, London, United Kingdom
business continuity across the group. Internal IT & Systems Manages internal IT and business systems administration (M365, NetSuite, SuccessFactors, SharePoint) -infrastructure, integrations, and IAM. Ensures observability and SRE capability is fit for purpose across cloud, hosted, and end-user environments. Vendor & Cost Management Drives cloud and vendor cost discipline - manages …/CD infrastructure requirements. • Head of Infrastructure & Cloud - Direct report. Hosting strategy, cloud platform, and FinOps execution. Head of SRE - Direct report. Observability, on-call, and DR/BCP processes. • Head of Internal Services - Direct report. Internal IT, business systems, and end-user support. Finance - Direct report. Cloud cost visibility ...

Senior UI Platform Engineer, London

Location
City Of London, England, United Kingdom
developer workflows Help define standards for application structure, testing, deployment, and supportability Partner with other engineering teams on shared concerns such as authentication, authorization, observability, and frontend/backend integration patterns Contribute to internal platform applications such as onboarding, access administration, and service discovery experiences Evaluate and maintain third-party … internal tooling Experience with Vite, Next.js, or both Experience with UI/component libraries such as Ant Design, AG Grid, or similar Experience with observability tooling such as Datadog, OpenTelemetry, or Grafana Experience with Microsoft Entra ID/Azure AD or similar identity platforms Experience publishing and maintaining internal ...