2,226 to 2,250 of 5,621 Observability Jobs

Principal Architect AI

Location
Greater London, England, United Kingdom
production AI solutions. Partner with data science teams to translate complex forecasting models and machine learning algorithms into production-ready, scalable architectures. Design robust observability and evaluation frameworks for GenAI systems, including system guardrails, continuous monitoring, and optimisation mechanisms. Your Skills Experienced background in software architecture, highlighted by a dedicated ...

Senior Software Engineer, Payments & Incentives

Location
Greater London, England, United Kingdom
relatively speedily. You have experience designing and implementing APIs and microservices in a high-velocity production environment. You possess a deep understanding of observability and the full software development life cycle (SDLC) in an Agile environment. You can effectively distill complex technical updates into clear, actionable information for both technical ...

Global IT Senior Director and Platform Team Lead - Gen AI Platforms and Agentic Development Kit

Hiring Organisation
The Boston Consulting Group
Location
London, UK
Employment Type
Full-time
delivery. Delivery & Operational ExcellenceOwn end to end delivery of capabilities & components and services owned by the platform teams. Establish and run enterprise‐grade AI observability frameworks to monitor performance, safety, and interpretability. Continuously enhance LLM Ops, Evaluation and Quality pipelines. Ensure robust AIOps/DevSecOps, CI/CD, Infrastructure ...

Head of Release - Cambridge

Location
Milton, England, United Kingdom
cloud-hosted and hybrid platforms, including technologies such as AWS, Kubernetes, OpenShift, EKS, EC2, Helm, Jenkins and Terraform.Ensures effective monitoring, alerting, logging and observability are in place so deployment progress, service health and customer impact can be understood quickly.Oversees release communications before, during and after deployment, ensuring stakeholders receive timely ...

Principal Cloud Engineer

Location
Greater London, England, United Kingdom
teammates’ work quickly. Our small team regularly deploys over a dozen times a day. Yes, we ship on Fridays. We are hooked on observability: we strive to get visibility over all parts of our system to make incident response as painless as possible and to know what needs to improve. ...

Java Lead Software Engineer

Location
Glasgow, Scotland, United Kingdom
standards Solid understanding of the Software Development Life Cycle and agile methodologies Familiarity with continuous integration and delivery, application resiliency, security best practices, and observability and monitoring Hands‐on experience using enterprise-authorized AI‐assisted software development tools with demonstrated ability to critically evaluate and validate AI-generated outputs Understanding ...

Lead Software Engineer - Java / Python - Equity Derivatives - Front Office Quant Developer

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
assisted development and automation capabilities, to maximize automation valueMaintain strong release discipline, including regression assessment, rollback/fallback planning, post-deployment verification, and observability improvementsGather and synthesize data and telemetry to develop reporting and metrics that improve stability, quality, and delivery predictabilityIdentify hidden failure patterns in production and drive improvements ...

Lead Software Engineer - FIXED INCOME UK

Location
Greater London, England, United Kingdom
. Preferred qualifications, capabilities, and skills In‐depth knowledge of the financial services industry and their IT systems Practical cloud native experience Experience with observability tooling and practices (structured logging, metrics, distributed tracing) and incident/problem management. Work Style/Ways of Workin g Strong communication skills, ownership mindset ...

Head of Release - Cambridge

Location
Cambridge, England, United Kingdom
cloud-hosted and hybrid platforms, including technologies such as AWS, Kubernetes, OpenShift, EKS, EC2, Helm, Jenkins and Terraform. Ensures effective monitoring, alerting, logging and observability are in place so deployment progress, service health and customer impact can be understood quickly. Oversees release communications before, during and after deployment, ensuring stakeholders ...

Specialist Solutions Architect (SSA) (Cloud Infrastructure)

Location
Greater London, England, United Kingdom
Private Link/GCP Private Service Connect), network routing, performance optimisation, and large-scale deployment management Platform Administration: High availability, disaster recovery, cluster orchestration, observability and audit (e.g. Amazon CloudWatch/CloudTrail, Azure Monitor, Google Cloud Operations Suite), and cloud cost management Infrastructure Automation (InfraOps): Hands‐on automation using ...

Databricks Architect - Lead/Principal Consultant

Location
Greater Manchester, England, United Kingdom
Engineering Fundamentals: Bring rigour to how we model, test, deploy and operate data platforms: dimensional and medallion modelling, CI/CD, infrastructure as code, observability, cost management and data quality by design. 1. Architectural Vision & Technical Ownership Lead from the Front: Build and lead high-performing delivery teams, setting direction ...

Infrastructure & Platform Senior Specialist Solutions Engineer

Hiring Organisation
DataBricks
Location
London, UK
Employment Type
Full-time
Azure Private Link/GCP Private Service Connect), network routing, performance optimisation, and large-scale deployment managementPlatform Administration: High availability, disaster recovery, cluster orchestration, observability and audit (e.g. Amazon CloudWatch/CloudTrail, Azure Monitor, Google Cloud Operations Suite), and cloud cost managementInfrastructure Automation (InfraOps): Hands-on automation using IaC tools ...

Infrastructure & Platform Senior Specialist Solutions Engineer

Location
Greater London, England, United Kingdom
Private Link/GCP Private Service Connect), network routing, performance optimisation, and large-scale deployment management Platform Administration: High availability, disaster recovery, cluster orchestration, observability and audit (e.g. Amazon CloudWatch/CloudTrail, Azure Monitor, Google Cloud Operations Suite), and cloud cost management Infrastructure Automation (InfraOps): Hands‐on automation using ...

Azure DevOps Engineer

Hiring Organisation
IMSERV EUROPE LIMITED
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent
privilege across the estate, and platform patching. Work with IT security on audit remediation and the security release plan. Monitoring, reliability and recovery. Own observability (Application Insights, Log Analytics), alerting, backups and disaster recovery, including periodic DR tests. Lead incident response for platform issues and run blameless root-cause reviews. ...

Associate, Java Engineer, BGM Data & Analytics

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
applications, including real-time streaming with high volume and concurrent workloads. Presenting designs to technical reviewers, contributing to reviewing designs/code Improving reliability, observability, and performance; Supporting production releases to help team meet SLAs. The role is within BGMs In-Business technology group in London. We are a community ...

AI Infrastructure Lead Architect

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
/CD strategies for reliable, repeatable, and scalable releases of AI systems, models, and data pipelines into production. Establish AI monitoring and observability strategy acrossInfraOpsandMLOps, defining SLAs, SLOs, alerting, and performance/cost tracking, and driving continuous optimization. Integrate AI/ML systems into enterprise environments, ensuring interoperability, security, compliance ...

Principal Developer - AI

Location
Salford, England, United Kingdom
lifecycle management on Kubernetes across AWS, GCP, and Azure. Establish and evangelize engineering best practices across services, including standards for scalability, fault tolerance, observability, and operational excellence. Mentor senior engineers and partner closely with data scientists to elevate technical quality and shorten the path from research to production. Partner with ...

Principal QA Engineer

Location
Witham, England, United Kingdom
built high-performing functions. Designed and implemented test automation frameworks, cloud-based testing tools and continuous testing solutions. Experience in performance testing, security testing, observability and non-functional assurance. Proven experience applying AI-driven, predictive, or intelligent testing capabilities. Worked within Agile, DevOps, Continuous Delivery, TDD and BDD environments. Influenced ...

Staff Software Engineer - Agentic Workflows

Hiring Organisation
Causaly
Location
London, UK
Employment Type
Full-time
built yourself. Python and LLM application frameworks (FastAPI, LangChain, LangGraph). Our GenAI service is written in Python and you will work in it. Observability and LLM instrumentation (OpenTelemetry, Langfuse). Multi-tenant SaaS where tenant isolation is a design constraint. Infrastructure as code (Terraform, Kubernetes). Search and retrieval ...

Senior Java Systems Engineer - Platform Exchange / Matching Engine

Hiring Organisation
Adaptive Financial Consulting Ltd
Location
Greater London, United Kingdom
Employment Type
Full Time
Chronicle, or LBM. An understanding of replicated state machines, event sourcing, or deterministic systems. Exposure to cloud environments (AWS/GCP), containers, Kubernetes, or observability tooling. YOU WILL: Design, build, and optimise low-latency Java services and protocols for an Exchange Accelerator product across the delivery lifecycle. Define order, session ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
joiner/mover/leaver Self-hosting and operating shared internal tooling that other teams depend on - deploy, upgrade, back up, own the uptime Observability and alerting design, plus on-call for a distributed team Cloud cost management/FinOps fundamentals Azure is a nice-to-have extension ...

Software Engineer III - Data Analytics Platform

Location
London, United Kingdom
degradation, autoscaling behaviors, incident follow-ups, and runbooks. Contribute to system design by breaking down ambiguous problems, proposing approaches, and making pragmatic tradeoffs. Add observability with metrics, tracing, logging, dashboards, and actionable alerts tied to SLOs. Support safe deployments through CI/CD improvements, canarying, feature flags, backward compatibility ...

Software Engineer III - Data Analytics Platform

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
degradation, autoscaling behaviors, incident follow-ups, and runbooks. Contribute to system design by breaking down ambiguous problems, proposing approaches, and making pragmatic tradeoffs. Add observability with metrics, tracing, logging, dashboards, and actionable alerts tied to SLOs. Support safe deployments through CI/CD improvements, canarying, feature flags, backward compatibility ...

Staff Software Engineer, Inference

Location
Greater London, England, United Kingdom
grade deployment pipelines for releasing new models to millions of users reliably Contributing to new inference features Supporting inference for new model architectures Analyzing observability data to tune performance based on real-world production workloads Managing multi-region deployments and geographic routing for global customers Deadline to apply: None. Applications ...

Data & AI Solution Architect – Generative AI Products & Delivery

Location
Greater London, England, United Kingdom
delivery, DevOps, MLOps, and LLMOps practices to enterprise AI programmes. Strong cloud architecture experience across Azure, AWS, or GCP. Experience in model deployment, monitoring, observability, governance, versioning, and operational support. Excellent stakeholder management skills with proven experience engaging clients and senior business leaders. Strong understanding of enterprise security, risk management ...