601 to 625 of 2,337 Observability Jobs in London

Director of Platform Engineering - SaaS & Observability

Location
Greater London, England, United Kingdom
ITRS is seeking a Director of Platform Engineering to lead a hands-on, high-availability SaaS platform that serves global financial institutions. You will supervise the Analytics SaaS stack, guide architecture decisions, and drive automation ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams.It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling.**Responsibilities*** Lead the design and evolution of scalable, secure, and highly available trading infrastructure.* Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows.* Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering.* Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices.* Lead and participate ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
monitor, diagnose, and auto-remediate global SaaS infrastructure. What you'll do Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands‐on experience building LLM pipelines, AI Agents, Model ...

Software Engineer

Location
Greater London, England, United Kingdom
diagnose, and auto-remediate global SaaS infrastructure. What You'll Do Technical Design & Architecture: Design and implement high-resilience software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design, build, and maintain production-grade AI agents, MCP tool integrations, and deterministic evaluation … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated root cause analysis (RCA). Preferred Qualifications AI & Agentic Systems: E xperience building LLM pipelines ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands‐on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement. We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve … teams to productionize AI solutions and ensure services are ready to operate reliably at scale. Establish engineering standards and reusable patterns for infrastructure, security, observability, resilience, documentation, and operational readiness. Lead architectural decisions and evaluate trade-offs across reliability, security, scalability, performance, cost, and maintainability. Take ownership of operational risks ...

DevOps Engineer, Blockchain Infra (Fully Remote)

Hiring Organisation
Binance
Location
London, UK
Employment Type
Full-time
pipelines for application and infrastructure deployment. Deploy and operate middleware platforms such as Kafka, Redis and NGINX.Automate operational tasks using Golang and Python. Improve observability through monitoring, logging, alerting, and distributed tracing. Ensure platform reliability, scalability, security, and disaster recovery. Troubleshoot production incidents and perform root cause analysis. Work closely …/Redis/NGINXStrong scripting and programming skills in: Golang/PythonExperience with Git, GitOps workflows, and CI/CD platforms. Strong understanding of observability tools such as Prometheus, OpenTelemetry. Familiarity with container technologies including Docker and Kubernetes. Strong troubleshooting and problem-solving skills. Preferred QualificationsExperience operating blockchain infrastructure ...

Senior SRE

Hiring Organisation
Pigment
Location
London, UK
Employment Type
Full-time
platform usage. Secure high availability and redundancy of the Pigment platform. Ensure that the platform's performance and correctness are monitored accordingly, spread observability best practices across the engineering team. Participate in incident response. Accompany Pigment geographical expansion as the company grows and we sign clients overseas. Work with … experience in software developments with languages such as C#, Java, C++, Golang, Rust, JavaScript, Python, or Ruby (this list is not exhaustive).Experience with observability tools (e.g. Datadog, Prometheus, ELK, Jaeger...)Great team spirit with a problem-solving attitude. A good dose of humility and the willingness to grow ...

AI Platform Engineer- Senior Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
Implement MLOps and LLMOps pipelines (model deployment, monitoring, retraining, and fine-tuning where relevant) using Infrastructure-as-Code, GitOps, and CI/CD• Establish observability, security, and governance frameworks specific to AI systems, including cost attribution and lifecycle management• Work with clients and internal teams to develop new opportunities … equivalent), including familiarity with the Model Context Protocol (MCP)• Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)• Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust)**MLOps & LLMOps**• Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone ...

Data Technical Lead

Hiring Organisation
PA Consulting
Location
London, UK
Employment Type
Full-time
Data governance: Catalogue, metadata, lineage, quality, security, privacy and access controls Data products: Reusable, discoverable and well-governed data products and marketplaces DataOps and observability: Testing, monitoring, operational controls and platform reliability Platform engineering: Terraform, CloudFormation, Azure Bicep and infrastructure-as-code CI/CD: GitHub Actions, Azure DevOps, Jenkins … serving. Applying strong software engineering practices to data platforms, including testing, CI/CD and infrastructure-as-code. Establishing effective approaches to data quality, observability, governance, metadata and security. You can lead data platform modernisation and migration, including coexistence, cutover and decommissioning. Making pragmatic technology choices and understanding the trade ...

Observability SRE

Location
Greater London, England, United Kingdom
want you to find your spark. Because that’s what drives you to be better, be more and ultimately, be more fulfilled. Job Title: Observability SRE Location: London, UK Employment Type: Fixed term contract (12 months duration) Job type: Onsite Job/Group Overview: SRE within the Group Platform Services …/support position responsible for administering and supporting Production environment as well as engineering reliability into the products/services we support i.e. monitoring & observability platform. The successful candidate will have a vital role in shaping future monitoring strategy and direction. A fantastic opportunity for somebody with 3+ years ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic … technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit ...

Senior Lead Data Platform Engineer - Java/Python

Location
Greater London, England, United Kingdom
develop secure, scalable production code in Java, Python, or other modern languages, and collaborate with agile teams to advance cloud-native data platforms and observability practices across the organization. #J-18808-Ljbffr ...

Remote Platform Engineer: AI, Cloud & CI/CD Automation

Location
Greater London, England, United Kingdom
design and operate pipelines, IaC, and cloud services across AWS/Azure and Kubernetes, aligning with GitOps and security best practices. You will implement observability, self‐service tooling, and automation while collaborating with developers to ship reliable software and AI workloads at scale. #J-18808-Ljbffr ...

Senior Network SRE: Cloud Reliability & IaC

Location
Greater London, England, United Kingdom
focus on cloud automation, IaC, and governance across our AWS infra, contributing to highly available services for millions of users. You will own automation, observability, and incident response, driving cross-functional projects in an Agile environment. Strong Linux skills, Python/Bash, and Terraform expertise are essential for success. #J ...

Global Cloud Networking & Security Architect

Location
Greater London, England, United Kingdom
multi-cloud strategy, implement Zero Trust, and build automation platforms using Terraform, Python, APIs, and CI/CD pipelines, advancing AI-enabled operations and observability across the #J-18808-Ljbffr ...

Contract Machine Learning Engineer - Production AI

Location
Greater London, England, United Kingdom
deploy and operate AI/ML architectures in production. The role emphasizes building high-volume batch systems, real-time microservices, automated retraining, and strong observability across cloud infrastructure. You will leverage Python, SQL, PyTorch, TensorFlow, and Scikit-learn, with Google Cloud Platform tools like BigQuery, Vertex AI, and Dataflow, plus ...

Senior Full-Stack Engineer, Industrial Emissions Platform

Location
Greater London, England, United Kingdom
mentor a multinational team in a hybrid work model. You will work across cloud platforms (Azure/AWS/GCP), implement CI/CD, observability, and security practices, and shape product delivery from discovery through production. #J-18808-Ljbffr ...

Application Engineer

Location
Greater London, England, United Kingdom
Experience working with financial services, customer facing digital platforms, product environment, KYC or highly regulated firms. Desired Skills Adobe Experience Manager (AEM) AWS certification Observability & monitoring tools #J-18808-Ljbffr ...

MLOps Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
Build and maintain orchestration pipelines using Dagster, Airflow, or Prefect. Deploy and manage ML workloads on Kubernetes and cloud platforms. Implement CI/CD, observability, and monitoring across production systems. Develop infrastructure using Terraform, Bicep, or similar IaC tools. Work closely with Data Scientists, ML Engineers, and AI teams ...

DevOps Engineer - SC Cleared

Hiring Organisation
CBSbutler Holdings Limited trading as CBSbutler
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £520/day
Build and manage AWS infrastructure using Terraform Automate environment provisioning and routine operational processes Maintain Docker-based deployment and runtime patterns Improve platform resilience, observability and security Lead incident investigation, remediation and platform upgrades If you have strong AWS, Terraform, Jenkins and DevOps experience and hold SC Clearance ...

AI Full Stack Engineer

Hiring Organisation
Randstad Technologies Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£350 - £388/day
Agentic Systems: Build multi-agent, stateful, and asynchronous workflows. LLM Management: Handle prompt engineering, context/memory management, hallucination controls, guardrails, and fallback logic. Observability: Implement tracing and evaluation frameworks, specifically utilizing LLM-as-a-Judge patterns. Backend Engineering: Develop secure Python APIs and maintain strong testing and deployment disciplines. ...

Engineering Lead AWS Platform, Infrastructure & Operations

Hiring Organisation
Interact Consulting Limited
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
developer experience. The role: Lead and develop a team of 5-7 Operations Engineers. Shape the AWS, infrastructure and DevOps strategy. Drive reliability, security, observability and operational excellence. Champion “You Build It, You Run It” engineering practices. Drive cloud cost optimisation and FinOps. Partner with Engineering, Product and Security ...