476 to 500 of 1,507 Observability Jobs

Distinguished AI Engineer

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more. Invent ...

Lead Backend Engineer (Routing Squad)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Technologies we use Backend languages: Python, Go Tech infrastructure: AWS, CDK TypeScript, Lambda, SQS, EventBridge, RDS, DynamoDB Data tooling: GCP, BigQuery, Looker, Looker Studio Observability: Loki, Tempo, Grafana, Prometheus Event‐driven architecture and domain‐driven design Our interviewing process Intro call with the hiring manager Live coding challenge solving ...

Staff Machine Learing Engineer

Hiring Organisation
Jobleads-UK
Location
Sunbury-on-Thames, England, United Kingdom
innovations from experimentation through to productised, maintainable solutions that deliver measurable value.* Drive engineering excellence across ML systems, including CI/CD, testing, observability, reliability, and MLOps guidelines.* Define technical standards, patterns, and protocols for ML engineering and applied ML science across teams.* Lead complex, multi-team technical initiatives ...

Machine Learning Engineer, Senior Manager

Hiring Organisation
Credit Acceptance Corporation
Location
United States
Employment Type
Permanent
Salary
USD 270,386 Annual
such infrastructure Hands-on expertise in scaling and maintaining production-grade ML services, with a strong focus on ML/LLM Operations (versioning, automation, observability, automated training and monitoring, etc.) and ability to balance ML model complexity with production requirements Passion for identifying new business opportunities and experience of using ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
implement platform engineering tools to deliver self-service capabilities via an IDP (e.g. Backstage) Define and enforce monitoring, alerting, and observability standards to ensure infrastructure reliability and performance Foster collaboration between development and operations teams, bridging gaps to drive a unified engineering culture Mentor junior engineers, share knowledge across … Developer Tooling : Experience implementing developer portals (Backstage), artifact repositories (JFrog Artifactory, Nexus), API gateways, and self-service platform engineering tools to support internal teams Observability & Reliability : Hands-on experience with monitoring and observability platforms (Honeycomb, Prometheus, Grafana), including defining alerting standards and ensuring infrastructure reliability at scale Architecture & Design: Strong ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
implement platform engineering tools to deliver self-service capabilities via an IDP (e.g. Backstage)* Define and enforce monitoring, alerting, and observability standards to ensure infrastructure reliability and performance* Foster collaboration between development and operations teams, bridging gaps to drive a unified engineering culture* Mentor junior engineers, share knowledge across … Developer Tooling**: Experience implementing developer portals (Backstage), artifact repositories (JFrog Artifactory, Nexus), API gateways, and self-service platform engineering tools to support internal teams* **Observability & Reliability**: Hands-on experience with monitoring and observability platforms (Honeycomb, Prometheus, Grafana), including defining alerting standards and ensuring infrastructure reliability at scale* **Architecture & Design:** Strong ...

Senior Software Engineer

Hiring Organisation
Permax Recruitment Limited
Location
West London, London, United Kingdom
Employment Type
Permanent, Work From Home
efficiently, this role spans cloud infrastructure, data platform engineering, and AI tooling. You'll manage our Snowflake environment and contribute to our Claude Enterprise observability alongside your core AWS and DevOps responsibilities. Beyond keeping systems running, we expect you to identify improvements, take ownership of them, and actively upskill colleagues … prod Lead the implementation of monitoring, logging, and alerting systems to ensure reliability in our solutions Collaborate in the management and optimisation of our observability dashboards, ensuring platform health is visible and actionable across the team Take ownership of our Snowflake environment: access controls, cost governance, performance, and data organisation ...

Senior Golang Engineer

Hiring Organisation
Talent Smart
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£650/annum
clean, maintainable, well-tested code following engineering best practices. Participate in code reviews, technical design sessions, and architectural discussions. Drive improvements in performance, scalability, observability, and reliability. Work within Agile delivery teams, contributing to continuous improvement and DevOps practices. Essential Skills & Experience Strong commercial experience developing backend applications using … Event Grid, or Event Hubs. Exposure to Kafka or other messaging technologies. Knowledge of Infrastructure as Code (Terraform or Bicep). Experience with observability platforms such as Prometheus, Grafana, or Azure Monitor. Exposure to blockchain or Web3 development, particularly Solana . Experience working across polyglot engineering environments. Personal Attributes Passionate ...

Azure Devops Engineer

Hiring Organisation
VIQU IT
Location
London, Candlewick, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £600/day Inside IR35
Develop and manage CI/CD pipelines using GitHub Actions or similar automation tools. Monitor platform performance, troubleshoot issues and provide insights using cloud observability tooling. Follow change management processes to support safe and controlled production deployments. Engage with a wide range of technical and non-technical stakeholders across … have proven commercial experience with: Microsoft Azure cloud (essential) Infrastructure as Code Terraform CI/CD pipelines Kubernetes clusters and containerised environments Monitoring and observability tools such as Prometheus, Grafana, Dynatrace, AppDynamics, Splunk or similar DevOps and/or Site Reliability Engineering (SRE) environments Networking fundamentals including DNS, VPNs, load ...

Kubernetes Specialist

Hiring Organisation
CGI
Location
Reading, United Kingdom
Employment Type
Full Time
related technologies •Support & Maintain Linux-based systems and platform environments •Troubleshoot & Resolve complex issues across applications, platforms, and infrastructure •Monitor & Analyse platform health using observability and logging tools including Elastic Required qualifications to be successful in this role To succeed in this role, you should have strong experience across DevOps … skills within Agile delivery environments Desirable experience: •Experience with AWS, Azure, GCP, or secure private cloud environments •Knowledge of Elastic for monitoring, logging, and observability •Understanding of service mesh technologies and concepts •Experience with Infrastructure-as-Code tooling #LI-SB2 Together, as owners, let's turn meaningful insights into action. ...

Secure Data Engineer

Hiring Organisation
Capgemini
Location
City and Borough of Birmingham, United Kingdom
Employment Type
Full Time
components that process, transform and expose data for analytical, operational or AI driven use cases. You will follow strong engineering practices, with testing and observability built in from the start. Streaming and Real Time Architecture Designing and implementing data ingestion and event driven patterns that support real time or near … using CI/CD principles adapted for Defence delivery. You will own your applications in production and contribute to secure patterns for deployment. Resilience, Observability and Compliance Implementing health monitoring, structured logging, metrics and lineage to meet Defence requirements for auditability, security and operational assurance. You will design systems that ...

Microsoft Azure Devops Engineer

Hiring Organisation
Jobleads-UK
Location
Park Central, England, United Kingdom
velocity and platform reliability.Innovation & Continuous ImprovementChampion platform modernisation and cloud-native ways of working.Explore and implement AI-enhanced capabilities across CI/CD, testing, observability, and operational workflows.Promote monitoring, observability, and operational excellence through modern tooling and best practices.Contribute to engineering standards, reusable accelerators, and DevOps reference architectures.Your Skills … scalability, resilience, and automation.Experience embedding DevSecOps practices, Azure security controls, identity management (Entra ID), and governance frameworks within cloud delivery programmes.Strong knowledge of monitoring, observability, and GitOps practices, leveraging tools such as Azure Monitor, Application Insights, Prometheus, and Grafana.Excellent leadership, stakeholder management, and communication skills, with experience leading engineering teams ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi‐step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root‐cause analysis, anomaly detection, and semantic … plus. Preferred Qualifications: AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.). Experience building self‐service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About the company At the company, we commit our people ...

Grafana Observability Engineer

Hiring Organisation
Shaw Daniels Solutions
Location
Other, United Kingdom
Employment Type
Permanent, Work From Home
Grafana Observability Engineer Location: Fully remote Our client They are delivering a company-wide Digital Transformation (DX) Programme that will modernise their technology landscape and transform how they deliver services. As part of this journey, their IT and Portfolio Delivery teams are implementing a new enterprise technology platform, making this … exciting opportunity to join the business and help shape their future. Role Overview Reporting to the Lead Platform Engineer, the Senior Observability Engineer will own and develop their observability capability, leading the design, implementation and continuous improvement of their monitoring and alerting platform. Working closely with infrastructure, platform and application ...

Senior SRE: GCP & Kubernetes, Automation Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
degree of autonomy and ownership. You will design, implement, and operate cloud-native infrastructure on GCP using Kubernetes, Terraform, and Helm, while championing automation, observability, and best practices across the SDLC. #J-18808-Ljbffr ...

Senior Cloud Platform Engineer: GCP, Kubernetes & DevSecOps

Hiring Organisation
Jobleads-UK
Location
Bolsterstone, England, United Kingdom
native platform for IDAM 2.0. You will build and maintain Kubernetes clusters and IaC, implement CI/CD pipelines, and manage security, networking and observability across environments. You will collaborate with IAM, security and architecture teams to deliver scalable platform services and drive automation and best practices across the stack. ...

Senior Lead Site Reliability / DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top‐notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability/DevOps Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part … significant business impact through your capabilities and contributions, and apply deep technical expertise and problem‐solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices to support the business ...

Logging & Inventory Cloud Engineer

Hiring Organisation
Diana Duggan UK Limited
Location
Bournemouth, Dorset, England, United Kingdom
Employment Type
Contractor
Contract Rate
£400 - £450 per day
scripts and tooling using Python and Shell scripting. Support cloud engineering activities across AWS and Google Cloud Platform (GCP). Drive improvements in cloud observability, monitoring and operational reporting. Manage and optimise containerised workloads where required. Work closely with engineering and platform teams to improve cloud governance and operational controls. ...

Senior Machine Learning Engineer Software engineering London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
proficiency in writing clear, production‐ready Python code. Experience with production ML models (online or offline) and standard MLOps practices. Experience with monitoring and observability of production systems, with a strong sense of ownership. Experience with training and operating models on Databricks. Familiarity with cloud‐based application development (AWS & Azure ...

Integration Engineer

Hiring Organisation
Searchability NS&D
Location
City of London, London, United Kingdom
exposure Background in enterprise or regulated environments (ideal) Desirable skills: Kafka/event-driven architecture Apigee or MuleSoft PostgreSQL/document stores AWS certifications Observability tools (CloudWatch, ELK, Grafana) Security Clearance Due to the nature of the work, candidates must either hold an active SC Clearance or be eligible ...

Senior Software Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering standards and the adoption of new tools, technologies and ways of working. Support the delivery and operation of production services, driving reliability, performance, observability and incident resolution. Communicate technical decisions, risks and progress effectively with both technical and non-technical stakeholders. Skills, Knowledge & Expertise Strong experience developing software using ...

Artificial Intelligence Engineer

Hiring Organisation
Discover International
Location
West Sussex, England, United Kingdom
integration, tool/function calling and agent orchestration Nice to have • Model Context Protocol (MCP) • Kafka or event-driven architectures • MLOps, LLM evaluation, observability and monitoring To apply or find out more, get in touch. #ContractJobs #AIJobs #GenAI #AIEngineer #LangGraph #Python #RAG #LLM #MachineLearning #Hiring #TechRecruitment #ContractRole ...

Principal Engineer

Hiring Organisation
Experis
Location
Knutsford, Cheshire, United Kingdom
Employment Type
Contract
Contract Rate
£490 - £500/day
orchestration * Drive modern backend engineering standards, including CI/CD, deployment automation, and operational best practices * Architect durable, fault-tolerant distributed systems with high observability * Develop integration strategies across enterprise APIs, event streams, and downstream systems * Build reusable frameworks, prototypes, and production solutions independently * Guide engineering teams on orchestration, system ...

Managing Engineer – Database, Platform

Hiring Organisation
Jobleads-UK
Location
Belfast, Northern Ireland, United Kingdom
product engineering, SRE, security, and cloud teams to define SLAs/SLOs, incident response, root cause analysis, and risk mitigation. Establish monitoring, alerting, and observability for database health and performance. Drive proactive incident prevention. Define and enforce standards, best practices, and governance for database access, security, backups, and compliance. Participate ...

Integration Solution Architect

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Contract
Azure Azure Integration Services API Management Event-Driven Architecture Kafka IBM MQ Java/Spring Boot Kubernetes Microservices CI/CD SQL Server MongoDB Observability and Monitoring Platforms If you receive suspicious outreach claiming to be from us, please contact us via the ManpowerGroup website. ...