451 to 475 of 502 Remote/Hybrid Observability Jobs

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Wales, United Kingdom
working on Mondays and Fridays Take technical ownership of the platforms that support Centerprise services and customer operations. You will lead improvements in reliability, observability and automation while remaining closely involved in complex engineering, major incidents and service recovery. Role Summary As Principal Platform Engineer, you will … hours escalation when required. Identify and address technical debt, operational risk and platform weaknesses. Ensure services remain supportable, recoverable and operationally efficient. Observability and service health Own monitoring and observability tooling, standards and operational dashboards. Develop service health metrics that provide clear and useful operational insight. Improve the quality ...

Staff Backend Engineer - Grafana Second Horizon | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … remote opportunity, and we would be interested in applicants located in Spain, Sweden, UK, Ireland or Germany. The Opportunity: At Grafana Labs, we build observability tools that help users understand, respond to, and improve their systems – regardless of scale, complexity, or tech stack. We recently started a skunkworks initiative with ...

Senior Director Technology

Hiring Organisation
Jobleads-UK
Location
Langley, England, United Kingdom
buildsourproductionAWSplatform, and onboards development teams and products to this platform.You’llbe responsible fordriving the move toinfrastructure as code, account management, automated pipelines, observability, resilience, operating coverage and cost control. You’llalso help define how the platform supports AI, data engineering and high-scale product demand, including where AWS Bedrock … standards for infrastructure as code, CI/CD, AWS account management, platform guardrails and developer enablement. Improve the operational model for the platform, including observability, incident response, reliability and 24/7 support coverage. Partner with Product and Commercial teams to get ahead of major demand changes, customer commitments ...

Operations Engineer

Hiring Organisation
ASCENT PROFESSIONAL SERVICES LTD
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
continuity. Key Responsibilities Provide operational support for enterprise platforms, applications, integrations, and associated technologies. Monitor system health, availability, and performance using monitoring, alerting, and observability tools. Analyse, troubleshoot, and resolve incidents affecting services and platforms. Perform root cause analysis and contribute to implementing permanent solutions to prevent recurring issues. Coordinate … within IT operations, support engineering, or service management environments. Experience supporting business-critical production services and operational platforms. Knowledge of monitoring, logging, alerting, and observability practices. Experience working with incident, problem, change, and release management processes. Excellent communication skills with the ability to collaborate effectively across multiple technical and business ...

Senior Cloud Infrastructure Engineer VMware

Hiring Organisation
100% IT Recruitment Ltd
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
infrastructure issues. Lead technical investigations and major incident recovery. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. Support … experience with Veeam Backup & Replication. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability ...

Cloud SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
databases. Day to day, you will monitor service health, resolve incidents, execute controlled changes, maintain patching and hardening standards, and improve reliability through automation, observability, and clear operational documentation. This is a 6-month contract, to work remotely (outside IR35) Key Responsibilities Infrastructure discovery and assessment - Build and maintain … AIOps and self-healing to detect and resolve issues earlier. Cloud administration - Administer compute, storage, and OS-layer services in Azure; support application infrastructure. Observability & monitoring - Implement and tune monitoring, alerting, and dashboards; improve signal quality and reduce noise. Security & compliance - Apply hardening, patching, and access controls in line with ...

Lead AI Engineer

Hiring Organisation
Capco
Location
Borough of Tameside, United Kingdom
Employment Type
Full Time
LLMs and multi-modal models at scale Strong engineering background in Python with proven backend and API development skills Solid understanding of scalable MLOps, observability, and cloud-native AI deployment Excellent communication, problem-solving, and project management skills in agile environments Bonus Points For Experience with agentic frameworks (e.g., LangChain … LlamaIndex) Experience in deep learning frameworks and front-end development Familiarity with Langfuse, Langsmith, or other LLM observability tools Understanding of Model Context Protocol and bias/hallucination mitigation techniques Previous success in integrating GenAI solutions into enterprise-scale systems Why Join Capco Deliver high-impact technology solutions for Tier ...

Senior Software Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
Impact and Responsibilities We are seeking a Senior Software Engineer who thrives on untangling complex systems and modernising core infrastructure without breaking production. This is an exciting opportunity to modernise core C#/SQL systems ...

Observability Software Engineer - Data & Dashboards

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Vercel is seeking a Software Engineer to join the Observability team in London. You will design, implement, and maintain observability features that help users monitor and understand their applications’ health and performance. The role offers in-office anchor days on Monday, Tuesday, and Friday for those within commuting distance; otherwise ...

Gen AI Engineer

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
based AI stack Experience with high volume document processing Familiarity with enterprise architecture security and compliance controls Exposure to monitoring model evaluation and AI observability tools Follow enterprise standards for security governance observability and performance We are a Disability Confident Employer : Capgemini is proud to be a Disability Confident Employer ...

platform engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Design, build, and operate core GenAI platform components, including an LLM routing gateway, vector search and RAG infrastructure, tool registry and MCP gateway, AI observability and evaluation tooling, and infrastructure for long-running agentic workflows; Own production-quality delivery of platform features from design through rollout, monitoring, and follow … Contribute to resilient system design with sensible APIs, failure handling, rate limiting, retries, idempotency, and safe change management; Improve reliability and observability through metrics, dashboards, alerting, incident follow-ups, and operational improvements; Partner with Applied AI Engineers and product teams to understand platform needs and help them build AI-powered ...

Gen AI Architect

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
production-grade AI systems using Amazon Bedrock, retrieval-augmented generation (RAG), agentic workflows, and cloud-native AWS services. Drive architecture standards, model orchestration, governance, observability, and operational excellence across the GenAI lifecycle while collaborating with engineering, security, compliance, and business stakeholders Hybrid working: The places that you work from … customization, prompt orchestration, retrieval pipelines, and agentic workflows Design agentic AI systems incorporating tool use, workflow orchestration, memory management, and autonomous decision flows Implement observability for prompts, model responses, vector retrieval quality, and agent execution workflows Integrate GenAI capabilities into enterprise applications, APIs, workflow platforms, and data ecosystems Work with ...

Lead Data Engineer

Hiring Organisation
Tenth Revolution Group
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £110,000 per annum
Python SQL Data Warehousing Data Lakes Data Pipelines APIs & Data Services Machine Learning Infrastructure AI & Machine Learning Systems Analytics Engineering Data Governance Data Observability Cloud-Native Architecture You'll Be Responsible For Designing and building scalable cloud-based data platforms. Creating robust ETL/ELT pipelines and data services. Developing … APIs and trusted data products for enterprise clients. Supporting machine learning and AI initiatives through high-quality data architecture. Establishing governance, lineage, quality and observability standards across the data estate. Working closely with Product, Engineering, Data Science and Leadership teams to influence strategic decisions. Helping define the long-term data ...

AI Engineer (Infrastructure SRE & Automation)

Hiring Organisation
Jobleads-UK
Location
Livingston, Scotland, United Kingdom
repetitive operational workflows across compute, storage, and networking layers. Integrate AI capabilities into CI/CD pipelines and infrastructure-as-code ecosystems. Data Engineering & Observability Build and maintain pipelines for high-volume telemetry data (metrics, logs, traces). Ensure data quality, labelling, and feature engineering for ML models. Leverage observability ...

Principal Cloud Engineer

Hiring Organisation
100 Percent
Location
Cardiff, South Glamorgan, United Kingdom
Employment Type
Permanent
Salary
£58000 - £60000/annum + Bonus + Benefits
infrastructure issues. Lead technical investigations and major incident recovery. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. Support … experience with Veeam Backup & Replication. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability ...

AI Engineer (Infrastructure SRE & Automation)

Hiring Organisation
Sky
Location
Livingston, West Lothian, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
Salary
GBP per hour
workflows across compute, storage, and networking layers . Integrate AI capabilities into CI/CD pipelines and infrastructure-as-code ecosystems . Data Engineering & Observability Build and maintain pipelines for high-volume telemetry data (metrics, logs, traces) . Ensure data quality, labelling, and feature engineering for ML models . Leverage … observability tools to inform AI systems . Reliability & Governance Define SLIs/SLOs for AI systems embedded in infrastructure workflows . Ensure robustness, explainability, and auditability of AI-driven decisions . Implement feedback loops to continuously improve model performance . AI-Driven SRE Enablement Develop ML/AI models ...

Senior Software Engineer, Billing & Revenue

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
right. You'll build systems that turn usage and commercial agreements into accurate, auditable invoices, and you'll build the checks and observability that keep them trustworthy as we scale. You'll also work well beyond billing itself, building integrations with other backend systems like our ERP and Data Platform … understand that billing and revenue systems have little room for error, and you design for it — with clear data models, good tests, and observability built in. You enjoy working across teams. You can talk to Finance or Business Development, understand what they need, and turn a commercial model into something ...

Staff Frontend Engineer (React Native / Mobile)

Hiring Organisation
Feeld
Location
Greater London, United Kingdom
Employment Type
Full Time
Salary
80000 to 110000 GBP Annually
performance, stability/crash rates, startup time, build/release velocity, or app size . Increased confidence in production through better observability, incident response practices, and ownership . Enabled other FE engineers to move faster through documentation, pairing/mentorship, reviews, and reusable platform components . What … traffic/user counts, complex feature sets) with a focus on reliability and performance. Demonstrated production ownership : incident response, debugging complex issues, and improving observability (metrics/logs/traces). Experience improving delivery systems (CI/CD, automated testing strategy, release process) and keeping teams moving. Staff-level ...

Azure Cloud SRE — Remote 6-Month Contract, Automation

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Korn Ferry is hiring a Cloud SRE to support our client’s Azure infrastructure estate, focusing on discovery, planning, and delivery across VMs, networking, storage, and databases. You will monitor service health, drive changes, patching ...

Head of AI Engineering

Hiring Organisation
Jobleads-UK
Location
Bexhill-on-Sea, England, United Kingdom
As a member of the CIO Technology Engineering senior leadership team, the Head of AI Engineering will lead the design, deployment, integration and continuous improvement of enterprise AI and machine learning capabilities, including ML platforms ...

Operational Architect

Hiring Organisation
Capgemini
Location
City and Borough of Birmingham, United Kingdom
Employment Type
Full Time
technical management activities during major incidents or key business events, supporting escalation, analysis and communication. Help maintain and evolve service improvement roadmaps, particularly for observability, monitoring and alerting capabilities. Support governance activities, including reviewing solution designs for cross-portfolio consistency and alignment with operational and observability standards. Work with architects … service teams and suppliers to support the handover of new or changed services into Live. Contribute to the development and maintenance of observability standards, guidance and blueprints. Assist in identifying opportunities to improve the use of observability data for service insight and continuous improvement. Your skills and experience Essential Experience ...

Gen AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
SkillsExperience with Azure based AI stackExperience with high volume document processingFamiliarity with enterprise architecture security and compliance controlsExposure to monitoring model evaluation and AI observability toolsFollow enterprise standards for security governance observability and performanceWe are a Disability Confident Employer:Capgemini is proud to be a Disability Confident Employer (Level ...

Senior Machine Learning Engineer, Developer Advocacy | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Senior ML Engineer Recommender Systems, Developer Advocacy | UK | Remote Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud … Experience with directed graphs, sequence models, or prerequisite‐aware recommendations Experience with contextual bandits or other exploration strategies Familiarity with Grafana or the broader observability ecosystem Experience with open source software or transparent development practices Experience working with privacy, fairness, explainability, or responsible personalization constraints Compensation & Rewards ...

OAT Quality Engineer

Hiring Organisation
Hays Technology
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£398/day £398 p/d Inside IR35
procedures, including incident, problem and change management readiness. Verify backup, restore, deployment, patching and infrastructure connectivity processes. Evaluate application and infrastructure monitoring, alerting and observability solutions. Analyse system logs, monitoring outputs and performance data to identify operational risks. Support release activities and ensure operational acceptance criteria are met. Work with … server testing. Experience testing cloud-hosted applications, ideally within AWS environments. Strong understanding of operational readiness, service transition and production supportability. Experience with monitoring, observability and log analysis tools such as Splunk, Dynatrace, New Relic or ELK. Background working within IT Operations, Service Management or mission-critical production environments. Strong ...

Staff Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering standards, best practices and reusable patterns while partnering with Enterprise Architecture and influencing technical direction Drive engineering excellence by improving code quality, testing, observability, reliability and operational practices Support end‐to‐end delivery by guiding teams through complex technical challenges, improving decision‐making, and contributing to planning and risk … data lakes/lakehouse architectures, Iceberg or similar table formats, as well as batch and streaming processing Knowledge of data quality, governance, cataloguing and observability tools (e.g. Datadog), with DBT or AI‐assisted engineering practices as a plus Your benefits 29 days holiday allowance + bank holidays Private medical ...