751 to 775 of 6,477 Permanent Observability Jobs

Remote Azure DevOps / Infrastructure Engineer

Location
Somerset, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
London, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Lossiemouth, Morayshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Daventry, Northamptonshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Southampton, Hampshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Bicester, Oxfordshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Kilmacolm, Renfrewshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Hassocks, Sussex, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Farnworth, Cheshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Cumnock, Ayrshire, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Kirriemuir, Angus, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Hiring Organisation
QuantumLoopAi
Location
Ceremonial county gwynedd, United Kingdom
Azure environment. Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. Experience designing and building custom dashboards for resource monitoring and performance tracking. Expertise in logging and telemetry frameworks, and setting up proactive alerting. Familiarity with ...

Remote Azure DevOps / Infrastructure Engineer

Location
Newcastle upon Tyne, Northumberland, United Kingdom
Azure environment. - Set up and manage users, groups and service principals to ensure secure, least-privilege access to all Azure resources. Monitoring and Observability - Build dashboards to monitor performance, utilisation and health of Azure resources, providing actionable insights to the engineering team. - Implement centralised logging and telemetry solutions using Azure … setting up and managing RBAC, users, groups and service principals within Azure. - Strong knowledge of Azure Monitor, Application Insights and related monitoring and observability tools. - Experience designing and building custom dashboards for resource monitoring and performance tracking. - Expertise in logging and telemetry frameworks, and setting up proactive alerting. - Familiarity with ...

Staff Software Engineer - AI

Location
Greater London, England, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event‐driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands‐on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Senior Data Platform Engineer (Python & Databricks)

Location
City of Edinburgh, Scotland, United Kingdom
ingesting, transforming, and delivering financial data. Design data models that support enterprise reporting, analytics, and downstream integrations. Improve the performance, reliability, maintainability, and observability of existing data workflows. Troubleshoot complex issues across data-processing and application layers. Python and API Development Design, build, and maintain production-grade applications and services … services or other data-intensive industries. Experience with cloud-based data architectures. Familiarity with infrastructure-as-code tools such as Terraform. Experience improving the observability and operational reliability of data pipelines. Familiarity with modern data governance, access-control, and data-quality practices. Experience working in an enterprise environment with strict ...

Lead Software Engineer - Application Owner

Location
Glasgow, Scotland, United Kingdom
after‐action reviews, and closure of follow‐up actions. Define and continuously improve production readiness standards, including release safety and rollback strategy, dependency awareness, observability requirements, and operational runbooks. Contribute hands‐on to design and delivery, including system design, code reviews, automation, and complex troubleshooting, with secure‐by‐design … with strong operational accountability, including controls, resiliency and recovery, and remediation tracking. Strong system design fundamentals and cloud‐native operational patterns, including scalability, reliability, observability, and dependency management. Hands‐on experience with Go‐based services and modern CI/CD practices. Experience operating workloads on AWS and Kubernetes ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US Our client ...

Safety Engineer - Free Tier Abuse

Location
Greater London, England, United Kingdom
models into production systems Architect robust APIs, data pipelines, and service architectures supporting real‐time and batch moderation workflows Implement comprehensive monitoring, alerting, and observability systems; establish SLIs, SLOs, and performance benchmarks Partner with ML engineers to translate research models into production‐ready systems and integrate them across our product … pipelines, and Python expertise (asynchronous Python, backend frameworks) Infrastructure & DevOps proficiency: cloud platforms (AWS/GCP), containerization (Docker/K8s), CI/CD pipelines Observability mindset with experience in monitoring tools (Prometheus, Grafana) and building observable systems Track record of taking products or systems from 0→1 with measurable impact ...

Lead DevOps Engineer

Location
Manchester, England, United Kingdom
lifecycle. Automation and reliability engineering Proficient in Python with solid scripting in PowerShell and Bash to drive automation and operational efficiency, combined with strong observability practice (for example Azure Monitor, Sentinel, Application Insights and Log Analytics) using SLIs, SLOs and distributed tracing to sustain and improve system reliability. Autonomy … Code and CI/CD. This is a hands‐on leadership role where you'll influence key technical decisions, champion DevSecOps, resilience and observability, and work closely with engineering teams to create high‐quality services that make a real impact. You'll use technologies including Azure, Azure DevOps and Terraform ...

Senior Service Reliability Engineer

Location
Knutsford, England, United Kingdom
reliability and customer experience expectations. Working across Engineering, Infrastructure, Security, and Product teams, you will identify opportunities to automate manual processes, improve monitoring and observability, strengthen resilience, and reduce operational risk. You will also contribute to service reviews, support governance and control activities, and provide clear communication to stakeholders during … reliability for critical business services. Cloud, Platform & Engineering Expertise –Strong hands-on knowledge of AWS cloud technologies, microservices, APIs, containerized platforms (OpenShift/Kubernetes), observability tooling, automation, CI/CD, Infrastructure as Code, and modern platform engineering practices. Senior Stakeholder & Operational Leadership–Proven ability to lead critical incidents, drive root ...

SRE Engineer

Location
Greater London, England, United Kingdom
engineering standards Driving infrastructure‐as‐code and automation across Azure and on‐prem Improving image bakery pipelines for secure, repeatable server builds Embedding observability using metrics, logs, traces, and effective alerting Ensuring all practices align with ISO 27001 and internal security frameworks Managing automated patching, vulnerability remediation and configuration compliance … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions) Experience with monitoring and observability stacks Solid understanding of OS fundamentals (Windows/Linux), security, networking Background in scripting or software development (PowerShell, Python, Go) Experience with containers and orchestration (Docker ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Paisley, Renfrewshire, UK
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US Our client ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US Our client ...

Senior Service Reliability Engineer

Hiring Organisation
Barclays
Location
Knutsford, Cheshire, United Kingdom
Salary
£ 70 K
meet reliability and customer experience expectations.Working across Engineering, Infrastructure, Security, and Product teams, you will identify opportunities to automate manual processes, improve monitoring and observability, strengthen resilience, and reduce operational risk. You will also contribute to service reviews, support governance and control activities, and provide clear communication to stakeholders during … service reliability for critical business services.Cloud, Platform & Engineering Expertise –Strong hands-on knowledge of AWS cloud technologies, microservices, APIs, containerized platforms (OpenShift/Kubernetes), observability tooling, automation, CI/CD, Infrastructure as Code, and modern platform engineering practices.Senior Stakeholder & Operational Leadership–Proven ability to lead critical incidents, drive root cause ...

Applied AI ML Lead - Python & Agentic AI

Location
Glasgow, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...