26 to 33 of 33 Prometheus Jobs in Manchester

Database Reliability Engineer

Location
Manchester, England, United Kingdom
multi-cloud—while ensuring rigorous data integrity and mobility A Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual intervention Engineering via Code: While ...

Infrastructure Engineer

Location
Greater Manchester, England, United Kingdom
ideally have experience across: Linux | Server Hardware | Virtualisation | Networking | Storage | Infrastructure Troubleshooting Experience with GPU/NVIDIA hardware, Proxmox, HPC, Zabbix/Prometheus/Grafana, Bash, Python or Ansible would be advantageous. For the Senior position , previous experience leading or managing technical team members is required. Important: You’ll need ...

Senior Platform Engineer: Kubernetes & Scalable SaaS

Location
Manchester, England, United Kingdom
based languages, shaping platforms as products and driving reliability and observability. You will own Kubernetes engineering, storage and data persistence, telemetry with OpenTelemetry, Prometheus and Grafana, and collaboration across teams to reduce technical debt. #J-18808-Ljbffr ...

Senior Platform Engineer

Location
Manchester, England, United Kingdom
just look atmetrics; you design the telemetry that allows for deep‐dive distributed tracing and root‐cause analysis. Observability Infrastructure: Hands‐on experience scaling Prometheus and Grafana tohandle high‐cardinality data. Your Skills BS degree in Computer Science, related technical field, or equivalent practical experience as a professional Platform Engineer. … operating workloads in Kubernetes. Experience with cloud storage S3, NFS, NetApp OnTap, AWS FSX Distributed Systems and programming techniques. Monitoring and metrics infrastructure (OTEL, Prometheus, Grafana). Cloud (AWS/GCP/Azure/Private DC). Bonus: Experience with AI/ML, Natural Language Technologies you’ll work with ...

Cloud Security Engineer London National Security West

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Our client is the digital and intelligence arm of a major defence company. Location(s):UK, Europe & Africa : UK : Manchester Our client is home to 4,500 digital, cyber and intelligence experts. We work collaboratively ...

Cloud Security Engineer Manchester National Security West

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Our client is the digital and intelligence arm of a major defence company. Location(s):UK, Europe & Africa : UK : Manchester Our client is home to 4,500 digital, cyber and intelligence experts. We work collaboratively ...

Senior Data Engineer

Location
Altrincham, England, United Kingdom
quality checks as first‐class concerns. Ensure data quality and observability: Embed data testing and validation processes alongside monitoring and alerting systems such as Prometheus and Grafana to ensure reliability, performance, and operational insight across data services. Work with metadata and standards: Implement metadata capture and validation within data pipelines … checks throughout pipelines to ensure accuracy, completeness, consistency, and reliability. Practical experience implementing monitoring, observability and alerting for data platforms using tools such as Prometheus and Grafana. Strong understanding of data architecture patterns, including data lakes, data warehouses, and event‐driven architectures. Awareness of GDPR and data protection practices, including ...

ML/AI Engineer

Location
Manchester, England, United Kingdom
Implement end‐to‐end observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace. Establish actionable alerting and runbooks for on‐call operations; drive incident reviews and reliability improvements. Operate a model registry (e.g., MLflow) with … stage pipelines; experience with GitOps, artefact repositories, and environment promotion. Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation. Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems. Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments. Expert ...