151 to 175 of 175 Remote Prometheus Jobs

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment
Location
North London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join the on-call … desirable Reference: BBBH27029A DevOps, DevOps Engineer, AWS, Cloud Security, Cyber, DevSecOps, Terraform, Ansible, Linux, Networking, GitHub Actions, CI/CD, Bash, Python, Go, Grafana, Prometheus, Woking, Remote, Surrey, London If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV. ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
Camden, London, Camden Town, United Kingdom
Employment Type
Permanent
Salary
£65000 - £70000/annum + Remote + Progression
cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join the on-call … desirable Reference: BBBH27029A DevOps, DevOps Engineer, AWS, Cloud Security, Cyber, DevSecOps, Terraform, Ansible, Linux, Networking, GitHub Actions, CI/CD, Bash, Python, Go, Grafana, Prometheus, Woking, Remote, Surrey, London If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV. ...

Consulting Principal - Solution Architect

Location
Greater London, England, United Kingdom
rollback and controlled artefact management using tools such as Jenkins, GitLab and AWS CodePipeline. Establish effective monitoring, logging and incident-management capabilities using CloudWatch, Prometheus, Grafana and the ELK stack, while advising stakeholders on AWS container best practices. Work model We believe hybrid work is the way forward … Experience implementing secure secrets management, Kubernetes security policies, network policies and controls for multi-tenant platforms. Knowledge of observability and centralized logging using CloudWatch, Prometheus, Grafana and ELK in high-security environments. Strong analytical and problem-solving skills, with the ability to create resilient solutions and manage technical ambiguity with ...

Cloud Security Engineer Gloucester- National Security West

Hiring Organisation
Hackajob Ltd
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
hackajob is partnering directly with BAE Systems Digital Intelligence to hire for this role. Location(s):UK, Europe & Africa : UK : Leeds BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts. ...

Cloud Security Engineer London National Security West

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
hackajob is partnering directly with BAE Systems Digital Intelligence to hire for this role. Location(s):UK, Europe & Africa : UK : Manchester BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts. ...

Cloud Security Engineer Manchester National Security West

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
hackajob is partnering directly with BAE Systems Digital Intelligence to hire for this role. Location(s):UK, Europe & Africa : UK : Manchester BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts. ...

Cloud Security Engineer Leeds National Security West

Hiring Organisation
Hackajob Ltd
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
hackajob is partnering directly with BAE Systems Digital Intelligence to hire for this role. Location(s):UK, Europe & Africa : UK : Leeds BAE Systems Digital Intelligence is home to 4,500 digital, cyber and intelligence experts. ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade … this role includes on-call responsibilities. Key Responsibilities Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level ...

Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) |

Hiring Organisation
Pure Resourcing Solutions
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £120,000 per annum
training versus inference, matrix multiplication, KV-caching, that level of detail Comfortable in profiling tools like Nsight or PyTorch Profiler, and monitoring stacks like Prometheus and Grafana Python for data work, Pandas and NumPy, plus general scripting Nice to have rather than essential: a postgraduate degree and research background (publications ...

Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)

Hiring Organisation
Pure Resourcing Solutions Limited
Location
Linton, Dry Drayton, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£55000 - £70000/annum
work with GPU or accelerator code, CUDA or similar Familiarity with profiling tools (Nsight, PyTorch Profiler) and ideally some exposure to monitoring stacks (Prometheus, Grafana) Strong Python for data work, Pandas and NumPy, genuine scripting ability Nice to have: exposure to inference serving frameworks like vLLM, published research, or open ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
infrastructure — monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues and ensuring service uptime. Conduct post‐incident reviews and help improve system … Strong experience with AWS services and containerised applications Strong experience operating operational data stores (Aurora MySQL, DynamoDB). Expertise in using monitoring tools(e.g. Prometheus, Grafana, CloudWatch) for real‐time platform performance insights. Strong understanding of network security and Cloudflare, VPC and networking fundamentals, with a clear grasp ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop … Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital ...

Senior DevSecOps Engineer

Location
Manchester, England, United Kingdom
DevOps AWS (EKS, Lambda, RDS/Postgres, DynamoDB, SQS, Kinesis, S3, Cognito, Route53, VPC, EC2) Terraform, Kubernetes, Helm, Ansible, Puppet GitLab CI Observability & Security Prometheus, Grafana, OpenSearch, CloudWatch Okta (SSO/IdP) Vulnerability management, secrets management, penetration test tooling Application layer (context, not expectation) Scala, Kotlin, TypeScript, Python, built … deployments stable. Own our platform security controls: vulnerability management, penetration test remediation, secrets management, and least-privilege access. Run our observability and monitoring platforms (Prometheus, Grafana, OpenSearch, CloudWatch), tuning alerts and driving improvements into the delivery pipeline. Continuously review and optimise our AWS infrastructure: Cost monitoring, right-sizing, capacity planning ...

Senior Data Engineer

Location
Altrincham, England, United Kingdom
quality checks as first‐class concerns. Ensure data quality and observability: Embed data testing and validation processes alongside monitoring and alerting systems such as Prometheus and Grafana to ensure reliability, performance, and operational insight across data services. Work with metadata and standards: Implement metadata capture and validation within data pipelines … checks throughout pipelines to ensure accuracy, completeness, consistency, and reliability. Practical experience implementing monitoring, observability and alerting for data platforms using tools such as Prometheus and Grafana. Strong understanding of data architecture patterns, including data lakes, data warehouses, and event‐driven architectures. Awareness of GDPR and data protection practices, including ...

Senior DevOps Engineer

Hiring Organisation
Granite Recruitment and Consulting
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
reliable services. You will get to work in a modern, cloud-based tech environment, and work with technologies such as Azure, Kubernetes, Terraform, Chef, Prometheus and Grafana. The team are passionate about working with cutting edge technology and adopt the latest DevOps/SRE methodologies. You will need the following …/CD tools and pipelines Azure DevOps or Argo CD It would be a bonus if you have experience with the following: Grafana Prometheus Chef Ansible Cloud platform security Benefits will include: 27 days holiday Private medical insurance Life assurance Flexible working A commitment to continue your professional development Career ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Hiring Organisation
Learning Technologies Group
Location
United Kingdom, UK
Employment Type
Full-time
configuration management (etcd)Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache)Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotationsEvaluating and implementing distributed storage solutions as the platform evolvesImproving deployment workflows and release processesCollaborating with internal … networking fundamentalsUnderstanding of distributed systems concepts: consensus, leader election, distributed locking, eventual consistency, and the tradeoffs involvedProficiency in building and maintaining observability pipelines (Prometheus, Grafana, Loki, or equivalent) in productionComfortable working in a GitLab-based CI/CD workflowClear communicator who can document architectural decisions and explain technical tradeoffs ...

Cloud Infrastructure Engineer (Open LMS) UK, Remote

Location
United Kingdom
configuration management (etcd) Managing and tuning a multi-tier caching strategy (Varnish, Redis/Valkey, PHP OPcache) Running and scaling our observability stack (Prometheus, Grafana, Loki, Fluentd, PagerDuty) and participating in on-call rotations Evaluating and implementing distributed storage solutions as the platform evolves Improving deployment workflows and release processes … fundamentals Understanding of distributed systems concepts: consensus, leader election, distributed locking, eventual consistency, and the tradeoffs involved Proficiency in building and maintaining observability pipelines (Prometheus, Grafana, Loki, or equivalent) in production Comfortable working in a GitLab-based CI/CD workflow Clear communicator who can document architectural decisions and explain ...

Business Consultant - Databricks

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
Collaborate with data scientists and platform teams to improve query performance on Delta Lake. Develop automated monitoring and alerting mechanisms using Databricks REST APIs, Prometheus, or Azure Monitor. Maintain best practices for data engineering performance, testing, and documentation. Technical Skills Required Strong proficiency with Databricks, Apache Spark (SQL, PySpark, Scala … offs. Version control (Git, GitHub Actions, DevOps pipelines) and CI/CD practices for Databricks. Desirable Skills Experience with Databricks monitoring and observability (Ganglia, Prometheus, or Datadog). Understanding of serverless Databricks clusters. Familiarity with Terraform and infrastructure as code for Databricks resource management. Exposure to MLflow and model-serving ...

Senior DevOps Engineer - Hybrid Cloud & Automation

Location
England, United Kingdom
Clayton Black is seeking a Senior DevOps Systems Administrator to lead automation, design hybrid cloud infrastructure, and shape our deployment practices. You’ll work with AWS, VMware/Proxmox, Terraform, Ansible and Kubernetes to deliver ...

Backend Server Engineer (MMORPG) - GCP

Location
Greater London, England, United Kingdom
real-time multiplayer gameplay Work with GKE, Helm, and Kubernetes to manage deployments and autoscaling for stateful game servers Improve observability: monitoring dashboards, alerting (Prometheus/Grafana or GCP Cloud Monitoring), and incident response Scale distributed databases (MongoDB, PostgreSQL) under high concurrent load Optimise Kafka event streaming for player actions … experience scaling distributed systems under load. Worked with at least some of: MongoDB, PostgreSQL, Kafka, Redis, pub/sub systems Comfortable with monitoring stacks (Prometheus, Grafana, Alertmanager, or equivalent) Infrastructure-as-code experience (Terraform, Helm) Self-directed. Github/linkedin portfolio What Sets You Apart Experience with game backends ...

Senior Site Reliability Engineer (AWS / EKS)

Location
West of England, England, United Kingdom
Improve Kubernetes scaling and efficiency using technologies such as Karpenter, KEDA and native Kubernetes autoscaling capabilities. Build and improve observability using technologies such as Prometheus, Grafana, OpenTelemetry, Datadog and/or ELK. Support highly available distributed and event-driven systems, including environments using technologies such as Kafka/MSK. Design … Terragrunt. Production experience with Kubernetes delivery and GitOps practices; Argo CD or FluxCD strongly preferred. Strong production observability experience with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog or ELK. Experience operating highly available, distributed production systems. Strong understanding of AWS networking, IAM, security, availability and resilience. Experience implementing and testing ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
Using BGP, EVPN-VXLAN, JunOS (Juniper QFX/MX), OSPF, Spine-Leaf/IP Fabric, VLANs, VRFs, Linux networking, Ansible, Terraform, Prometheus/Grafana, Observability tooling The adventures that await you after becoming Senior Network Engineer at Hack The Box: Design and implement spine-leaf network architectures across multiple data … running JunOS Develop and maintain network automation using Ansible, Terraform or similar infrastructure‐as‐code tooling Establish and improve network monitoring, alerting and observability (Prometheus, Grafana, SNMP, streaming telemetry) Plan and execute network capacity upgrades, site bring‐ups and hardware refresh cycles Collaborate with the platform/systems engineering team ...

Software Engineer

Location
Uxbridge, England, United Kingdom
Full/part time : Full time Business area : Tech, Software Location : Uxbridge (can be remote with monthly office visits) A bit about giffgaff Do you want to join a connectivity provider that’s up to ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
building efficient, scalable databases and APIs (e.g. Django, FastAPI) a huge plus; Experience in Kubernetes, Docker, Spark and related monitoring tools (e.g. DataDog, Grafana, Prometheus) for DataOps a huge plus; Experience with Airflow a huge plus; Experience with dbt for pipeline modelling also beneficial; Skilled at shaping needs into … file formats on S3 for data lake storage Terraform for our infrastructure definition Kubernetes for data services and task orchestration Datadog/Grafana/Prometheus for platform monitoring Django for custom databases and frameworks Postgres/Aurora for our relational databases Notion for documentation We have a preference for candidates ...

Senior .NET Backend Developer

Hiring Organisation
Oscar Associates (UK) Limited
Location
York, North Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
Senior .NET Backend Developer Location: York (Hybrid - 2 days per week) Salary: Competitive + Benefits We're partnering with an established technology business undergoing a significant platform modernisation programme and are looking to appoint a ...