801 to 825 of 962 Grafana Jobs

Junior Platform Engineer Telemetry, SIEM, Observability

Hiring Organisation
apto solutions
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£40,000
customers find out the hard way. Our job is to make sure they never do . We work across Splunk, Cribl , Microsoft Sentinel, Grafana, Datadog and OpenTelemetry . We are a partner-led business with deep vendor relationships, and we build a lot of our own platform tooling on top. ...

Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)

Hiring Organisation
Pure Resourcing Solutions Limited
Location
Linton, Dry Drayton, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£55000 - £70000/annum
work with GPU or accelerator code, CUDA or similar Familiarity with profiling tools (Nsight, PyTorch Profiler) and ideally some exposure to monitoring stacks (Prometheus, Grafana) Strong Python for data work, Pandas and NumPy, genuine scripting ability Nice to have: exposure to inference serving frameworks like vLLM, published research, or open ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Birchanger, Hertfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 40,000 - 50,000 Annual
both technical and non-technical teams Desirable qualifications Microsoft certifications (AZ-900, AZ-104, AZ-305, AZ-500) or similar Experience with LogicMonitor admin, Grafana or other observability tools Familiarity with SRE concepts (SLIs, SLOs, error budgets) Understanding of ITIL processes Who are Solus? Solus, who are owned by Aviva ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, South East, United Kingdom
Employment Type
Permanent
Salary
£50,000
both technical and non-technical teams Desirable qualifications Microsoft certifications (AZ-900, AZ-104, AZ-305, AZ-500) or similar Experience with LogicMonitor admin, Grafana or other observability tools Familiarity with SRE concepts (SLIs, SLOs, error budgets) Understanding of ITIL processes Who are Solus? Solus, who are owned by Aviva ...

Senior Data Analyst - Content protection and discoverability

Hiring Organisation
Springer Nature
Location
Greater London, United Kingdom
Employment Type
Full Time
data security and governance Desirable Python or other scripting experience for data analysis and automation Experience with dashboarding and visualisation platforms such as Looker, Grafana or similar tools Exposure to security analytics, threat intelligence or fraud detection Understanding of SEO, scholarly discovery services, content metadata or digital content distribution ecosystems ...

Lead, Business Analysis

Location
Greater London, England, United Kingdom
metrics and benchmarks, plus experience with growth experiments, A/B testing, and funnel optimisation Deep SQL knowledge. Proficiency in product analytics tools (Mixpanel, Grafana, Tableau, Looker, etc) Deep understanding of data analytics, user behaviour analysis and segmentation Leadership & Soft Skills Strong communication skills to interact with teams, stakeholders ...

Support Manager

Location
Glasgow, Scotland, United Kingdom
errors are reused. Performance & improvement Own the team's SLA performance across response time, resolution time and customer satisfaction, monitoring through HubSpot, Jira and Grafana and acting early on risks. Produce weekly reporting on volumes, SLA performance, escalation trends and quality, and drive continuous improvement of workflows, escalation paths ...

Senior Product Manager, Observability New UK

Location
United Kingdom
surfaces logs, metrics, and traces at scale, and you understand the architectural and UX tradeoffs involved. Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry. Experience with deployment tooling in a data centre or infrastructure context, including provisioning workflows, networking automation, or zero-touch deployment pipelines. Experience building ...

Technical Support Lead

Location
Glasgow, Scotland, United Kingdom
errors are reused. Performance & improvement Own the team's SLA performance across response time, resolution time and customer satisfaction, monitoring through HubSpot, Jira and Grafana and acting early on risks. Produce weekly reporting on volumes, SLA performance, escalation trends and quality, and drive continuous improvement of workflows, escalation paths ...

Tech Lead, Autonomy Performance - Robotaxi

Location
Greater London, England, United Kingdom
defining release processes, gates, rollback plans, and regression handling Familiarity with modern data/observability stacks used in autonomy programs (e.g., Databricks/SQL, Grafana, internal eval/sim tooling, dashboarding/reporting) Track record of building scalable triage systems and performance programs (metrics taxonomy, prioritisation frameworks, quality bars ...

Customer Success Engineer

Location
Greater London, England, United Kingdom
with mobile SDKs, webviews, server-side integrations, cloud platforms, networking, or observability tools. Experience with data analysis or querying systems such as Snowflake, Looker, Grafana, or similar platforms. Experience designing customer-facing documentation, implementation playbooks, health checks, or technical enablement programs. Experience using automation or AI-assisted tools to improve ...

Engineering Manager, Runtime Platform, Robot Software

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
automotive, robotics, or another safety-relevant real-time domain. Familiarity with profiling toolchains (pprof/gperftools, perf, Nsight, NVLumo) and observability stacks (OpenTelemetry, Grafana, Datadog).Experience with NVIDIA (Orin/Thor) and/or Qualcomm compute platforms. Experience with a micro-kernel and/or real-time OS (e.g. ...

Lead Site Reliability Engineer

Location
Southampton, England, United Kingdom
with DevOps and engineering teams to establish and enforce SLOs, SLAs, and error budgets Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor. Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc. Developing bicep modules for monitoring infrastructure … Experience with infrastructure/configuration as code and version control (ARM, BICEP, Git) Strong Experience managing monitoring, alerting and dashboarding platforms (Azure Monitor, Prometheus, Grafana, Elasticsearch) Demonstrable experience of supporting live cloud services and platforms Expert in developing queries for dashboards and alerting for microservices. Collaborate with DevOps and engineering ...

Senior Mobile Engineer - Grafana Ops - IRM | Germany | Remote arbeitnow grafanalabs · 10/6/2026

Location
United Kingdom
Type: Remote Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. We are a 100% remote company with team members across 40+ countries, backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 70,000 Annual
private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join … certification desirable Reference: BBBH27029A DevOps, DevOps Engineer, AWS, Cloud Security, Cyber, DevSecOps, Terraform, Ansible, Linux, Networking, GitHub Actions, CI/CD, Bash, Python, Go, Grafana, Prometheus, Woking, Remote, Surrey, London If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV. ...

DevOps Engineer (AWS & Cloud Security)

Hiring Organisation
Ernest Gordon Recruitment Limited
Location
Camden, London, Camden Town, United Kingdom
Employment Type
Permanent
Salary
£65000 - £70000/annum + Remote + Progression
private cloud environments Automate infrastructure using Terraform and Ansible Build and maintain CI/CD pipelines using GitHub Actions Implement monitoring and observability using Grafana, Prometheus and CloudWatch Manage hybrid networking, IAM, firewalls and VPNs Improve infrastructure security, reliability and performance Support Kubernetes environments, including AWS EKS Join … certification desirable Reference: BBBH27029A DevOps, DevOps Engineer, AWS, Cloud Security, Cyber, DevSecOps, Terraform, Ansible, Linux, Networking, GitHub Actions, CI/CD, Bash, Python, Go, Grafana, Prometheus, Woking, Remote, Surrey, London If you're interested in this role, click 'apply now' to forward an up-to-date copy of your CV. ...

Senior Dev Ops Engineer

Location
Greater London, England, United Kingdom
clusters (GKE, manual Kubernetes, etc.) Extensive experience designing and building high‐availability, resilient infrastructure Experience setting up and maintaining monitoring, alerting, log collection (rsyslog, Grafana, Prometheus, ELK) Automation and scripting experience (Terraform, Packer, Ansible, Bash) Extensive Linux administration experience (Ubuntu LTS) Bachelor’s Degree in Computer Science or equivalent experience … solve problems, and work well under pressure Experience working independently and without direct supervision Willing to work outside standard business hours Desirable Experience with Grafana Loki Experience with Google, Amazon, Microsoft cloud platforms Experience with Python, Go languages Experience with MySQL, PostgreSQL databases Experience with CI/CD platforms (GitLab ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent
Salary
£85,000
service level indicators and error budgets. Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices. Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry. Develop custom application and platform metrics to improve operational visibility. Create advanced queries, dashboards and alerts for distributed microservices. Develop … ideally using AKS. Extensive experience in platform engineering, cloud provisioning and observability. Strong monitoring, alerting and dashboarding experience using technologies such as: Azure Monitor, Grafana, Prometheus, OpenTelemetry, Elasticsearch Experience creating custom metrics, queries, dashboards and alerts for microservices. Advanced scripting or software development skills using PowerShell, Python, C# ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Northam, Devon, UK
service level indicators and error budgets. Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices. Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry. Develop custom application and platform metrics to improve operational visibility. Create advanced queries, dashboards and alerts for distributed microservices. Develop … ideally using AKS. Extensive experience in platform engineering, cloud provisioning and observability. Strong monitoring, alerting and dashboarding experience using technologies such as: Azure Monitor, Grafana, Prometheus, OpenTelemetry, Elasticsearch Experience creating custom metrics, queries, dashboards and alerts for microservices. Advanced scripting or software development skills using PowerShell, Python, C# ...

DevOps and Automation Engineer (Contract)

Location
Greater London, England, United Kingdom
self-service environment provisioning using Infrastructure-as-Code (Terraform) and pipeline-driven automation.* Implement advanced observability and monitoring: Use platforms such as Datadog, Prometheus, Grafana, and OpenTelemetry to provide real-time insights into system health, deployments, and business metrics.* Embed security and compliance by design: Integrate security into every stage … generate new automation ideas and create user stories for rapid prototyping.* Use technologies like Terraform, Ansible, Azure DevOps, Github, OctopusDeploy, Kubernetes, OpenTelemetry, Datadog, Grafana, Prometheus, low-code automation platforms (e.g., Power Automate, UiPath) to help evolve the team's capabilities.* Coordinate planned outage and environment refreshes in collaboration with project ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
with third-party custodians. Build out Kubernetes and Docker deployments as we containerize more of the custody stack. Set up monitoring and alerting (Prometheus, Grafana, or equivalent) so we know about problems before customers do. Apply security controls and standards throughout - access control, key management, incident response. Provide on-call … with real production experience. A strong security focus - you've worked on systems where key management and access control are critical. Experience with Prometheus, Grafana, or an equivalent monitoring stack. Experience integrating with third-party custodians like BitGo or Fireblocks - this is a key differentiator for this role. What ...

Senior Software Engineer, Custody Services

Hiring Organisation
Robinhood Financial
Location
London, UK
Employment Type
Full-time
with third-party custodians. Build out Kubernetes and Docker deployments as we containerize more of the custody stack. Set up monitoring and alerting (Prometheus, Grafana, or equivalent) so we know about problems before customers do. Apply security controls and standards throughout — access control, key management, incident response. Provide on-call … with real production experience. A strong security focus — you've worked on systems where key management and access control are critical. Experience with Prometheus, Grafana, or an equivalent monitoring stack. Bonus pointsExperience integrating with third-party custodians like BitGo or Fireblocks — this is a key differentiator for this role. What ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues and ensuring service uptime. Conduct post‐incident reviews and help improve system resiliency … experience with AWS services and containerised applications Strong experience operating operational data stores (Aurora MySQL, DynamoDB). Expertise in using monitoring tools(e.g. Prometheus, Grafana, CloudWatch) for real‐time platform performance insights. Strong understanding of network security and Cloudflare, VPC and networking fundamentals, with a clear grasp of how traffic ...

Forward Deployed Engineer - Lead Platform Engineer

Hiring Organisation
Kyndryl
Location
London, UK
Employment Type
Full-time
guardrails, with security as a first-class concern (policy-as-code/OPA, access controls, secrets management, compliance-driven engineering) Instrument platforms for observability (Grafana, Prometheus, OpenTelemetry) Provide architectural oversight across multi-disciplinary workstreams, staying close enough to unblock the team directly Capture field learnings, codify reusable patterns and blueprints … access to endpoints Hands-on experience with CI/CD pipelines, Git-based workflows, and microservices/API architectures Practical experience with observability stacks (Grafana, Prometheus, OpenTelemetry) Experience with generative AI platforms: LLM hosting, LLM gateways (e.g. LiteLLM, Portkey, Kong AI Gateway), and MLOps/LLMOps practices for deploying ...

Staff Infrastructure Engineer (GCP) - Engine by Starling

Location
Manchester, England, United Kingdom
workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity Experience with automation using a scripting language like Python or Go Experience implementing CI/… native Container-based architecture Kubernetes (GKE on GCP, EKS on AWS) TeamCity for CI/CD (with multiple production releases per day) Terraform and Grafana RDS and CloudSQL for PostgreSQL Our Interview Process Interviewing is a two-way process and we want you to have the time and opportunity ...